Baseten on Hugging Face Inference Providers ๐ฅ
Baseten, an AI infrastructure platform, has joined the Hugging Face Hub as a supported Inference Provider. This integration enhances the Hub's serverless inference capabilities and makes it easier for developers to use a wide variety of models directly from model pages. Baseten offers support for a broad spectrum of model types, including LLMs and text-to-speech models, such as Kimi K3, DeepSeek V4 Flash, GLM-5.2, and more.
Users can access these models by utilizing the Hugging Face SDKs (huggingface_hub = 1.26.1 for Python and @huggingface/inference for JavaScript). The integration is also available in various Agent Harnesses, allowing seamless use of baseten-hosted models in tools like Pi, OpenCode, Hermes Agents, and OpenClaw. Inference requests are automatically routed to Baseten, and users are billed by the corresponding provider based on their usage.
PRO users receive $2 worth of Inference credits monthly, which can be used across providers. The Hugging Face Hub offers free inference with a small quota for signed-in free users, but upgrading to PRO unlocks additional benefits like ZeroGPU, Spaces Dev Mode, higher limits, and more. Feedback is welcomed at https://huggingface.co/spaces/huggingface/HuggingDiscussions/discussions/49.
Written by urgent.news from Hugging Face's reporting โ not their text. Machine-written โ may contain errors; check the original before relying on it.