Inference Endpoints
Deploys a model from the Hub to a dedicated managed endpoint, billed by the hour.
Find alternatives to Inference Endpoints
Free Hub access, PRO at $9 a month, Team at $20 and Enterprise at $50 per user, plus hourly compute for Spaces and Inference Endpoints.
What it does
- Model inference
Sources
- Pricing
-
Free Hub access, PRO at $9 a month, Team at $20 and Enterprise at $50 per user, plus hourly compute for Spaces and Inference Endpoints.
- Sold within
-
Its own hourly rate card on the pricing page: "Dedicated inference, starting at $0.033/hour.", followed by per-instance CPU/accelerator/GPU prices. Docs (huggingface.co/docs/inference-endpoints/pricing) add a prerequisite that does not name a plan: "🤗 Inference Endpoints is accessible to Hugging Face accounts with an active subscription and credit card on file."
Also from Hugging Face
-
Hugging Face Hub Platform
Hosts and versions open models, datasets and demo applications, publicly or privately.
-
Spaces Platform
Hosts runnable demo applications for models on free or paid hardware.
-
HuggingChat Application
Chat app powered by open-source AI models with an Omni router that automatically selects the most suitable model, or lets users pick directly from 140+ open models.
-
Inference Providers API service
Marketplace connecting users to multiple third-party inference providers for hosted AI models, billed per input/output token with provider- and model-specific rates shown side by side.
-
Hugging Face Jobs Infrastructure service
Pay-as-you-go compute for running AI training, fine-tuning, synthetic data generation and batch inference jobs on Hugging Face's own CPU/GPU/TPU infrastructure via CLI or Python API.
Open Hugging Face in the directory
Something wrong here? Send a correction — quote this product id: huggingface-inference-endpoints.