Baseten Model APIs
Pre-optimised hosted endpoints for open-source frontier models, called without deploying or managing a deployment first.
Find alternatives to Baseten Model APIs
Per-million-token rates are now shown directly on the product page for specific models, e.g. GLM-5.3 $1.40 input / $4.40 output, DeepSeek-V4-Flash-0731 $0.13 input / $0.26 output, Kimi K3 $3.00 input / $15.00 output.
What it does
- Model inference
- Model hosting
How it compares
-
Kimi
latest 1.6T-parameter open frontier model from DeepSeek AI Model API Kimi K3 The open frontier model. 2.8T parameters, 1M-token context, top be
Official documentation · 5 Sep 2026 -
Modal Inference
Modal differentiates through sub-second cold starts and pure Python integration, while competitors like Baseten rely on REST APIs or YAML configurations.
Research report · 19 Sep 2026
Sources
- Pricing
-
Per-million-token rates shown directly on the product page for named models.
- Sold within
-
Baseten's pricing page publishes per-model Model API token rates with a "Try Model API" control (for example "GLM-5.3" at "$1.40" input and "$4.40" output), and lists "Model APIs" under "Included in Basic:", the "$0 per month, pay as you go" plan.
- Deployment
-
Inferred 2026-09-08, not read off a page: every already-decided product of type 'model-api' in this catalogue includes "saas" in its deployment list, zero exceptions, per item 198's measured cross-tab. May be incomplete if this product also offers a self-hosted or private-cloud option; not wrong either way.
Also from Baseten
-
Baseten Platform
Deploys and serves machine-learning models as production endpoints on managed GPU infrastructure.
-
Baseten Training Platform
Trains and fine-tunes models with reinforcement learning through the Loops SDK, deploying the result onto Baseten's inference stack.
-
Baseten Frontier Gateway API service
Production-grade API launch platform for model labs: takes a lab's model from research to a reliable, scalable, white-labelled API in days, distinct from Baseten's Distribution Platform (model marketplace listing).
-
Baseten Inference Runtime Infrastructure service
Named inference runtime (automatic TensorRT/SGLang/vLLM builds, speculative decoding, custom kernel fusion, KV-cache optimisation) that underlies Baseten's Dedicated Inference, Model APIs and Training products.
Something wrong here? Send a correction — quote this product id: baseten-model-apis.