Baseten Inference Runtime
Named inference runtime (automatic TensorRT/SGLang/vLLM builds, speculative decoding, custom kernel fusion, KV-cache optimisation) that underlies Baseten's Dedicated Inference, Model APIs and Training products.
Find alternatives to Baseten Inference Runtime
Not sold on its own; usage draws from whichever Baseten product (Dedicated Inference, Model APIs, Training) it runs underneath. Whether a recurring free tier exists is not stated either way.
What it does
- Model inference
Sources
- Description
- What it does
- Sold within
-
Official documentation · 14 Sep 2026
Page states it is 'the technical foundation... supporting Baseten's broader platform offerings -- including Dedicated Inference, Model APIs, and Training'.
- Name
- URL
Also from Baseten
-
Baseten Platform
Deploys and serves machine-learning models as production endpoints on managed GPU infrastructure.
-
Baseten Model APIs Model API
Pre-optimised hosted endpoints for open-source frontier models, called without deploying or managing a deployment first.
-
Baseten Training Platform
Trains and fine-tunes models with reinforcement learning through the Loops SDK, deploying the result onto Baseten's inference stack.
-
Baseten Frontier Gateway API service
Production-grade API launch platform for model labs: takes a lab's model from research to a reliable, scalable, white-labelled API in days, distinct from Baseten's Distribution Platform (model marketplace listing).
Something wrong here? Send a correction — quote this product id: baseten-inference-runtime.