Together Inference
Hosted API serving open-weight text, image and audio models, billed per token.
Find alternatives to Together Inference
What it does
- Model inference
How it compares
-
Modal
Specialized AI platforms like Modal, Together.ai, and Fireworks.ai
Official documentation · 19 Sep 2026
Sources
- Pricing
- Sold within
-
Together's pricing page publishes a per-model rate card headed "Price per 1M tokens", and the serverless inference page says to "pay only for the tokens you use".
- Deployment
-
Inferred 2026-09-08, not read off a page: every already-decided product of type 'model-api' in this catalogue includes "saas" in its deployment list, zero exceptions, per item 198's measured cross-tab. May be incomplete if this product also offers a self-hosted or private-cloud option; not wrong either way.
- Interfaces
-
Official documentation · 29 Sep 2026
Page states: 'Same API, better models. No code changes required. Drop in your API key and access hundreds of open-source models through the same interface you're already using', with a Python code sample against the API.
Also from Together AI
-
Together GPU Clusters Infrastructure service
Reserved NVIDIA GPU clusters for training and large-scale inference.
-
Together Fine-Tuning API service
A managed service for fine-tuning open-source models on a customer's own data and serving the result on Together's infrastructure.
-
Together Batch Inference API service
Asynchronous bulk inference for workloads that do not need a real-time response, priced below Together's serverless rate.
-
Together Custom Training Developer tool
Custom model training service covering supervised fine-tuning and direct preference optimization, billed per token.
-
Together Dedicated Container Inference Infrastructure service
Dedicated, reserved GPU containers for model inference with guaranteed performance, billed per GPU-hour.
Open Together AI in the directory
Something wrong here? Send a correction — quote this product id: together-inference.