Together Dedicated Container Inference
Dedicated, reserved GPU containers for model inference with guaranteed performance, billed per GPU-hour.
Find alternatives to Together Dedicated Container Inference
Per-GPU-hour billing, from $3.99/hr (H100 promotional) to $8.99/hr on-demand; reserved options available. Recurring free access is not stated either way.
What it does
- Model hosting
- Model inference
Sources
- Pricing
- Description
- What it does
- Sold within
-
Own published per-GPU-hour rate card confirms standalone purchase.
- Name
- URL
Also from Together AI
-
Together GPU Clusters Infrastructure service
Reserved NVIDIA GPU clusters for training and large-scale inference.
-
Together Inference Model API
Hosted API serving open-weight text, image and audio models, billed per token.
-
Together Fine-Tuning API service
A managed service for fine-tuning open-source models on a customer's own data and serving the result on Together's infrastructure.
-
Together Batch Inference API service
Asynchronous bulk inference for workloads that do not need a real-time response, priced below Together's serverless rate.
-
Together Custom Training Developer tool
Custom model training service covering supervised fine-tuning and direct preference optimization, billed per token.
Open Together AI in the directory
Something wrong here? Send a correction — quote this product id: together-dedicated-inference.