Crusoe Managed Inference
Hosted endpoints serving open-weight and customer-supplied models on Crusoe's own hardware, billed per token or as dedicated hourly deployments.
Find alternatives to Crusoe Managed Inference
Serverless per-1M-token rates vary by model with separate input, output and cached-token tiers (e.g. Llama 3.3 70B $0.25 input, DeepSeek V3 $0.50 input, GLM 5.2 $1.40 input). Self-Serve Deployments are billed hourly: NVIDIA H100 80GB HGX $5.50/hr, H200 141GB HGX $6.00/hr. Provisioned throughput is transacted via AI Model Units (AMUs) on commitment - contact sales.
What it does
- Model inference
- Model hosting
Sources
- Pricing
- Sold within
-
Official documentation · 6 Sep 2026
https://www.crusoe.ai/developers - 'Fine-tune and serve models in Crusoe Intelligence Foundry with $5 in free credits.'
Also from Crusoe
-
Crusoe Cloud Infrastructure service
GPU cloud running on Crusoe's own energy-colocated data centres, rented by the hour.
-
Serverless Fine-Tuning API service
Managed LoRA-based fine-tuning of open-weight models, billed per token of training data rather than per GPU-hour.
-
Crusoe Edge Zones Infrastructure service
Crusoe Cloud GPU capacity delivered at edge and sovereign locations via prefabricated Crusoe Spark modular units.
-
Crusoe Spark Hardware
Prefabricated modular AI data centre integrating power, cooling, fire suppression and GPU racks for deployment at a customer site.
-
Crusoe Command Center Developer tool
Unified operations and observability console for GPU fleets running AI workloads on Crusoe Cloud.
Something wrong here? Send a correction — quote this product id: crusoe-managed-inference.