Modular Cloud
Modular's hosted inference service — shared and dedicated frontier-model endpoints, and the same stack deployed into a customer's own VPC.
Find alternatives to Modular Cloud
Three editions on the pricing page: Modular Cloud billed 'PER TOKEN OR GPU HOUR', Bring Your Own Cloud billed 'PAY PER MINUTE', and Enterprise as a custom engagement. Sample published endpoint rates include DeepSeek V4 at $1.74 input / $3.48 output per 1M tokens and FLUX.2-dev at $10 per 1K tokens.
What it does
- Model inference
- Model hosting
Sources
- Pricing
-
Rates and billing units quoted from the page.
- Description
-
Official documentation · 18 Sep 2026
Editions, billing units and sample endpoint rates read from Modular's own pricing page, 2026-09-18; docs.modular.com's own heading for it is 'Deploy with Modular Cloud — A simple and powerful inference solution.'
- Sold within
-
Bought directly as metered inference, not as an add-on to MAX.
- URL
-
Fetched 2026-09-18, 200. Modular publishes no standalone marketing page for the cloud service — modular.com/platform returns 404 — so the pricing page, which names the editions, is the record's own page. Worth repointing if Modular ships one.
Also from Modular
-
MAX Infrastructure service
Open-source AI serving and modelling framework: a Python API, model pipelines and GPU kernels for NVIDIA, AMD and Apple hardware.
-
Mojo Developer tool
Apache-2.0 systems language for writing fast code across CPUs, GPUs and other accelerators without vendor lock-in, published by Modular.
Something wrong here? Send a correction — quote this product id: modular-cloud.