Serverless
Autoscaling GPU API endpoints for AI inference, billed per second with scale-to-zero and sub-200ms cold starts.
Find alternatives to Serverless
Per-second GPU billing; published tiers run from $0.58/hr (A4000/A4500/RTX 4000/RTX 2000 class) to $9.98/hr (B300). H100 $4.79/hr, A100 $2.72/hr.
What it does
- GPU cloud
- Model inference
- Model hosting
How it compares
-
Modal
Runpod vs Modal: Python-native serverless versus portable GPU compute
Official documentation · 19 Sep 2026
Sources
- Pricing
-
Fetched 2026-09-06; page content returned. Numeric HTTP status unavailable this run — see the batch's method note. Serverless table quotes 'B300: $9.98/hr', 'H100: $4.79/hr', 'A100: $2.72/hr', 'A4000/A4500/RTX 4000/RTX 2000: $0.58/hr'; billing 'per second, metered from worker start to full stop'.
- Description
-
Official documentation · 6 Sep 2026
Fetched 2026-09-06; page content returned. Numeric HTTP status unavailable this run — see the batch's method note. Headline: 'Dedicated Serverless GPU API endpoints'; body: 'Runpod Serverless runs AI inference with sub-200ms FlashBoot cold starts, per-second billing, and scale to zero.'
- Sold within
-
runpod.io/pricing (already cited pricing-page elsewhere on this record); own top-level pricing entry.
Also from RunPod
-
Pods Infrastructure service
Per-hour GPU pods and per-hour serverless endpoints across both datacentre accelerators and consumer cards, sold on price — the company's own claim is compute up to 90% below traditional cloud providers.
-
Runpod Clusters Infrastructure service
Multi-node GPU environments with high-speed InfiniBand interconnect for distributed training and large batch workloads.
-
Runpod Hub Developer tool
A catalog of templates, models and open-source AI apps that can be forked and deployed onto Runpod Serverless in one click.
-
Public Endpoints Model API
Instant API access to pre-deployed third-party AI models for image, video, audio and text generation, billed per request or per token with no infrastructure setup.
-
Runpod Hybrid Cloud Platform
Brings customer-owned or rented GPU hardware under Runpod's console, CLI and APIs as a single control plane, with Runpod cloud used for overflow capacity.
Something wrong here? Send a correction — quote this product id: runpod-serverless.