Lambda Inference
Hosted inference endpoints for open-weight models on Lambda's own GPU fleet, billed per token.
Find alternatives to Lambda Inference
Token-metered. The inference page carries no figure in the fetched HTML; the rate card is on lambda.ai/pricing, which an applying session should read in a browser before writing a number.
What it does
- Model inference
- Model hosting
Sources
- Pricing
-
Token-metered. The inference page carries no figure in the fetched HTML; the rate card is on lambda.ai/pricing and needs a browser read (JS-rendered) before a number can be written.
- Description
-
Official documentation · 5 Sep 2026
Lambda's own product nav and its /inference page, fetched 2026-09-05, HTTP 200.
- Sold within
-
lambda.ai/pricing (already cited pricing-page elsewhere on this record); nothing packages it inside the cloud subscription.
Also from Lambda
-
Lambda Cloud Infrastructure service
On-demand NVIDIA instances for training, fine-tuning and serving — 1 to 8 GPUs launched in minutes with self-serve access — billed by the minute with no egress charge, alongside the 1-Click Clusters and liquid-cooled superclusters Lambda sells for larger runs.
-
Lambda 1-Click Clusters Infrastructure service
Self-serve GPU clusters from 16 to 2,000+ NVIDIA B200 or H100 GPUs, provisioned without a sales cycle for training, fine-tuning and large inference runs.
-
Lambda On-Demand Instances Infrastructure service
Hourly NVIDIA GPU instances — H100, H200, B200 and A100 — launched in minutes with no egress fees.
-
Lambda Superclusters Infrastructure service
Rents dedicated large-scale AI training and inference GPU clusters at supercomputer scale.
Something wrong here? Send a correction — quote this product id: lambda-inference.