Qdrant Cloud Inference

API service

Embedding generation run inside Qdrant Cloud so text or images are vectorised and searched in a single API call.

Find alternatives to Qdrant Cloud Inference

TypeAPI service
RoleBundled component
AvailabilitySold
DeploymentSaaS
Intended for—
SecurityNot recorded
StatusActive
Sold withinQdrant Cloud — Included with the parentOfficial documentation · 18 Sep 2026
PricingUsage-basedPricing page · 18 Sep 2026
Sources1 of 1 field

Enabled by default on new Qdrant Cloud clusters at no extra cost; inference is then billed per token at a fixed per-model rate, with free models free and paid Cloud accounts carrying an allowance of up to 5 million tokens per model per month. No per-model rate card is printed on the page.

What it does

qdrant.tech

Sources

Pricing

Pricing page · 18 Sep 2026

Token billing, the free-model carve-out and the 5-million-token allowance quoted from the page.

Description

Official documentation · 18 Sep 2026

The page's own words: 'Embed and Search in One API Call. Qdrant Cloud runs inference alongside your vector search, so you can simplify your data pipeline.'

Sold within

Official documentation · 18 Sep 2026

The page states it runs inside Qdrant Cloud and is enabled by default on new clusters — a Qdrant Cloud account is what a buyer has.

URL

Official documentation · 18 Sep 2026

Fetched 2026-09-18, 200; listed in Qdrant's own Products nav.

Also from Qdrant

Open Qdrant in the directory

Something wrong here? Send a correction — quote this product id: qdrant-cloud-inference.