Nebius Token Factory
Managed inference endpoint serving open-weight models on Nebius hardware, billed per token.
Find alternatives to Nebius Token Factory
Per-token pricing with no seat fee; the floor is 0 because a caller pays only for tokens used.
What it does
- Model inference
- Model hosting
How it compares
-
Kimi
Top open-source models available Text and multimodal DeepSeek-V4-Pro Kimi-K3 GLM-5.1 MiniMax-M2.5 Nemotron-3-Ultra INTELLECT-3 Qwen3-235B-A22B
Official documentation · 5 Sep 2026
Sources
- Pricing
- Sold within
-
Official documentation · 13 Sep 2026
Fetched in a real browser 2026-09-13: tokenfactory.nebius.com is a live, distinct signup and billing surface (own workspace, own "Get API key"/log-in flow, own model catalog) separate from nebius.com's own AI Cloud pricing -- a buyer does not have to purchase Nebius AI Cloud compute to use Token Factory.
Also from Nebius Group
-
Nebius AI Cloud Infrastructure service
Rented GPU clusters with storage and networking for training and serving models, billed by the GPU-hour.
-
Managed Soperator Infrastructure service
Nebius-managed Kubernetes operator that runs Slurm clusters on GPU infrastructure for fault-tolerant large-scale AI training, with topology-aware scheduling and automatic node health checks and recovery.
-
Serverless AI Infrastructure service
Nebius AI Cloud's on-demand GPU runtime that runs containerised AI workloads as Jobs and hosts custom models behind HTTP Endpoints without provisioning or managing clusters.
-
Managed Service for MLflow Developer tool
Fully managed MLflow deployment on Nebius AI Cloud for tracking experiments, metrics and artifacts across the machine-learning lifecycle without maintaining tracking-server infrastructure.
Open Nebius Group in the directory
Something wrong here? Send a correction — quote this product id: nebius-token-factory.