Cogito
Cogito is Decart's OpenAI-compatible LLM inference API, serving frontier open-weight models including Kimi K2.6, Kimi K2.7 Code, GLM-5.2, Qwen3 235B and GPT-OSS 120B across Trainium, TPU and GPU capacity. A self-serve Standard tier is billed per token, while an ultra-fast reserved tier advertises 1,000+ tokens per second. It is built on Decart's own DOS optimization stack.
Metered per-token rate card, Standard tier, verified twice (researching pass and independent verification pass, both 2026-09-05): Kimi K2.6 $0.66/1M input, $0.144/1M cached input, $3.41/1M output; Kimi K2.6 Fast $0.68 / $0.144 / $3.41; GLM-5.2 $1.20 / $0.20 / $4.00. GPT-OSS 120B, Qwen3 235B and Kimi K2.7 Code are listed as 'Contact sales'. Standard is 'self-serve with no contracts or minimums' and new users get '$5 in starter credits to try Standard, no card'. The ultra-fast 1,000+ tps tier is 'priced per workload rather than per token' and requires contacting sales - this is why it is bought on its own rather than inside the Decart API Platform.
What it does
- Model inference
- Model hosting
- Text generation
Sources
- Pricing
-
Per-token table for the Standard tier; ultra-fast tier 'priced per workload'; '$5 in starter credits to try Standard, no card'.
- Description
-
Company disclosure · 5 Sep 2026
Title 'Cogito | Ultra-fast Inference'. 'Cogito is Decart's LLM inference API: frontier models at ultra-fast speed, on capacity the GPU crunch can't reach.' Page credits throughput to DOS (Decart Optimization Stack). Independently re-fetched by the verification pass 2026-09-05. Note this product is NOT mentioned anywhere on decart.ai - which is why the directory was missing it.
- What it does
-
Official documentation · 5 Sep 2026
'Cogito is OpenAI-compatible. Use the OpenAI SDK in any language; just point base_url' at https://api.cogito.decart.ai/v1.
- Sold within
-
cogito.decart.ai/pricing (already cited pricing-page elsewhere on this record); self-serve signup with public per-token pricing.
Also from Decart
-
Decart API Platform Model API
Hosted API for Decart's real-time video and world models.
-
Decart Optimization Stack Infrastructure service
The Decart Optimization Stack (DOS) is Decart's inference and training optimization infrastructure, spanning hardware-aware model design, kernel tooling, proprietary compilers and inference optimization. It is sold to hardware providers and AI teams as engagements covering benchmark optimization, customer-defined kernel and compiler work, cross-workload efficiency gains, and profiler and simulator licensing, and is marketed as hardware-agnostic across GPUs, TPUs, Trainium and AMD accelerators.
-
Delulu (by Decart) Application
Delulu is Decart's consumer mobile app for AI photo transformation, published on iOS and Android by Decart.AI, Inc. Its store listing describes it as a way to 'Turn any photo into something fun, weird, or just really good-looking, with a single tap', browsing preset styles and creating stickers to share.
-
Lucy Application
Real-time video editing platform that edits and transforms live video at streaming speed, for enterprises, brands, creators and live/streaming/gaming use cases; built on Decart's DOS infrastructure.
-
Oasis 3 API service
Interactive world model for Physical AI, generating realistic, controllable real-time simulation environments for training autonomous systems, distributed as an API/SDK product built on Decart's Optimization Stack (DOS) infrastructure.
Something wrong here? Send a correction — quote this product id: decart-cogito.