Kimi
Moonshot AI's family, most of it open-weight.
At a glance
Current release
From the release row Kimi K3. Uncached input; cache writes are billed separately ($3.00 for a 5-minute TTL, $6.00 for 1 hour).
Family record: Press release · 5 Sep 2026
Built on it
7 products in this catalogue name this family on their own pages.
-
Baseten Model APIs
Baseten
Official documentation · 5 Sep 2026 -
Cerebras Inference
Cerebras Systems
Official documentation · 5 Sep 2026 -
Fireworks AI
Fireworks AI
Official documentation · 5 Sep 2026 -
Kimi
Moonshot AI
Official documentation · 5 Sep 2026 -
Kimi Open Platform
Moonshot AI
Official documentation · 5 Sep 2026 -
Nebius Token Factory
Nebius Group
Official documentation · 5 Sep 2026 -
Vanchin
Kuaishou
Official documentation · 5 Sep 2026
Releases
-
Kimi k1.5
Multimodal reasoning model trained with long-context RL; the paper claims 77.5 on AIME and 96.2 on MATH-500, and 'long2short' methods beating GPT-4o and Claude 3.5 Sonnet.
-
Kimi-VL
16B MoE vision-language model with 2.8B active, MoonViT native-resolution encoder, 128K context and a Thinking variant under MIT; the lab reports 61.7 MMMU for the thinking model.
-
Kimi-Dev
Kimi-Dev-72B, an open coding model built on Qwen2.5-72B and optimised with large-scale RL to patch repositories; the lab claims 60.4% on SWE-bench Verified, MIT licence.
-
Kimi K2
1T-parameter MoE with 32B active, optimised for agentic tool use as a non-thinking model, released as K2-Base and K2-Instruct; the 0905 update (2025-09-05) brought 256K context.
-
Kimi K2 Thinking
Thinking agent on K2 that reasons step by step across 200-300 sequential tool calls, 256K context, native INT4 QAT; claims 44.9% HLE with tools and 71.3% SWE-bench Verified.
-
Kimi K2.5
Native multimodal agentic model (1T total / 32B active, MoonViT encoder, 256K context) with thinking and instant modes and an agent swarm of up to 100 sub-agents; Modified MIT.
-
Kimi K2.6
Open-sourced update focused on long-horizon coding, agent swarm scaled to 300 sub-agents and better tool calling; the lab reports 58.6% SWE-Bench Pro and 66.7% Terminal-Bench 2.0.
-
Kimi K3
2.8T-parameter model (104B active) with native vision and 1M-token context, max thinking effort by default; announced 2026-07-16, weights promised by 2026-07-27 (Kimi K3 License).