DeepSeek
DeepSeek's family, most of it open-weight.
At a glance
Current release
From the release row DeepSeek-V4. Peak-hour rate, cache miss; off-peak is 50% lower ($0.66 / $1.98). V4-Pro left preview 2026-08-13.
Family record: Press release · 5 Sep 2026
Built on it
9 products in this catalogue name this family on their own pages.
-
Baseten Model APIs
Baseten
Official documentation · 5 Sep 2026 -
DeepSeek
DeepSeek
Official documentation · 5 Sep 2026 -
DeepSeek API Platform
DeepSeek
Official documentation · 5 Sep 2026 -
Fireworks AI
Fireworks AI
Official documentation · 5 Sep 2026 -
MaaS
Huawei Cloud
Official documentation · 5 Sep 2026 -
Nebius Token Factory
Nebius Group
Official documentation · 5 Sep 2026 -
SambaCloud
SambaNova
Official documentation · 5 Sep 2026 -
Together Inference
Together AI
Official documentation · 5 Sep 2026 -
Vanchin
Kuaishou
Official documentation · 5 Sep 2026
Releases
-
DeepSeek Coder
Code models from 1B to 33B trained from scratch on 2T tokens (87% code), 16K window for project-level completion and infilling; commercial use permitted under the Model License.
-
DeepSeek LLM
First general model: 7B and 67B base and chat trained on 2T tokens; the lab reports 67B Chat at 73.78 HumanEval and 84.1 GSM8K, commercial use permitted.
-
DeepSeek-V2
236B MoE with 21B active, 128K context, MLA attention and DeepSeekMoE; the lab claims 42.5% lower training cost and 93.3% smaller KV cache than DeepSeek 67B; V2-Lite on 2024-05-16.
-
DeepSeek-Coder-V2
Code MoE continued from V2 on 6T more tokens, 16B-Lite and 236B variants, 338 languages, 128K context; the paper claims performance comparable to GPT4-Turbo on code tasks.
-
DeepSeek-R1-Lite-Preview
First reasoning model, live on chat.deepseek.com with a visible thought process; the lab claims o1-preview-level AIME and MATH results, with open weights and API 'coming soon'.
-
DeepSeek-V3
671B MoE with 37B active pre-trained on 14.8T tokens, 60 tokens/s (3x V2), API-compatible with V2.5; fully open weights and paper on GitHub.
-
DeepSeek-R1
Reasoning model with MIT-licensed weights; the lab claims performance on par with OpenAI-o1 in math, code and reasoning, plus six distilled models, 32B and 70B on par with o1-mini.
-
DeepSeek-V3.1
Hybrid inference: one model with Think and Non-Think modes, 128K context on both API endpoints, post-training for tool use and agents; Base and instruct weights on Hugging Face.
-
DeepSeek-V3.2
Successor to V3.2-Exp (2025-09-29, DeepSeek Sparse Attention debut); first DeepSeek model with thinking inside tool use; API-only V3.2-Speciale claims IMO and IOI 2025 gold level.
-
DeepSeek-V4
Preview of V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B), 1M context now default across DeepSeek services, open weights; claims open-source SOTA in agentic coding.
-
DeepSeek-V4-Flash-Vision-Exp
Experimental multimodal model keeping V4-Flash text ability, image input via base64, URL or the new Files API at up to 384 tokens per image; claims a big leap on multimodal agents.