Who’s winning the AI model race?

Every frontier model from the 15 labs that set the pace, on one timeline — dated from the lab’s own announcement, one source per row. Size tiers, snapshots and point updates fold into the generation they belong to. This page does not pick a winner — it shows who shipped what, and when, so you can judge the race yourself.

  1. 2018JunGPT-1NovBERT
  2. 2019FebGPT-2AprERNIE
  3. 2020FebT5MayGPT-3
  4. 2021MayLaMDAJulERNIE 3.0
  5. 2022AprPaLMAugGLM-130B
  6. 2023FebLLaMAMarClaudeMarGPT-4MarChatGLM-6BMarERNIE BotMayPaLM 2JunChatGLM2-6BJulClaude 2JulLlama 2AugQwen-7BAugQwen-VLAugCode LlamaSepMistral 7BOctERNIE 4.0OctChatGLM3-6BNovDeepSeek CoderNovGrok-1NovDeepSeek LLMDecGemini 1.0DecMixtral 8x7B
  7. 2024JanGLM-4FebQwen1.5FebGemini 1.5FebGemmaFebMistral LargeMarClaude 3MarCommand RMarGrok-1 (open release)MarGrok-1.5AprCommand R+AprGrok-1.5VAprLlama 3MayDeepSeek-V2MayGPT-4oMayCodestralJunQwen2JunDeepSeek-Coder-V2JunClaude 3.5 SonnetJulLlama 3.1JulMistral Large 2AugEXAONE 3.0AugGrok-2Sepo1SepPixtral 12BSepQwen2.5SepLlama 3.2NovDeepSeek-R1-Lite-PreviewNovQwQ-32B-PreviewDecEXAONE 3.5DecGemini 2.0DecCommand R7BDecDeepSeek-V3
  8. 2025JanMiniMax-Text-01JanDeepSeek-R1JanKimi k1.5JanQwen2.5-MaxJanMistral Small 3FebGrok 3FebClaude 3.7 SonnetMarCommand AMarERNIE 4.5MarERNIE X1MarEXAONE DeepMarGemini 2.5AprLlama 4AprKimi-VLAprGPT-4.1AprGLM-Z1-32B-0414Apro3AprQwen3MayMistral Medium 3MayClaude 4JunMagistralJunMiniMax-M1JunKimi-DevJunERNIE 4.5 (open-source family)JulStep 3JulGrok 4JulKimi K2JulEXAONE 4.0JulQwen3-CoderJulGLM-4.5JulCommand A VisionAuggpt-ossAugGPT-5AugCommand A ReasoningAugDeepSeek-V3.1SepQwen3-NextSepGrok 4 FastSepQwen3-VLSepClaude Sonnet 4.5SepGLM-4.6OctMiniMax-M2NovKimi K2 ThinkingNovERNIE 5.0NovGrok 4.1NovGemini 3NovClaude Opus 4.5DecDeepSeek-V3.2DecMistral 3DecDevstral 2DecGLM-4.7
  9. 2026JanKimi K2.5FebMiniMax-M2.5FebClaude Opus 4.6FebStep 3.5 FlashFebGLM-5FebQwen3.5FebGemini 3.1 ProMarGrok 4.20MarGPT-5.4MarMistral Small 4MarMiniMax-M2.7AprGemma 4AprGLM-5.1AprMuse SparkAprEXAONE 4.5AprClaude Opus 4.7AprKimi K2.6AprDeepSeek-V4MayStep 3.7 FlashMayERNIE 5.1MayGemini 3.5 FlashMayCommand A+MayMistral Medium 3.5JunMiniMax-M3JunClaude Fable 5 and Claude Mythos 5JunGLM-5.2JunClaude Sonnet 5JulMuse Spark 1.1JulGPT-5.6JulKimi K3JulGrok 4.5JulClaude Opus 5JulK-EXAONE 2.0AugMuse GlimmerAugQwen3.8AugGrok 4.6AugGLM-5.3AugDeepSeek-V4-Flash-Vision-ExpSepClaude Fable 5.1 and Claude Mythos 5.1SepGPT-6 AstraSepStep 5 Preview

flagship generation other headline release · every name links to its dated row below

OpenAI

13 releases · 2018 → 2026

GPT for the flagship line (GPT-1 to GPT-6 Astra); o-series for reasoning models (o1, o3), merged into GPT-5 in 2025; gpt-oss for open weights.

201820192020202120222023202420252026 GPT-1 · Jun 2018GPT-2 · Feb 2019o1 · Sep 2024o3 · Apr 2025gpt-oss · Aug 2025GPT-3 · May 2020GPT-4 · Mar 2023GPT-4o · May 2024GPT-4.1 · Apr 2025GPT-5 · Aug 2025GPT-5.4 · Mar 2026GPT-5.6 · Jul 2026GPT-6 Astra · Sep 2026
  1. GPT-1ResearchOpen weights

    Code and model for the paper 'Improving Language Understanding by Generative Pre-Training', the first generative pre-trained Transformer; the paper itself carries no date.

    Official documentation
  2. GPT-2ResearchOpen weights

    Paper 'Language Models are Unsupervised Multitask Learners' with a staged release: only a smaller model was made public at launch over misuse concerns.

    News reporting
  3. GPT-3FlagshipResearch

    Paper 'Language Models are Few-Shot Learners': 175 billion parameters, 10x any previous non-sparse model, learning tasks from text demonstrations without fine-tuning.

    Research report
  4. GPT-4FlagshipMultimodal

    Accepts image and text input; OpenAI said it passes a simulated bar exam around the top 10% of test takers. ChatGPT (GPT-3.5) had launched Nov 2022; GPT-4 Turbo followed Nov 2023.

    News reporting
  5. GPT-4oFlagshipMultimodal

    Described by OpenAI as 'our fastest and most affordable flagship model'; GPT-4o mini, a small model for lightweight tasks, followed on Jul 18, 2024.

    Official documentation
  6. o1Reasoning

    o-series, not GPT-numbered: o1-preview and o1-mini, 'trained with reinforcement learning' to reason before answering; the full o1 reached the API on Dec 17, 2024.

    Official documentation
  7. GPT-4.1Flagship

    gpt-4.1, gpt-4.1-mini and gpt-4.1-nano added to the API; preceded by the GPT-4.5 research preview released Feb 27, 2025.

    Official documentation
  8. o3Reasoning

    o-series: o3 and o4-mini launched together in the API; o3-mini had preceded them on Jan 31, 2025.

    Official documentation
  9. gpt-ossOpen weightsReasoning

    gpt-oss-120b (117B params, 5.1B active) and gpt-oss-20b (21B, 3.6B active), Apache 2.0, text-only reasoning models with adjustable reasoning effort and 128K context.

    News reporting
  10. GPT-5FlagshipReasoning

    gpt-5, gpt-5-mini and gpt-5-nano in the API; 400K context, text and image in. GPT-5 Pro (Oct 6), GPT-5.1 (Nov 13) and GPT-5.2 (Dec 11, 2025) were later flagships in the family.

    Official documentation
  11. GPT-5.4Flagship

    'Our newest frontier model for professional work' with a 1,050,000-token context; GPT-5.4 mini and nano followed Mar 17 and GPT-5.5 / GPT-5.5 Pro on Apr 24, 2026.

    Official documentation
  12. GPT-5.6Flagship

    A named family rather than tiers: GPT-5.6 Sol (frontier), Terra and Luna, plus GPT-5.6 Cyber; Sol has a 1,050,000-token context with text and image input.

    Official documentation
  13. GPT-6 AstraFlagshipReasoning

    'Our most capable model, built for the hardest end-to-end work'; 1,050,000-token context, reasoning effort low to max, rolling out first to Trusted Access Program enterprises.

    Official documentation
Anthropic

14 releases · 2023 → 2026

Claude for every model: numbered generations (Claude 1-4), then tier-first names (Opus, Sonnet, Haiku) with a version, and from June 2026 the Fable/Mythos 5 line above Opus.

201820192020202120222023202420252026 Claude · Mar 2023Claude 2 · Jul 2023Claude 3 · Mar 2024Claude 3.5 Sonnet · Jun 2024Claude 3.7 Sonnet · Feb 2025Claude 4 · May 2025Claude Sonnet 4.5 · Sep 2025Claude Opus 4.5 · Nov 2025Claude Opus 4.6 · Feb 2026Claude Opus 4.7 · Apr 2026Claude Fable 5 and Claude Mythos 5 · Jun 2026Claude Sonnet 5 · Jun 2026Claude Opus 5 · Jul 2026Claude Fable 5.1 and Claude Mythos 5.1 · Sep 2026
  1. ClaudeFlagship

    First public Claude, offered as Claude (high-performance) and Claude Instant (lighter, cheaper, faster), for summarization, search, writing, Q&A and coding.

    Press release
  2. Claude 2Flagship

    100K-token input, launched with the claude.ai website; Anthropic reported 76.5% on the Bar exam multiple choice and 71.2% on Codex HumanEval, up from Claude 1.3.

    Press release
  3. Claude 3FlagshipMultimodal

    Three tiers introduced, Haiku, Sonnet and Opus, with vision input and a 200K context window; Anthropic claimed Opus outperformed peers on MMLU, GPQA and GSM8K.

    Press release
  4. Claude 3.5 SonnetFlagship

    First of the 3.5 family, 200K context, claimed to beat Claude 3 Opus at twice the speed and 64% on an agentic coding eval; launched alongside Artifacts.

    Press release
  5. Claude 3.7 SonnetFlagshipReasoning

    Presented as 'the first hybrid reasoning model on the market': near-instant or extended step-by-step thinking with a token budget up to 128K; Claude Code launched in preview.

    Press release
  6. Claude 4FlagshipReasoning

    Claude Opus 4 and Claude Sonnet 4, with extended thinking plus tool use and 72.5% / 72.7% on SWE-bench as reported by Anthropic; Claude Opus 4.1 followed on Aug 5, 2025.

    Press release
  7. Claude Sonnet 4.5Flagship

    Anthropic reported 77.2% on SWE-bench Verified and 61.4% on OSWorld, with 30-hour task focus; the Claude Agent SDK shipped with it and Claude Haiku 4.5 followed Oct 15, 2025.

    Press release
  8. Claude Opus 4.5Flagship

    Opus priced at $5/$25 per million tokens, a new effort parameter, and Anthropic's claim of state-of-the-art SWE-bench Verified with 76% fewer output tokens than Sonnet 4.5.

    Press release
  9. Claude Opus 4.6FlagshipReasoning

    1M-token context (beta), adaptive thinking, four effort levels, context compaction and agent teams; Anthropic claimed the top Terminal-Bench 2.0 and Humanity's Last Exam scores.

    Press release
  10. Claude Opus 4.7FlagshipMultimodal

    New tokenizer, images up to 2,576 px on the long edge, an 'xhigh' effort level and better instruction following; Claude Opus 4.8 followed on May 28, 2026.

    Press release
  11. Claude Fable 5 and Claude Mythos 5FlagshipReasoning

    A new tier above Opus: Fable 5 for general use with safety classifiers, Mythos 5 the same model by invitation for defensive cyber work; 1M context, adaptive thinking always on.

    Press release
  12. Claude Sonnet 5Flagship

    1M context, adaptive thinking on by default, $2/$10 per million tokens; Anthropic said it approaches Opus 4.8 with gains on BrowseComp and OSWorld-Verified over Sonnet 4.6.

    Press release
  13. Claude Opus 5Flagship

    Opus tier for long-running agents; Anthropic claimed state of the art on Frontier-Bench and GDPval-AA coding, with a fast mode at about 2.5x speed for double the price.

    Press release
  14. Claude Fable 5.1 and Claude Mythos 5.1FlagshipReasoning

    Same price as Fable 5 with cache reads at a quarter of the cost; Anthropic reported 55.8% on Terminal-Bench 4.0 and 52.6% on Terminal-Bench-Science, Mythos 5.1 by invitation only.

    Official documentation
Alphabet (Google)

14 releases · 2018 → 2026

BERT, T5, LaMDA and PaLM were research-era lines; Gemini is the flagship since Dec 2023 (1.0, 1.5, 2.0, 2.5, 3, 3.1, then 3.5-3.8 Flash), with Gemma as the open-weights sibling.

201820192020202120222023202420252026 BERT · Nov 2018T5 · Feb 2020LaMDA · May 2021Gemma · Feb 2024Gemma 4 · Apr 2026PaLM · Apr 2022PaLM 2 · May 2023Gemini 1.0 · Dec 2023Gemini 1.5 · Feb 2024Gemini 2.0 · Dec 2024Gemini 2.5 · Mar 2025Gemini 3 · Nov 2025Gemini 3.1 Pro · Feb 2026Gemini 3.5 Flash · May 2026
  1. BERTResearchOpen weights

    Open-sourced bidirectional pre-training for NLP: TensorFlow code plus pre-trained models, with state-of-the-art results claimed on 11 NLP tasks (SQuAD 93.2 F1).

    Press release
  2. T5ResearchOpen weights

    Text-to-Text Transfer Transformer with code, pre-trained models up to 11B parameters and the C4 dataset released; near-human SuperGLUE score claimed for the 11B model.

    Press release
  3. LaMDAResearch

    Language Model for Dialogue Applications, trained on dialogue and unveiled at I/O 2021; presented as research, not released as a product or as weights.

    Press release
  4. PaLMFlagshipResearch

    Pathways Language Model: a 540-billion-parameter dense decoder-only Transformer; state-of-the-art few-shot results claimed, including 58% on GSM8K.

    Press release
  5. PaLM 2Flagship

    Unveiled at I/O 2023 in four sizes (Gecko, Otter, Bison, Unicorn) with improved multilingual, reasoning and coding; said to power Bard and 25+ Google products.

    Press release
  6. Gemini 1.0FlagshipMultimodal

    First Gemini generation, built natively multimodal, in Ultra, Pro and Nano sizes; Ultra claimed as the first model to outperform human experts on MMLU (90.0%).

    Press release
  7. Gemini 1.5FlagshipMultimodal

    Gemini 1.5 Pro, a Mixture-of-Experts model with a 1 million token context window (1 hour of video, 11 hours of audio); 1.5 Flash followed in May 2024.

    Press release
  8. GemmaOpen weightsSmall

    First open-weights line from the Gemini research, released as Gemma 2B and 7B with terms permitting commercial use; Gemma 3 (27B) followed March 12, 2025.

    Press release
  9. Gemini 2.0FlagshipMultimodal

    Gemini 2.0 Flash for the 'agentic era': native image and audio output plus native tool use; said to beat 1.5 Pro on key benchmarks at twice the speed, GA Feb 2025.

    Press release
  10. Gemini 2.5FlagshipReasoning

    Gemini 2.5 Pro Experimental, the first 'thinking' Gemini, with a 1M-token context; claimed top of LMArena and 18.8% on Humanity's Last Exam. Stable 2.5 Pro/Flash June 17, 2025.

    Press release
  11. Gemini 3FlagshipReasoning

    Gemini 3 Pro plus a Deep Think mode; claimed 1501 Elo on LMArena and 37.5% on Humanity's Last Exam, shipped with Google Antigravity. Gemini 3 Flash preview followed Dec 17, 2025.

    Press release
  12. Gemini 3.1 ProFlagshipReasoning

    Preview release presented as a smarter baseline for complex problem-solving; claimed a verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro.

    Press release
  13. Gemma 4Open weightsReasoningMultimodal

    Four sizes (E2B, E4B, 26B MoE, 31B dense) under Apache 2.0, with reasoning, function calling, image and video input, 140+ languages and 128K-256K context.

    Press release
  14. Gemini 3.5 FlashFlagship

    Start of the rapid Flash cadence: gemini-3.5-flash GA, then 3.6 Flash (Jul 21), 3.7 Flash (Aug 13) and 3.8 Flash (Sep 2, 2026) each presented as the top 'workhorse' model.

    Official documentation
Meta Platforms

10 releases · 2023 → 2026

LLaMA/Llama was the open-weights flagship from 2023 through Llama 4 (2025); in 2026 Meta Superintelligence Labs began a closed Muse line (Muse Spark), with Muse Glimmer as its open model.

201820192020202120222023202420252026 LLaMA · Feb 2023Code Llama · Aug 2023Llama 3.2 · Sep 2024Muse Glimmer · Aug 2026Llama 2 · Jul 2023Llama 3 · Apr 2024Llama 3.1 · Jul 2024Llama 4 · Apr 2025Muse Spark · Apr 2026Muse Spark 1.1 · Jul 2026
  1. LLaMAResearchOpen weights

    Foundational model in 7B, 13B, 33B and 65B sizes, released under a noncommercial research license with access granted case by case to researchers.

    Press release
  2. Llama 2FlagshipOpen weights

    Pretrained and fine-tuned chat weights released 'free for research and commercial use', with Microsoft as preferred partner on Azure and Windows.

    Press release
  3. Code LlamaCodeOpen weights

    Code-specialised Llama 2 in 7B, 13B and 34B with Python and Instruct variants under the Llama 2 community license; a 70B size was added January 29, 2024.

    Press release
  4. Llama 3FlagshipOpen weights

    8B and 70B models pretrained on over 15T tokens, claimed the best open models of their class; a 400B+ model was still training at release.

    Press release
  5. Llama 3.1FlagshipOpen weights

    Llama 3.1 405B, presented as the first frontier-level open model, with 128K context, eight languages and a license allowing outputs to train other models.

    Press release
  6. Llama 3.2MultimodalOpen weightsSmall

    First vision Llama models (11B and 90B) plus 1B and 3B text models for edge and mobile devices with 128K context, announced at Connect 2024.

    Press release
  7. Llama 4FlagshipMultimodalOpen weights

    Natively multimodal mixture-of-experts models Scout (16 experts, 10M context claimed) and Maverick (128 experts) downloadable at release; Behemoth previewed, still training.

    Press release
  8. Muse SparkFlagshipReasoningMultimodal

    First of the Muse family from Meta Superintelligence Labs: a natively multimodal reasoning model with a Contemplating mode, claimed 58% on Humanity's Last Exam; closed, at meta.ai.

    Press release
  9. Muse Spark 1.1FlagshipReasoningMultimodal

    1 million token context and gains in tool and computer use, shipped with the Meta Model API public preview; 1.2 (Aug 5, with Muse Code) and 1.3 (Sep 2, 2026) followed.

    Press release
  10. Muse GlimmerOpen weightsSmall

    First open-weights Muse model: a 30-billion-parameter agentic model under Apache 2.0, sized to run on a single consumer GPU or Mac for always-on local agents.

    Press release
xAI

12 releases · 2023 → 2026

Grok for every model, numbered Grok-1 through Grok 4.6, with Fast, Heavy and Think variants; the lab's 2026 posts sign as SpaceXAI.

201820192020202120222023202420252026 Grok-1 (open release) · Mar 2024Grok-1.5V · Apr 2024Grok 4 Fast · Sep 2025Grok-1 · Nov 2023Grok-1.5 · Mar 2024Grok-2 · Aug 2024Grok 3 · Feb 2025Grok 4 · Jul 2025Grok 4.1 · Nov 2025Grok 4.20 · Mar 2026Grok 4.5 · Jul 2026Grok 4.6 · Aug 2026
  1. Grok-1Flagship

    First Grok, with real-time knowledge via X; xAI reported 73% MMLU and 63.2% HumanEval and said it surpassed ChatGPT-3.5 in its compute class, after a 33B Grok-0 prototype.

    Press release
  2. Grok-1 (open release)Open weights

    Base-model weights and architecture of Grok-1 released under Apache 2.0: a 314B-parameter mixture-of-experts with 25% of weights active per token, not fine-tuned for dialogue.

    Press release
  3. Grok-1.5Flagship

    128K-token context, a 16x increase, with xAI reporting 50.6% on MATH, 90% on GSM8K and 74.1% on HumanEval.

    Press release
  4. Grok-1.5VMultimodal

    xAI's 'first-generation multimodal model', reading documents, diagrams, charts, screenshots and photos, introduced with the RealWorldQA benchmark.

    Press release
  5. Grok-2FlagshipMultimodal

    Grok-2 and Grok-2 mini in beta, text and vision, with xAI citing 56.0% GPQA, 76.1% MATH and 88.4% HumanEval; image generation via Black Forest Labs' FLUX.1 on X.

    Press release
  6. Grok 3FlagshipReasoning

    Grok 3 and Grok 3 mini plus a reasoning mode, Grok 3 (Think), and the DeepSearch agent; xAI claimed a 1M-token context and 93.3% on AIME 2025 for Think.

    Press release
  7. Grok 4FlagshipReasoning

    Native tool use and real-time search, 256K API context, with Grok 4 Heavy using parallel test-time compute; xAI claimed 50.7% on Humanity's Last Exam.

    Press release
  8. Grok 4 FastSmallReasoning

    Cost-efficient model uniting reasoning and non-reasoning modes with a 2M-token context; xAI claimed Grok 4-level benchmark results at a 98% lower price.

    Press release
  9. Grok 4.1FlagshipReasoning

    Grok 4.1 Thinking and a non-reasoning mode, with xAI citing #1 and #2 on LMArena, better emotional intelligence and creative writing, and a lower hallucination rate.

    Press release
  10. Grok 4.20Flagship

    Grok 4.20 in reasoning and non-reasoning forms plus Grok 4.20 Multi-agent, listed in the API docs with a 1M-token context.

    Official documentation
  11. Grok 4.5FlagshipCode

    Built for coding, agentic tasks and knowledge work with Cursor; xAI claimed 62.0% on DeepSWE 1.0 at 80 tokens per second, priced $2/$6 per million tokens.

    Press release
  12. Grok 4.6Flagship

    Aimed at long-running agents and interactive and visual work; xAI said it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index; a fast variant costs double.

    Press release
Mistral AI

13 releases · 2023 → 2026

Open-weights Mistral 7B and Mixtral first, then a Large/Medium/Small ladder plus Codestral/Devstral (code), Pixtral (vision) and Magistral (reasoning), unified in Mistral 3, Small 4 and Medium 3.5.

201820192020202120222023202420252026 Mistral 7B · Sep 2023Codestral · May 2024Pixtral 12B · Sep 2024Mistral Small 3 · Jan 2025Magistral · Jun 2025Devstral 2 · Dec 2025Mistral Small 4 · Mar 2026Mixtral 8x7B · Dec 2023Mistral Large · Feb 2024Mistral Large 2 · Jul 2024Mistral Medium 3 · May 2025Mistral 3 · Dec 2025Mistral Medium 3.5 · May 2026
  1. Mistral 7BOpen weightsSmall

    First release: a 7B model under Apache 2.0 using grouped-query and sliding-window attention, claimed to outperform Llama 2 13B on all benchmarks; shipped with 7B Instruct.

    Press release
  2. Mixtral 8x7BFlagshipOpen weights

    Sparse mixture-of-experts (46.7B total, 12.9B active per token) under Apache 2.0 with 32k context; claimed to beat Llama 2 70B and match or beat GPT-3.5 on most benchmarks.

    Press release
  3. Mistral LargeFlagship

    First closed flagship, API-only via la Plateforme and Azure, 32K context, five European languages; claimed second-ranked API model after GPT-4. Launched with Le Chat.

    Press release
  4. CodestralCodeOpen weights

    First code model: 22B, 80+ programming languages, 32k context, fill-in-the-middle; weights under the Mistral AI Non-Production License. Codestral 25.08 update July 30, 2025.

    Press release
  5. Mistral Large 2FlagshipOpen weights

    123B dense model with 128k context, 84.0% MMLU claimed, weights on Hugging Face under the Mistral Research License (commercial license required for self-deployment).

    Press release
  6. Pixtral 12BMultimodalOpen weights

    First natively multimodal model: a 400M vision encoder plus 12B decoder, 128k context, Apache 2.0, 52.5% MMMU claimed; Pixtral Large followed November 18, 2024.

    Press release
  7. Mistral Small 3SmallOpen weights

    Latency-optimised 24B model under Apache 2.0, claimed over 81% MMLU at 150 tokens/s and competitive with Llama 3.3 70B while more than 3x faster.

    Press release
  8. Mistral Medium 3Flagship

    Enterprise mid-tier model claimed at or above 90% of Claude Sonnet 3.7 at $0.4/$2 per M tokens, deployable on four GPUs; closed weights, hybrid and on-prem options.

    Press release
  9. MagistralReasoningOpen weights

    First reasoning model, released as Magistral Small (24B, Apache 2.0) and Magistral Medium (enterprise); Medium claimed 73.6% on AIME 2024, 90% with majority voting.

    Press release
  10. Mistral 3FlagshipOpen weightsMultimodal

    Mistral Large 3 (sparse MoE, 675B total / 41B active) plus Ministral 3 at 14B, 8B and 3B, all Apache 2.0, with image understanding and 40+ languages.

    Press release
  11. Devstral 2CodeOpen weights

    Agentic coding models Devstral 2 (123B, modified MIT) and Devstral Small 2 (24B, Apache 2.0), claimed 72.2% on SWE-bench Verified, shipped with the Mistral Vibe CLI.

    Press release
  12. Mistral Small 4SmallReasoningOpen weights

    119B-total / 6B-active MoE under Apache 2.0 that merges Magistral reasoning, Pixtral vision and Devstral coding into one model with configurable reasoning effort.

    Press release
  13. Mistral Medium 3.5FlagshipOpen weights

    Called the first flagship merged model: dense 128B, 256k context, instruction, reasoning and coding in one set of weights, modified MIT license; 77.6% SWE-Bench Verified claimed.

    Press release
DeepSeek

11 releases · 2023 → 2026

DeepSeek-numbered flagship line (LLM 67B, V2, V3, V4) with R1 as the reasoning branch folded into the V3.x hybrid-thinking models; DeepSeek Coder was a parallel code line.

201820192020202120222023202420252026 DeepSeek Coder · Nov 2023DeepSeek-Coder-V2 · Jun 2024DeepSeek-R1-Lite-Preview · Nov 2024DeepSeek-R1 · Jan 2025DeepSeek-V4-Flash-Vision-Exp · Aug 2026DeepSeek LLM · Nov 2023DeepSeek-V2 · May 2024DeepSeek-V3 · Dec 2024DeepSeek-V3.1 · Aug 2025DeepSeek-V3.2 · Dec 2025DeepSeek-V4 · Apr 2026
  1. DeepSeek CoderCodeOpen weights

    Code models from 1B to 33B trained from scratch on 2T tokens (87% code), 16K window for project-level completion and infilling; commercial use permitted under the Model License.

    Official documentation
  2. DeepSeek LLMFlagshipOpen weights

    First general model: 7B and 67B base and chat trained on 2T tokens; the lab reports 67B Chat at 73.78 HumanEval and 84.1 GSM8K, commercial use permitted.

    Official documentation
  3. DeepSeek-V2FlagshipOpen weights

    236B MoE with 21B active, 128K context, MLA attention and DeepSeekMoE; the lab claims 42.5% lower training cost and 93.3% smaller KV cache than DeepSeek 67B; V2-Lite on 2024-05-16.

    Official documentation
  4. DeepSeek-Coder-V2CodeOpen weights

    Code MoE continued from V2 on 6T more tokens, 16B-Lite and 236B variants, 338 languages, 128K context; the paper claims performance comparable to GPT4-Turbo on code tasks.

    Research report
  5. DeepSeek-R1-Lite-PreviewReasoning

    First reasoning model, live on chat.deepseek.com with a visible thought process; the lab claims o1-preview-level AIME and MATH results, with open weights and API 'coming soon'.

    Press release
  6. DeepSeek-V3FlagshipOpen weights

    671B MoE with 37B active pre-trained on 14.8T tokens, 60 tokens/s (3x V2), API-compatible with V2.5; fully open weights and paper on GitHub.

    Press release
  7. DeepSeek-R1ReasoningOpen weights

    Reasoning model with MIT-licensed weights; the lab claims performance on par with OpenAI-o1 in math, code and reasoning, plus six distilled models, 32B and 70B on par with o1-mini.

    Press release
  8. DeepSeek-V3.1FlagshipReasoningOpen weights

    Hybrid inference: one model with Think and Non-Think modes, 128K context on both API endpoints, post-training for tool use and agents; Base and instruct weights on Hugging Face.

    Press release
  9. DeepSeek-V3.2FlagshipReasoningOpen weights

    Successor to V3.2-Exp (2025-09-29, DeepSeek Sparse Attention debut); first DeepSeek model with thinking inside tool use; API-only V3.2-Speciale claims IMO and IOI 2025 gold level.

    Press release
  10. DeepSeek-V4FlagshipReasoningOpen weights

    Preview of V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B), 1M context now default across DeepSeek services, open weights; claims open-source SOTA in agentic coding.

    Press release
  11. DeepSeek-V4-Flash-Vision-ExpMultimodal

    Experimental multimodal model keeping V4-Flash text ability, image input via base64, URL or the new Files API at up to 384 tokens per image; claims a big leap on multimodal agents.

    Press release
Alibaba

13 releases · 2023 → 2026

Qwen (Tongyi Qianwen) is the whole line, numbered 1 to 1.5, 2, 2.5, 3, 3.5, 3.8 and mostly open-weights, with closed -Max/-Plus API tiers and QwQ as the first reasoning branch.

201820192020202120222023202420252026 Qwen-VL · Aug 2023QwQ-32B-Preview · Nov 2024Qwen3-Coder · Jul 2025Qwen3-Next · Sep 2025Qwen3-VL · Sep 2025Qwen-7B · Aug 2023Qwen1.5 · Feb 2024Qwen2 · Jun 2024Qwen2.5 · Sep 2024Qwen2.5-Max · Jan 2025Qwen3 · Apr 2025Qwen3.5 · Feb 2026Qwen3.8 · Aug 2026
  1. Qwen-7BFlagshipOpen weights

    First open Qwen release, 7B base and chat on ModelScope and Hugging Face; Qwen-14B followed 2023-09-25, and Qwen-72B plus Qwen-1.8B (3T tokens, 32K context) on 2023-11-30.

    Official documentation
  2. Qwen-VLMultimodalOpen weights

    Vision-language model on Qwen-7B taking image, text and bounding boxes as input; base and chat weights released with a paper, commercial use allowed; Qwen-VL-Max API 2024-01-18.

    Official documentation
  3. Qwen1.5FlagshipOpen weights

    Base and chat models from 0.5B to 72B (32B, 110B and an MoE added later), all with 32K context; code merged into Hugging Face transformers 4.37 so no custom code is needed.

    Press release
  4. Qwen2FlagshipOpen weights

    Five sizes 0.5B to 72B including a 57B-A14B MoE, 128K context on instruct models, 27 more languages; all but 72B under Apache 2.0, and the lab claims 72B beats Qwen1.5-110B.

    Press release
  5. Qwen2.5FlagshipOpen weights

    Seven sizes 0.5B to 72B trained on up to 18T tokens, 128K context, shipped with Qwen2.5-Coder and Qwen2.5-Math; Apache 2.0 except 3B and 72B; API-only Qwen-Plus and Qwen-Turbo.

    Press release
  6. QwQ-32B-PreviewReasoningOpen weights

    Experimental 32B reasoning model that reflects and self-questions; the lab reports 65.2% GPQA, 50.0% AIME and 90.6% MATH-500, weights on Hugging Face. QwQ-32B followed 2025-03-06.

    Press release
  7. Qwen2.5-MaxFlagship

    Large-scale MoE pretrained on over 20T tokens with SFT and RLHF, API-only via Alibaba Cloud (qwen-max-2025-01-25) and Qwen Chat; the lab claims it beats DeepSeek V3 on Arena-Hard.

    Press release
  8. Qwen3FlagshipReasoningOpen weights

    Two MoE (235B-A22B, 30B-A3B) and six dense models 0.6B to 32B with switchable thinking and non-thinking modes, 119 languages, up to 128K context, Apache 2.0; claims R1/o1 parity.

    Press release
  9. Qwen3-CoderCodeOpen weights

    480B MoE with 35B active for agentic coding, 256K context natively and 1M with extrapolation, released with the open-source Qwen Code CLI; claims open-model SOTA on agentic coding.

    Press release
  10. Qwen3-NextOpen weights

    Qwen3-Next-80B-A3B, an ultra-sparse MoE with hybrid Gated DeltaNet and gated attention, 262K native context extensible to 1M; said to match Qwen3-235B-A22B-Instruct-2507.

    Official documentation
  11. Qwen3-VLMultimodalOpen weights

    Qwen3-VL-235B-A22B Instruct and Thinking launched the series; 30B-A3B, 4B, 8B, 2B and 32B variants followed through 2025-10-21, with a technical paper on 2025-11-27.

    Official documentation
  12. Qwen3.5FlagshipMultimodalOpen weights

    Natively multimodal 397B-A17B MoE with early-fusion vision-language training, 262K context, thinking on by default, 201 languages, Apache 2.0; Qwen3.6 sizes followed in April 2026.

    Official documentation
  13. Qwen3.8FlagshipReasoningOpen weights

    Qwen3.8-2.4T-A95B (2.4T total / 95B active, 262K context, text-only, thinking required, reasoning_effort levels), called Qwen3.8-Max in the blog; 27B multimodal on 2026-08-14.

    Official documentation
Moonshot AI

8 releases · 2025 → 2026

Kimi names both the product and the models: Kimi Chat, then k1.5 for reasoning, the open K2 line (K2, K2 Thinking, K2.5, K2.6, K2.7-Code) and K3 in 2026.

201820192020202120222023202420252026 Kimi k1.5 · Jan 2025Kimi-VL · Apr 2025Kimi-Dev · Jun 2025Kimi K2 Thinking · Nov 2025Kimi K2 · Jul 2025Kimi K2.5 · Jan 2026Kimi K2.6 · Apr 2026Kimi K3 · Jul 2026
  1. Kimi k1.5ReasoningResearchMultimodal

    Multimodal reasoning model trained with long-context RL; the paper claims 77.5 on AIME and 96.2 on MATH-500, and 'long2short' methods beating GPT-4o and Claude 3.5 Sonnet.

    Press release
  2. Kimi-VLMultimodalOpen weightsSmall

    16B MoE vision-language model with 2.8B active, MoonViT native-resolution encoder, 128K context and a Thinking variant under MIT; the lab reports 61.7 MMMU for the thinking model.

    Press release
  3. Kimi-DevCodeOpen weights

    Kimi-Dev-72B, an open coding model built on Qwen2.5-72B and optimised with large-scale RL to patch repositories; the lab claims 60.4% on SWE-bench Verified, MIT licence.

    Press release
  4. Kimi K2FlagshipOpen weights

    1T-parameter MoE with 32B active, optimised for agentic tool use as a non-thinking model, released as K2-Base and K2-Instruct; the 0905 update (2025-09-05) brought 256K context.

    Press release
  5. Kimi K2 ThinkingReasoningOpen weights

    Thinking agent on K2 that reasons step by step across 200-300 sequential tool calls, 256K context, native INT4 QAT; claims 44.9% HLE with tools and 71.3% SWE-bench Verified.

    Press release
  6. Kimi K2.5FlagshipMultimodalOpen weights

    Native multimodal agentic model (1T total / 32B active, MoonViT encoder, 256K context) with thinking and instant modes and an agent swarm of up to 100 sub-agents; Modified MIT.

    Press release
  7. Kimi K2.6FlagshipOpen weights

    Open-sourced update focused on long-horizon coding, agent swarm scaled to 300 sub-agents and better tool calling; the lab reports 58.6% SWE-Bench Pro and 66.7% Terminal-Bench 2.0.

    Press release
  8. Kimi K3FlagshipMultimodalOpen weights

    2.8T-parameter model (104B active) with native vision and 1M-token context, max thinking effort by default; announced 2026-07-16, weights promised by 2026-07-27 (Kimi K3 License).

    Press release
Z.ai (Zhipu AI)

13 releases · 2022 → 2026

GLM for every generation: GLM-130B, then the ChatGLM-6B open chat models, then GLM-4 through GLM-4.7 and GLM-5.x; GLM-Z1 was the first reasoning line, GLM-4.5V/4.6V the vision variants.

201820192020202120222023202420252026 GLM-130B · Aug 2022ChatGLM-6B · Mar 2023ChatGLM2-6B · Jun 2023ChatGLM3-6B · Oct 2023GLM-Z1-32B-0414 · Apr 2025GLM-4 · Jan 2024GLM-4.5 · Jul 2025GLM-4.6 · Sep 2025GLM-4.7 · Dec 2025GLM-5 · Feb 2026GLM-5.1 · Apr 2026GLM-5.2 · Jun 2026GLM-5.3 · Aug 2026
  1. GLM-130BResearchOpen weights

    130B-parameter bilingual (English/Chinese) pre-trained model released with weights on application, presented as runnable on a single A100 server; paper followed 2022-10-05.

    Press release
  2. ChatGLM-6BOpen weightsSmall

    First ChatGLM: the aligned ChatGLM-130B went live at chatglm.cn and its 6.2B open-weights version ChatGLM-6B, runnable in 6GB GPU memory, shipped the same day.

    Official documentation
  3. ChatGLM2-6BOpen weightsSmall

    Second-generation 6B chat model: context 2K to 32K, MMLU +23% and GSM8K +571% over ChatGLM-6B per the lab, 42% faster inference; a 32K variant followed 2023-07-31.

    Official documentation
  4. ChatGLM3-6BOpen weightsSmall

    Third-generation 6B open model adding native function call, code interpreter and agent prompting, shipped with Base, 32K and 128K variants.

    Research report
  5. GLM-4Flagship

    First closed GLM-4 (0116) via API, claimed to match GPT-4 Turbo (128K) on long context; GLM-4 (0520), GLM-4-Air and the open GLM-4-9B (128K/1M) followed by 2024-06-05.

    Research report
  6. GLM-Z1-32B-0414ReasoningOpen weights

    First GLM reasoning model, built on GLM-4-32B-0414 with cold start and extended RL on math, code and logic; shipped MIT-licensed with the GLM-4-32B-0414 series and a 9B variant.

    Official documentation
  7. GLM-4.5FlagshipOpen weights

    355B/32B-active MoE with thinking and non-thinking modes, 128K context, MIT licence, ranked 3rd across 12 benchmarks per the lab; shipped with GLM-4.5-Air, GLM-4.5V followed 2025-08-11.

    Official documentation
  8. GLM-4.6FlagshipOpen weights

    Context expanded from 128K to 200K, higher code-benchmark scores and tool use during inference, claimed gains over GLM-4.5 on eight benchmarks; GLM-4.6V vision followed 2025-12-08.

    Official documentation
  9. GLM-4.7FlagshipOpen weights

    355B/32B-active MIT model with interleaved and preserved thinking, claiming 73.8% SWE-bench Verified; the 30B/3B-active GLM-4.7-Flash followed 2026-01-19.

    Official documentation
  10. GLM-5FlagshipOpen weights

    Scaled to 744B/40B-active on 28.5T tokens with DeepSeek Sparse Attention, MIT licence, claimed best among open models on reasoning, coding and agentic tasks.

    Official documentation
  11. GLM-5.1FlagshipOpen weights

    754B MIT model for agentic engineering, 200K context, presented as able to work autonomously on one task for up to 8 hours across thousands of tool calls.

    Official documentation
  12. GLM-5.2FlagshipOpen weights

    753B MIT model introducing a 1M-token context for long-horizon tasks, with IndexShare attention claimed to cut per-token FLOPs 2.9x at 1M context.

    Official documentation
  13. GLM-5.3FlagshipOpen weights

    753B model with 1M context, claimed 50% better coding than GLM-5.2 on Z.ai Code Bench and emergent cybersecurity capability; GLM-5.3-Flash with native vision followed 2026-08-26.

    Official documentation
Baidu

9 releases · 2019 → 2026

ERNIE for everything: research models ERNIE 1.0-3.0, the ERNIE Bot product on ERNIE 3.5/4.0, then ERNIE 4.5 (open-sourced June 2025), X1 for reasoning, and ERNIE 5.0/5.1.

201820192020202120222023202420252026 ERNIE · Apr 2019ERNIE 3.0 · Jul 2021ERNIE X1 · Mar 2025ERNIE 4.5 (open-source family) · Jun 2025ERNIE Bot · Mar 2023ERNIE 4.0 · Oct 2023ERNIE 4.5 · Mar 2025ERNIE 5.0 · Nov 2025ERNIE 5.1 · May 2026
  1. ERNIEResearch

    First ERNIE (Enhanced Representation through Knowledge Integration): a BERT-style Chinese model with entity- and phrase-level masking, claimed to beat BERT on Chinese NLP benchmarks.

    Research report
  2. ERNIE 3.0Research

    10B-parameter model fusing auto-regressive and auto-encoding networks trained on 4TB of text plus a knowledge graph, claimed to surpass human performance on SuperGLUE.

    Research report
  3. ERNIE BotFlagshipMultimodal

    Baidu's first chatbot product, built on ERNIE and PLATO, presented with literary creation, business writing, maths, Chinese understanding and multi-modal generation.

    Press release
  4. ERNIE 4.0Flagship

    Presented at Baidu World as the most advanced ERNIE foundation model, with upgraded understanding, generation, reasoning and memory to power AI-native applications.

    Press release
  5. ERNIE 4.5FlagshipMultimodal

    Native multimodal foundation model, free to individuals in ERNIE Bot and priced from RMB 0.004 per 1K input tokens via API; ERNIE 4.5 Turbo followed 2025-04-25.

    Press release
  6. ERNIE X1Reasoning

    Baidu's first deep-thinking reasoning model with tool use, launched alongside ERNIE 4.5 at half its price; X1 Turbo followed 2025-04-25 and X1.1 on 2025-09-09.

    Press release
  7. ERNIE 4.5 (open-source family)Open weightsMultimodal

    Ten open-weight ERNIE 4.5 models from 0.3B to 424B under Apache 2.0, MoE text and vision-language variants, 131K context on the 300B-A47B text model.

    Press release
  8. ERNIE 5.0FlagshipMultimodal

    Natively omni-modal 2.4T-parameter model jointly handling text, image, audio and video, in public preview via ERNIE Bot and Qianfan; the full-release blog followed 2026-02-06.

    Press release
  9. ERNIE 5.1FlagshipMultimodal

    Compresses total parameters to about a third of ERNIE 5.0 while claiming leading performance at lower cost; preceded by ERNIE-5.1-Preview on LMArena 2026-04-30.

    Press release
Cohere

7 releases · 2024 → 2026

Command for the enterprise line: Command, then Command R and R+ (2024), Command A (2025) with Vision, Reasoning and Translate variants, and Command A+ (2026); Aya is the separate multilingual research line.

201820192020202120222023202420252026 Command R7B · Dec 2024Command A Vision · Jul 2025Command A Reasoning · Aug 2025Command R · Mar 2024Command R+ · Apr 2024Command A · Mar 2025Command A+ · May 2026
  1. Command RFlagshipOpen weights

    35B model built for production-scale retrieval-augmented generation with 128K context, released with CC-BY-NC research weights on Hugging Face; refreshed as command-r-08-2024.

    Press release
  2. Command R+FlagshipOpen weights

    104B RAG-optimised model for enterprise workloads with 128K context and multi-step tool use, first available on Microsoft Azure, CC-BY-NC weights on Hugging Face; refreshed 08-2024.

    Press release
  3. Command R7BSmallOpen weights

    Smallest model in the R series, 7B with 128K context, aimed at commodity GPUs and edge devices, released as CC-BY-NC open weights.

    Press release
  4. Command AFlagshipOpen weights

    111B model with 256K context claimed to match or beat GPT-4o and DeepSeek-V3 on enterprise agentic tasks at far lower compute; CC-BY-NC weights on Hugging Face.

    Press release
  5. Command A VisionMultimodalOpen weights

    Cohere's first image-capable model, 112B, for enterprise chart, graph and diagram analysis at a low compute footprint; CC-BY-NC weights on Hugging Face.

    Press release
  6. Command A ReasoningReasoningOpen weights

    Cohere's first reasoning model, 111B with 256K context and a controllable thinking budget for agents; open-weights research release under CC-BY-NC; Command A Translate shipped the same month.

    Press release
  7. Command A+FlagshipOpen weightsMultimodal

    Cohere's first Mixture-of-Experts model, 218B total/25B active, 128K context, vision input, reasoning and 48 languages in one model, released under Apache 2.0.

    Press release
MiniMax

6 releases · 2025 → 2026

MiniMax-Text-01 (2025) began the open line, then M1 for reasoning (2025), the agent-focused M2 line (M2, M2.1, M2.5, M2.7; 2025-26) and the multimodal M3 in 2026.

201820192020202120222023202420252026 MiniMax-Text-01 · Jan 2025MiniMax-M1 · Jun 2025MiniMax-M2 · Oct 2025MiniMax-M2.5 · Feb 2026MiniMax-M2.7 · Mar 2026MiniMax-M3 · Jun 2026
  1. MiniMax-Text-01FlagshipOpen weights

    456B-parameter mixture-of-experts language model (45.9B active) using Lightning Attention with a 4 million token context, released with MiniMax-VL-01.

    Press release
  2. MiniMax-M1FlagshipReasoningOpen weights

    Open-weight hybrid-attention reasoning model, 456B parameters (45.9B active), 1 million token context and up to 80,000 tokens of reasoning output; 40k and 80k variants.

    Press release
  3. MiniMax-M2FlagshipReasoningCodeOpen weights

    Model built for agentic and coding work, with open weights on MiniMax's Hugging Face organisation; M2.1 followed on 22 Dec 2025.

    Official documentation
  4. MiniMax-M2.5FlagshipReasoningCodeOpen weights

    M2 line update (with M2.5-highspeed) aimed at programming, tool calling, search and office work; only the month was stated. M2.1 (Dec. 22, 2025) preceded it.

    Official documentation
  5. MiniMax-M2.7FlagshipReasoningCodeOpen weights

    M2.7 and a M2.7-highspeed variant, described as emphasizing recursive self-improvement.

    Official documentation
  6. MiniMax-M3FlagshipReasoningMultimodalOpen weights

    Natively multimodal model, about 428B parameters (23B active) with MiniMax Sparse Attention and up to 1M token context; API on release, weights on Hugging Face later.

    Official documentation
StepFun

4 releases · 2025 → 2026

Step for every generation: Step-1 and Step-2 first, then the open-weights Step 3 (2025), Step 3.5 Flash and 3.7 Flash (2026), and Step 5 Preview (2026), API-first with weights announced for October.

201820192020202120222023202420252026 Step 3 · Jul 2025Step 3.5 Flash · Feb 2026Step 3.7 Flash · May 2026Step 5 Preview · Sep 2026
  1. Step 3FlagshipReasoningMultimodalOpen weights

    321B-parameter (38B active) mixture-of-experts vision-language model with 65,536-token context; Apache-2.0 weights; only the month is stated here (paper submitted 2025-07-25).

    Research report
  2. Step 3.5 FlashFlagshipReasoningOpen weights

    196B-parameter, 11B-active sparse mixture-of-experts model with 256K context under Apache-2.0.

    Research report
  3. Step 3.7 FlashFlagshipReasoningMultimodalOpen weights

    About 198B-parameter sparse mixture-of-experts vision-language model with 256K context, Apache-2.0 weights and three selectable reasoning levels; only the month is confirmed.

    Official documentation
  4. Step 5 PreviewFlagshipReasoningMultimodal

    600B-parameter, 27B-active sparse mixture-of-experts model with 1M-token context and text, image and video input, API-only at launch with open weights announced for October 15, 2026.

    News reporting
LG Corp

6 releases · 2024 → 2026

EXAONE 3.0 (2024, first open release), 3.5 (2024), Deep (2025, reasoning), 4.0 (2025, hybrid reasoning), 4.5 (2026, vision-language) and the sovereign-project K-EXAONE line, K-EXAONE 2.0 in 2026.

201820192020202120222023202420252026 EXAONE Deep · Mar 2025EXAONE 3.0 · Aug 2024EXAONE 3.5 · Dec 2024EXAONE 4.0 · Jul 2025EXAONE 4.5 · Apr 2026K-EXAONE 2.0 · Jul 2026
  1. EXAONE 3.0FlagshipOpen weights

    First open EXAONE: a 7.8B instruction-tuned bilingual (English/Korean) model released on Hugging Face for research use.

    Official documentation
  2. EXAONE 3.5FlagshipOpen weights

    Open-sourced in 2.4B on-device, 7.8B general and 32B sizes, announced with the ChatEXAONE enterprise agent.

    Press release
  3. EXAONE DeepReasoningOpen weights

    Reasoning models in 2.4B, 7.8B and 32B sizes for math, science and coding.

    Official documentation
  4. EXAONE 4.0FlagshipReasoningOpen weights

    Hybrid model with non-reasoning and reasoning modes in 32B and 1.2B sizes, 131,072-token context, tool use, under a non-commercial licence.

    Official documentation
  5. EXAONE 4.5FlagshipMultimodalOpen weights

    First open-weight EXAONE vision-language model, 33B parameters (32B language model plus 1.2B vision encoder), 256K context, non-commercial licence.

    Official documentation
  6. K-EXAONE 2.0FlagshipOpen weights

    750B-parameter mixture-of-experts model (A37B active) from Korea's sovereign foundation model project, moved to Apache-2.0, supporting 10 languages.

    Press release