Mistral
Mistral AI's family, most of it open-weight.
At a glance
Current release
From the release row Mistral Medium 3.5.
Family record: Press release · 5 Sep 2026
Built on it
4 products in this catalogue name this family on their own pages.
-
Fireworks AI
Fireworks AI
Official documentation · 5 Sep 2026 -
Mistral AI Studio
Mistral AI
Official documentation · 5 Sep 2026 -
Nebius Token Factory
Nebius Group
Official documentation · 5 Sep 2026 -
Together Inference
Together AI
Official documentation · 5 Sep 2026
Releases
-
Mistral 7B
First release: a 7B model under Apache 2.0 using grouped-query and sliding-window attention, claimed to outperform Llama 2 13B on all benchmarks; shipped with 7B Instruct.
-
Mixtral 8x7B
Sparse mixture-of-experts (46.7B total, 12.9B active per token) under Apache 2.0 with 32k context; claimed to beat Llama 2 70B and match or beat GPT-3.5 on most benchmarks.
-
Mistral Large
First closed flagship, API-only via la Plateforme and Azure, 32K context, five European languages; claimed second-ranked API model after GPT-4. Launched with Le Chat.
-
Codestral
First code model: 22B, 80+ programming languages, 32k context, fill-in-the-middle; weights under the Mistral AI Non-Production License. Codestral 25.08 update July 30, 2025.
-
Mistral Large 2
123B dense model with 128k context, 84.0% MMLU claimed, weights on Hugging Face under the Mistral Research License (commercial license required for self-deployment).
-
Pixtral 12B
First natively multimodal model: a 400M vision encoder plus 12B decoder, 128k context, Apache 2.0, 52.5% MMMU claimed; Pixtral Large followed November 18, 2024.
-
Mistral Small 3
Latency-optimised 24B model under Apache 2.0, claimed over 81% MMLU at 150 tokens/s and competitive with Llama 3.3 70B while more than 3x faster.
-
Mistral Medium 3
Enterprise mid-tier model claimed at or above 90% of Claude Sonnet 3.7 at $0.4/$2 per M tokens, deployable on four GPUs; closed weights, hybrid and on-prem options.
-
Magistral
First reasoning model, released as Magistral Small (24B, Apache 2.0) and Magistral Medium (enterprise); Medium claimed 73.6% on AIME 2024, 90% with majority voting.
-
Mistral 3
Mistral Large 3 (sparse MoE, 675B total / 41B active) plus Ministral 3 at 14B, 8B and 3B, all Apache 2.0, with image understanding and 40+ languages.
-
Devstral 2
Agentic coding models Devstral 2 (123B, modified MIT) and Devstral Small 2 (24B, Apache 2.0), claimed 72.2% on SWE-bench Verified, shipped with the Mistral Vibe CLI.
-
Mistral Small 4
119B-total / 6B-active MoE under Apache 2.0 that merges Magistral reasoning, Pixtral vision and Devstral coding into one model with configurable reasoning effort.
-
Mistral Medium 3.5
Called the first flagship merged model: dense 128B, 256k context, instruction, reasoning and coding in one set of weights, modified MIT license; 77.6% SWE-Bench Verified claimed.