Qwen
Alibaba's family, overwhelmingly open-weight, served through Model Studio.
At a glance
Current release
From the release row Qwen3.8. Price of qwen3.8-max, the API model Alibaba calls 'the official version based on' these weights (1M context, vision input). International deployment; Global and Beijing are $1.65 / $4.951.
Family record: Official documentation · 5 Sep 2026
Built on it
6 products in this catalogue name this family on their own pages.
-
Alibaba Cloud Model Studio
Alibaba
Official documentation · 5 Sep 2026 -
Nebius Token Factory
Nebius Group
Official documentation · 5 Sep 2026 -
Qwen
Alibaba
Official documentation · 5 Sep 2026 -
Qwen Code
Alibaba
Official documentation · 5 Sep 2026 -
SambaCloud
SambaNova
Official documentation · 5 Sep 2026 -
Together Inference
Together AI
Official documentation · 5 Sep 2026
Releases
-
Qwen-7B
First open Qwen release, 7B base and chat on ModelScope and Hugging Face; Qwen-14B followed 2023-09-25, and Qwen-72B plus Qwen-1.8B (3T tokens, 32K context) on 2023-11-30.
-
Qwen-VL
Vision-language model on Qwen-7B taking image, text and bounding boxes as input; base and chat weights released with a paper, commercial use allowed; Qwen-VL-Max API 2024-01-18.
-
Qwen1.5
Base and chat models from 0.5B to 72B (32B, 110B and an MoE added later), all with 32K context; code merged into Hugging Face transformers 4.37 so no custom code is needed.
-
Qwen2
Five sizes 0.5B to 72B including a 57B-A14B MoE, 128K context on instruct models, 27 more languages; all but 72B under Apache 2.0, and the lab claims 72B beats Qwen1.5-110B.
-
Qwen2.5
Seven sizes 0.5B to 72B trained on up to 18T tokens, 128K context, shipped with Qwen2.5-Coder and Qwen2.5-Math; Apache 2.0 except 3B and 72B; API-only Qwen-Plus and Qwen-Turbo.
-
QwQ-32B-Preview
Experimental 32B reasoning model that reflects and self-questions; the lab reports 65.2% GPQA, 50.0% AIME and 90.6% MATH-500, weights on Hugging Face. QwQ-32B followed 2025-03-06.
-
Qwen2.5-Max
Large-scale MoE pretrained on over 20T tokens with SFT and RLHF, API-only via Alibaba Cloud (qwen-max-2025-01-25) and Qwen Chat; the lab claims it beats DeepSeek V3 on Arena-Hard.
-
Qwen3
Two MoE (235B-A22B, 30B-A3B) and six dense models 0.6B to 32B with switchable thinking and non-thinking modes, 119 languages, up to 128K context, Apache 2.0; claims R1/o1 parity.
-
Qwen3-Coder
480B MoE with 35B active for agentic coding, 256K context natively and 1M with extrapolation, released with the open-source Qwen Code CLI; claims open-model SOTA on agentic coding.
-
Qwen3-Next
Qwen3-Next-80B-A3B, an ultra-sparse MoE with hybrid Gated DeltaNet and gated attention, 262K native context extensible to 1M; said to match Qwen3-235B-A22B-Instruct-2507.
-
Qwen3-VL
Qwen3-VL-235B-A22B Instruct and Thinking launched the series; 30B-A3B, 4B, 8B, 2B and 32B variants followed through 2025-10-21, with a technical paper on 2025-11-27.
-
Qwen3.5
Natively multimodal 397B-A17B MoE with early-fusion vision-language training, 262K context, thinking on by default, 201 languages, Apache 2.0; Qwen3.6 sizes followed in April 2026.
-
Qwen3.8
Qwen3.8-2.4T-A95B (2.4T total / 95B active, 262K context, text-only, thinking required, reasoning_effort levels), called Qwen3.8-Max in the blog; 27B multimodal on 2026-08-14.