GLM
Z.ai's family, overwhelmingly open-weight.
At a glance
Current release
From the release row GLM-5.3.
Family record: Official documentation · 5 Sep 2026
Built on it
8 products in this catalogue name this family on their own pages.
-
Baseten Model APIs
Baseten
Official documentation · 5 Sep 2026 -
Fireworks AI
Fireworks AI
Official documentation · 5 Sep 2026 -
MaaS
Huawei Cloud
Official documentation · 5 Sep 2026 -
Nebius Token Factory
Nebius Group
Official documentation · 5 Sep 2026 -
Together Inference
Together AI
Official documentation · 5 Sep 2026 -
Vanchin
Kuaishou
Official documentation · 5 Sep 2026 -
Z.ai
Z.ai (Zhipu AI)
Official documentation · 5 Sep 2026 -
Z.ai Model API
Z.ai (Zhipu AI)
Official documentation · 5 Sep 2026
Releases
-
GLM-130B
130B-parameter bilingual (English/Chinese) pre-trained model released with weights on application, presented as runnable on a single A100 server; paper followed 2022-10-05.
-
ChatGLM-6B
First ChatGLM: the aligned ChatGLM-130B went live at chatglm.cn and its 6.2B open-weights version ChatGLM-6B, runnable in 6GB GPU memory, shipped the same day.
-
ChatGLM2-6B
Second-generation 6B chat model: context 2K to 32K, MMLU +23% and GSM8K +571% over ChatGLM-6B per the lab, 42% faster inference; a 32K variant followed 2023-07-31.
-
ChatGLM3-6B
Third-generation 6B open model adding native function call, code interpreter and agent prompting, shipped with Base, 32K and 128K variants.
-
GLM-4
First closed GLM-4 (0116) via API, claimed to match GPT-4 Turbo (128K) on long context; GLM-4 (0520), GLM-4-Air and the open GLM-4-9B (128K/1M) followed by 2024-06-05.
-
GLM-Z1-32B-0414
First GLM reasoning model, built on GLM-4-32B-0414 with cold start and extended RL on math, code and logic; shipped MIT-licensed with the GLM-4-32B-0414 series and a 9B variant.
-
GLM-4.5
355B/32B-active MoE with thinking and non-thinking modes, 128K context, MIT licence, ranked 3rd across 12 benchmarks per the lab; shipped with GLM-4.5-Air, GLM-4.5V followed 2025-08-11.
-
GLM-4.6
Context expanded from 128K to 200K, higher code-benchmark scores and tool use during inference, claimed gains over GLM-4.5 on eight benchmarks; GLM-4.6V vision followed 2025-12-08.
-
GLM-4.7
355B/32B-active MIT model with interleaved and preserved thinking, claiming 73.8% SWE-bench Verified; the 30B/3B-active GLM-4.7-Flash followed 2026-01-19.
-
GLM-5
Scaled to 744B/40B-active on 28.5T tokens with DeepSeek Sparse Attention, MIT licence, claimed best among open models on reasoning, coding and agentic tasks.
-
GLM-5.1
754B MIT model for agentic engineering, 200K context, presented as able to work autonomously on one task for up to 8 hours across thousands of tool calls.
-
GLM-5.2
753B MIT model introducing a 1M-token context for long-horizon tasks, with IndexShare attention claimed to cut per-token FLOPs 2.9x at 1M context.
-
GLM-5.3
753B model with 1M context, claimed 50% better coding than GLM-5.2 on Z.ai Code Bench and emergent cybersecurity capability; GLM-5.3-Flash with native vision followed 2026-08-26.