What is GLM-5.2?
GLM-5.2 is the flagship model from Z.ai (zai-org on Hugging Face), built for long-horizon tasks: the long agent runs where a model has to stay coherent across hours of work. It is a 753.3B-parameter model, and for the first time Z.ai delivers that long-horizon capability on a solid 1M-token context, a window it says stably sustains long-horizon work instead of degrading as it fills. The weights landed on June 16, 2026 under the MIT license, and the repo has passed 2.5 million downloads.
| Specification | GLM-5.2 |
|---|---|
| Developer | Z.ai (zai-org on Hugging Face) |
| Total parameters | 753.3B |
| Context window | 1M tokens |
| Attention | Sparse attention with IndexShare, one indexer shared across every four layers |
| Per-token FLOPs at 1M context | Cut by 2.9x by IndexShare |
| Multi-Token Prediction | Improved MTP layer, acceptance length up by as much as 20% |
| Coding | Multiple thinking effort levels to balance performance and latency |
| Task | Text generation |
| Modalities | Text input and output |
| Languages | English and Chinese |
| Serving frameworks | SGLang, vLLM, Transformers, KTransformers, Unsloth |
| Technical report | GLM-5, arXiv 2602.15763 |
| Release date | June 16, 2026 |
| License | MIT |
Most of the architecture work targets that 1M window. IndexShare reuses the same indexer across every four sparse attention layers, which cuts per-token FLOPs by 2.9x at a 1M context length, and Z.ai published it as its own paper (arXiv 2603.12201). Z.ai also reworked the MTP layer that drafts tokens for speculative decoding, raising acceptance length by up to 20%. Coding runs with multiple thinking effort levels to balance performance and latency.
GLM-5.2 benchmarks
Z.ai's launch numbers, from the model card, compare GLM-5.2 with its predecessor GLM-5.1, DeepSeek-V4-Pro, Claude Opus 4.8 and GPT-5.5 (the vendor table also includes Qwen3.7-Max, MiniMax M3 and Gemini 3.1 Pro):
| Benchmark | GLM-5.2 | GLM-5.1 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 |
|---|---|---|---|---|---|
AIME 2026 Competition math | 99.2 | 95.3 | 94.6 | 95.7 | 98.3 |
IMOAnswerBench Olympiad math | 91.0 | 83.8 | 89.8 | 83.5 | - |
SWE-bench Pro Harder engineering | 62.1 | 58.4 | 55.4 | 69.2 | 58.6 |
Terminal Bench 2.1 Terminal agents | 81.0 | 63.5 | 64 | 85 | 84 |
DeepSWE Agentic coding | 46.2 | 18 | 8 | 58 | 70 |
FrontierSWE Frontier engineering | 74.4 | 30.5 | 29.0 | 75.1 | 72.6 |
MCP-Atlas MCP tools | 76.8 | 71.8 | 73.6 | 77.8 | 75.3 |
GLM-5.2 takes both math rows and posts the largest generational jumps on the agent side: Terminal Bench climbs from 63.5 to 81.0, DeepSWE from 18 to 46.2 and FrontierSWE from 30.5 to 74.4 over GLM-5.1. Claude Opus 4.8 leads SWE-bench Pro, Terminal Bench, FrontierSWE and MCP-Atlas, and GPT-5.5 is far ahead on DeepSWE.
Z.ai calls the release a substantial leap in long-horizon task capability over GLM-5.1, and the rest of its table follows that line: SWE-Marathon goes from 1.0 to 13.0, PostTrainBench from 20.1 to 34.3 and NL2Repo from 42.7 to 48.9. Humanity's Last Exam is left out of the table above because Z.ai reports the text-only subset by default and marks full-set results with an asterisk, so the competitor scores in that row are not the same evaluation as GLM-5.2's 40.5. The long-horizon rows are not Z.ai's own runs either: Proximal scored FrontierSWE, Abundant AI scored SWE-Marathon, each at 1M context, max effort level and 128K maximum output tokens.
GLM-5.2 hardware requirements
The system requirement to check is memory, and at this size it is measured in hundreds of gigabytes. Quantized builds come from Unsloth's unsloth/GLM-5.2-GGUF; each build ships as a multi-part download, and the sizes below are the full totals.
| Memory | Build to pick | File size |
|---|---|---|
| 256 GB | UD-Q2_K_XL | 253.9 GB |
| 320 GB | UD-IQ3_S | 308.7 GB |
| 384 GB | UD-IQ4_NL | 372.7 GB |
| 448 GB | UD-Q4_K_S | 436.4 GB |
| 512 GB | UD-Q4_K_XL | 467.3 GB |
| 640 GB | UD-Q6_K | 625.9 GB |
| 768 GB | UD-Q6_K_XL | 684.4 GB |
| 1 TB and up | UD-Q8_K_XL | 819.7 GB |
Every row is the largest build in that repo that fits the memory number next to it: when two builds both fit, take the larger one. The smallest build on offer, UD-IQ1_S, is still a 217 GB download, so this is workstation and server territory, not a laptop model. Sizes are download totals, so leave room above them for context. The unquantized BF16 conversion in the same repo is 1.5 TB across 33 files, the number to plan disk around if you skip quantization.
For the original safetensors weights, Z.ai lists SGLang v0.5.13.post1+, vLLM v0.23.0+, Transformers v0.5.12+, KTransformers v0.5.12+ and Unsloth v0.1.47-beta+ as supported serving stacks, plus vLLM-Ascend, xLLM and SGLang for Ascend NPU hardware. Z.ai ran its reasoning evaluations at temperature 1.0 and top_p 0.95 with generation lengths up to 163,840 tokens, a sane starting point for your own sampling. If the GGUF format is new to you, start with what GGUF is.
How to run GLM-5.2 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for GLM-5.2 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every GLM model you can run locally.
GLM-5.2 license
GLM-5.2 ships under the MIT license, tagged as license: mit on the Hugging Face model card. Z.ai frames it as pure open: an MIT open-source license, no regional limits, technical access without borders. Those are the only license terms the card states for the weights you download and run yourself.
