GLM-5.2

Updated
05.10.2026
Thinking
Reasoning
Code

GLM-5.2 is Z.ai’s 753.3B-parameter flagship with a solid 1M-token context, released under MIT with no regional limits.

At a glance

  • License: MIT
  • Parameters: 753.3B
  • Context length: 1M tokens
  • Modalities: Text (English and Chinese)
  • Minimum hardware: 256 GB of memory (the smallest GGUF build, UD-IQ1_S, is 216.7 GB)

What is GLM-5.2?

GLM-5.2 is the flagship model from Z.ai (zai-org on Hugging Face), built for long-horizon tasks: the long agent runs where a model has to stay coherent across hours of work. It is a 753.3B-parameter model, and for the first time Z.ai delivers that long-horizon capability on a solid 1M-token context, a window it says stably sustains long-horizon work instead of degrading as it fills. The weights landed on June 16, 2026 under the MIT license, and the repo has passed 2.5 million downloads.

SpecificationGLM-5.2
DeveloperZ.ai (zai-org on Hugging Face)
Total parameters753.3B
Context window1M tokens
AttentionSparse attention with IndexShare, one indexer shared across every four layers
Per-token FLOPs at 1M contextCut by 2.9x by IndexShare
Multi-Token PredictionImproved MTP layer, acceptance length up by as much as 20%
CodingMultiple thinking effort levels to balance performance and latency
TaskText generation
ModalitiesText input and output
LanguagesEnglish and Chinese
Serving frameworksSGLang, vLLM, Transformers, KTransformers, Unsloth
Technical reportGLM-5, arXiv 2602.15763
Release dateJune 16, 2026
LicenseMIT

Most of the architecture work targets that 1M window. IndexShare reuses the same indexer across every four sparse attention layers, which cuts per-token FLOPs by 2.9x at a 1M context length, and Z.ai published it as its own paper (arXiv 2603.12201). Z.ai also reworked the MTP layer that drafts tokens for speculative decoding, raising acceptance length by up to 20%. Coding runs with multiple thinking effort levels to balance performance and latency.

GLM-5.2 benchmarks

Z.ai's launch numbers, from the model card, compare GLM-5.2 with its predecessor GLM-5.1, DeepSeek-V4-Pro, Claude Opus 4.8 and GPT-5.5 (the vendor table also includes Qwen3.7-Max, MiniMax M3 and Gemini 3.1 Pro):

BenchmarkGLM-5.2GLM-5.1DeepSeek-V4-ProClaude Opus 4.8GPT-5.5
AIME 2026
Competition math
99.295.394.695.798.3
IMOAnswerBench
Olympiad math
91.083.889.883.5-
SWE-bench Pro
Harder engineering
62.158.455.469.258.6
Terminal Bench 2.1
Terminal agents
81.063.5648584
DeepSWE
Agentic coding
46.21885870
FrontierSWE
Frontier engineering
74.430.529.075.172.6
MCP-Atlas
MCP tools
76.871.873.677.875.3

GLM-5.2 takes both math rows and posts the largest generational jumps on the agent side: Terminal Bench climbs from 63.5 to 81.0, DeepSWE from 18 to 46.2 and FrontierSWE from 30.5 to 74.4 over GLM-5.1. Claude Opus 4.8 leads SWE-bench Pro, Terminal Bench, FrontierSWE and MCP-Atlas, and GPT-5.5 is far ahead on DeepSWE.

Z.ai calls the release a substantial leap in long-horizon task capability over GLM-5.1, and the rest of its table follows that line: SWE-Marathon goes from 1.0 to 13.0, PostTrainBench from 20.1 to 34.3 and NL2Repo from 42.7 to 48.9. Humanity's Last Exam is left out of the table above because Z.ai reports the text-only subset by default and marks full-set results with an asterisk, so the competitor scores in that row are not the same evaluation as GLM-5.2's 40.5. The long-horizon rows are not Z.ai's own runs either: Proximal scored FrontierSWE, Abundant AI scored SWE-Marathon, each at 1M context, max effort level and 128K maximum output tokens.

GLM-5.2 hardware requirements

The system requirement to check is memory, and at this size it is measured in hundreds of gigabytes. Quantized builds come from Unsloth's unsloth/GLM-5.2-GGUF; each build ships as a multi-part download, and the sizes below are the full totals.

MemoryBuild to pickFile size
256 GBUD-Q2_K_XL253.9 GB
320 GBUD-IQ3_S308.7 GB
384 GBUD-IQ4_NL372.7 GB
448 GBUD-Q4_K_S436.4 GB
512 GBUD-Q4_K_XL467.3 GB
640 GBUD-Q6_K625.9 GB
768 GBUD-Q6_K_XL684.4 GB
1 TB and upUD-Q8_K_XL819.7 GB

Every row is the largest build in that repo that fits the memory number next to it: when two builds both fit, take the larger one. The smallest build on offer, UD-IQ1_S, is still a 217 GB download, so this is workstation and server territory, not a laptop model. Sizes are download totals, so leave room above them for context. The unquantized BF16 conversion in the same repo is 1.5 TB across 33 files, the number to plan disk around if you skip quantization.

For the original safetensors weights, Z.ai lists SGLang v0.5.13.post1+, vLLM v0.23.0+, Transformers v0.5.12+, KTransformers v0.5.12+ and Unsloth v0.1.47-beta+ as supported serving stacks, plus vLLM-Ascend, xLLM and SGLang for Ascend NPU hardware. Z.ai ran its reasoning evaluations at temperature 1.0 and top_p 0.95 with generation lengths up to 163,840 tokens, a sane starting point for your own sampling. If the GGUF format is new to you, start with what GGUF is.

How to run GLM-5.2 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for GLM-5.2 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every GLM model you can run locally.

GLM-5.2 license

GLM-5.2 ships under the MIT license, tagged as license: mit on the Hugging Face model card. Z.ai frames it as pure open: an MIT open-source license, no regional limits, technical access without borders. Those are the only license terms the card states for the weights you download and run yourself.

Get the weights from Hugging Face

huggingface-cli download zai-org/GLM-5.2
from transformers import AutoModel
model = AutoModel.from_pretrained("zai-org/GLM-5.2")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

GLM-5.2 is a 753.3B-parameter open-weight model from zai-org (Z.ai), built on a mixture-of-experts architecture and released under the MIT license. It targets coding, reasoning, and agentic tasks, supports English and Chinese, and has a 1,048,576-token context window. The weights are public on Hugging Face, so you can run it yourself instead of through a hosted API.

The full BF16 weights need hundreds of gigabytes of memory, so most people run a quantized build. A 2-bit dynamic quant is around 241 GB and fits a 256 GB unified-memory Mac, or a single 24 GB GPU combined with 256 GB of system RAM using MoE offloading. On consumer hardware with low-bit quants, expect roughly 3 to 9 tokens per second.

Yes. GLM-5.2 is released under the MIT license, which makes the weights free to download, run, modify, and use commercially. Running it locally through Atomic Chat has no per-token cost since the model executes on your own machine. You only pay for the hardware and electricity.

Yes. Once the weights are downloaded, GLM-5.2 runs entirely on your device with no internet connection required. Nothing is sent to an external server, so your prompts and code stay local. Atomic Chat loads the model on-device for this kind of private, offline use.

Its strongest area is coding and agentic engineering, where it ranks among the top open-weight models on benchmarks like SWE-bench Pro and Terminal-Bench. The 1,048,576-token context also makes it good for reasoning over a full codebase or a long technical document. Its thinking capability suits multi-step problems that need the model to work through logic before answering.