What is GLM-5.1?
GLM-5.1 is a 753.9B-parameter model from Z.ai (zai-org on Hugging Face), the successor to GLM-5 and the flagship of the family, built for agentic engineering. The pitch is endurance rather than first-pass answers: the model is built to break a problem down, run experiments, read the results and revise its strategy over hundreds of rounds and thousands of tool calls. Z.ai published the weights on Hugging Face on April 3, 2026 under the MIT license.
| Specification | GLM-5.1 |
|---|---|
| Total parameters | 753.9B (753,864,139,008) |
| Transformers integration | glm_moe_dsa model doc |
| Predecessor | GLM-5 |
| Focus | Agentic engineering and coding |
| Languages | English and Chinese |
| Modality | Text generation |
| Weight format | safetensors |
| Local serving | SGLang, vLLM, xLLM, Transformers, KTransformers |
| Hosted access | Z.ai API Platform |
| Technical report | GLM-5: from Vibe Coding to Agentic Engineering (arXiv 2602.15763) |
| Release date | April 3, 2026 |
| License | MIT |
The claim that matters is what happens after the first hour. Z.ai says earlier models, GLM-5 included, apply familiar techniques for quick initial gains and then plateau, and that giving them more time does not help, while GLM-5.1 handles ambiguous problems with better judgment, identifies blockers with real precision and keeps improving the longer it runs. On the launch numbers that shows up as state of the art on SWE-Bench Pro and a wide lead over GLM-5 on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks). The weights ship as safetensors and the vendor lists no context window on the model card.
GLM-5.1 benchmarks
Z.ai's launch table, from the model card, compares GLM-5.1 with eight other models, among them Qwen3.6-Plus, MiniMax M2.7, DeepSeek-V3.2 and Kimi K2.5; here it is against its predecessor GLM-5 and the closed frontier trio of Claude Opus 4.6, Gemini 3.1 Pro and GPT-5.4:
| Benchmark | GLM-5.1 | GLM-5 | Claude Opus 4.6 | Gemini 3.1 Pro | GPT-5.4 |
|---|---|---|---|---|---|
SWE-Bench Pro Harder engineering | 58.4 | 55.1 | 57.3 | 54.2 | 57.7 |
CyberGym Security reasoning | 68.7 | 48.3 | 66.6 | 38.8 | 66.3 |
BrowseComp Web research | 68.0 | 62.0 | - | - | - |
NL2Repo Repo-level coding | 42.7 | 35.9 | 49.8 | 33.4 | 41.3 |
Terminal-Bench 2.0 Terminal agents | 63.5 | 56.2 | 65.4 | 68.5 | - |
τ³-Bench Tool agents | 70.6 | 69.2 | 72.4 | 67.1 | 72.9 |
HLE Expert questions | 31.0 | 30.5 | 36.7 | 45.0 | 39.8 |
GLM-5.1 takes SWE-Bench Pro, CyberGym and BrowseComp outright and improves on GLM-5 in every row shown. The closed models keep NL2Repo, the terminal run, tool use and HLE, so this is parity on agentic coding, not a clean sweep.
The same table separates plain runs from tool-assisted ones, and the gap between the two columns is wide: HLE goes from 31.0 to 52.3 once tools are allowed, and BrowseComp from 68.0 to 79.3 with context management, though Gemini 3.1 Pro leads that second column at 85.9. On Terminal-Bench 2.0 the best self-reported figure is 69.0 driving Claude Code, against 56.2 for GLM-5 on the same harness and 75.1 for GPT-5.4 on Codex. On the tool rows it sits mid-pack: 71.8 on MCP-Atlas (Public Set), where Qwen3.6-Plus leads at 74.1, and 40.7 on Tool-Decathlon, where GPT-5.4 leads at 54.6. On the math and knowledge rows the picture is flatter: 95.3 on AIME 2026 and 86.2 on GPQA-Diamond, both within a few tenths of GLM-5. Vending Bench 2 pays out $5,634.41, against $8,017.59 for Claude Opus 4.6 at the top of that row.
GLM-5.1 hardware requirements
The system requirement to check is memory: at 753.9B parameters even the smallest usable build is over 200 GB. Z.ai ships safetensors and names five serving stacks with a floor version for each: SGLang v0.5.10, vLLM v0.19.0, xLLM v0.8.0, and Transformers or KTransformers v0.5.3, with the Transformers path documented under the glm_moe_dsa model doc. The GGUF builds below come from unsloth/GLM-5.1-GGUF, which carries 22 quantizations, from UD-IQ1_M at 205.5 GB up to UD-Q8_K_XL at 810.6 GB, plus an unquantized BF16 conversion at 1.5 TB; each size is the total of the split files in that folder.
| Memory | Build to pick | File size |
|---|---|---|
| 256 GB | UD-IQ2_XXS | 220.6 GB |
| 320 GB | UD-IQ3_S | 279.6 GB |
| 384 GB | UD-Q3_K_XL | 340.1 GB |
| 512 GB | UD-Q4_K_XL | 466.0 GB |
| 640 GB | UD-Q5_K_XL | 560.1 GB |
| 768 GB | UD-Q6_K_XL | 684.3 GB |
| 1 TB and up | UD-Q8_K_XL | 810.6 GB |
The floor is UD-IQ1_M at 205.5 GB, only 15 GB under UD-IQ2_XXS, so keep it for a machine that cannot hold anything else. When two builds both fit, take the larger one, and leave headroom for the KV cache on top of the file size. If the format is new to you, start with what GGUF is.
How to run GLM-5.1 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for GLM-5.1 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every GLM model you can run locally, or the sibling GLM-5.2.
GLM-5.1 license
GLM-5.1 is released under the MIT license, one of the most permissive there is. It allows commercial use, modification, fine-tuning and redistribution with no royalties and almost no conditions beyond keeping the copyright notice, so you can build products on the model and run it on your own hardware without a usage fee.
