GLM-5.1

Updated
05.10.2026
Tools
Thinking
Reasoning
Code

GLM-5.1 is Z.ai’s 753.9B-parameter MIT-licensed flagship for agentic engineering, topping SWE-Bench Pro at 58.4. Run it locally in Atomic Chat.

At a glance

  • License: MIT
  • Parameters: 753.9B total (753,864,139,008)
  • Context length: Not stated on the vendor model card
  • Modalities: Text in, text out (English and Chinese)
  • Minimum hardware: 256 GB of memory; the smallest GGUF build is 205.5 GB

What is GLM-5.1?

GLM-5.1 is a 753.9B-parameter model from Z.ai (zai-org on Hugging Face), the successor to GLM-5 and the flagship of the family, built for agentic engineering. The pitch is endurance rather than first-pass answers: the model is built to break a problem down, run experiments, read the results and revise its strategy over hundreds of rounds and thousands of tool calls. Z.ai published the weights on Hugging Face on April 3, 2026 under the MIT license.

SpecificationGLM-5.1
Total parameters753.9B (753,864,139,008)
Transformers integrationglm_moe_dsa model doc
PredecessorGLM-5
FocusAgentic engineering and coding
LanguagesEnglish and Chinese
ModalityText generation
Weight formatsafetensors
Local servingSGLang, vLLM, xLLM, Transformers, KTransformers
Hosted accessZ.ai API Platform
Technical reportGLM-5: from Vibe Coding to Agentic Engineering (arXiv 2602.15763)
Release dateApril 3, 2026
LicenseMIT

The claim that matters is what happens after the first hour. Z.ai says earlier models, GLM-5 included, apply familiar techniques for quick initial gains and then plateau, and that giving them more time does not help, while GLM-5.1 handles ambiguous problems with better judgment, identifies blockers with real precision and keeps improving the longer it runs. On the launch numbers that shows up as state of the art on SWE-Bench Pro and a wide lead over GLM-5 on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks). The weights ship as safetensors and the vendor lists no context window on the model card.

GLM-5.1 benchmarks

Z.ai's launch table, from the model card, compares GLM-5.1 with eight other models, among them Qwen3.6-Plus, MiniMax M2.7, DeepSeek-V3.2 and Kimi K2.5; here it is against its predecessor GLM-5 and the closed frontier trio of Claude Opus 4.6, Gemini 3.1 Pro and GPT-5.4:

BenchmarkGLM-5.1GLM-5Claude Opus 4.6Gemini 3.1 ProGPT-5.4
SWE-Bench Pro
Harder engineering
58.455.157.354.257.7
CyberGym
Security reasoning
68.748.366.638.866.3
BrowseComp
Web research
68.062.0---
NL2Repo
Repo-level coding
42.735.949.833.441.3
Terminal-Bench 2.0
Terminal agents
63.556.265.468.5-
τ³-Bench
Tool agents
70.669.272.467.172.9
HLE
Expert questions
31.030.536.745.039.8

GLM-5.1 takes SWE-Bench Pro, CyberGym and BrowseComp outright and improves on GLM-5 in every row shown. The closed models keep NL2Repo, the terminal run, tool use and HLE, so this is parity on agentic coding, not a clean sweep.

The same table separates plain runs from tool-assisted ones, and the gap between the two columns is wide: HLE goes from 31.0 to 52.3 once tools are allowed, and BrowseComp from 68.0 to 79.3 with context management, though Gemini 3.1 Pro leads that second column at 85.9. On Terminal-Bench 2.0 the best self-reported figure is 69.0 driving Claude Code, against 56.2 for GLM-5 on the same harness and 75.1 for GPT-5.4 on Codex. On the tool rows it sits mid-pack: 71.8 on MCP-Atlas (Public Set), where Qwen3.6-Plus leads at 74.1, and 40.7 on Tool-Decathlon, where GPT-5.4 leads at 54.6. On the math and knowledge rows the picture is flatter: 95.3 on AIME 2026 and 86.2 on GPQA-Diamond, both within a few tenths of GLM-5. Vending Bench 2 pays out $5,634.41, against $8,017.59 for Claude Opus 4.6 at the top of that row.

GLM-5.1 hardware requirements

The system requirement to check is memory: at 753.9B parameters even the smallest usable build is over 200 GB. Z.ai ships safetensors and names five serving stacks with a floor version for each: SGLang v0.5.10, vLLM v0.19.0, xLLM v0.8.0, and Transformers or KTransformers v0.5.3, with the Transformers path documented under the glm_moe_dsa model doc. The GGUF builds below come from unsloth/GLM-5.1-GGUF, which carries 22 quantizations, from UD-IQ1_M at 205.5 GB up to UD-Q8_K_XL at 810.6 GB, plus an unquantized BF16 conversion at 1.5 TB; each size is the total of the split files in that folder.

MemoryBuild to pickFile size
256 GBUD-IQ2_XXS220.6 GB
320 GBUD-IQ3_S279.6 GB
384 GBUD-Q3_K_XL340.1 GB
512 GBUD-Q4_K_XL466.0 GB
640 GBUD-Q5_K_XL560.1 GB
768 GBUD-Q6_K_XL684.3 GB
1 TB and upUD-Q8_K_XL810.6 GB

The floor is UD-IQ1_M at 205.5 GB, only 15 GB under UD-IQ2_XXS, so keep it for a machine that cannot hold anything else. When two builds both fit, take the larger one, and leave headroom for the KV cache on top of the file size. If the format is new to you, start with what GGUF is.

How to run GLM-5.1 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for GLM-5.1 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every GLM model you can run locally, or the sibling GLM-5.2.

GLM-5.1 license

GLM-5.1 is released under the MIT license, one of the most permissive there is. It allows commercial use, modification, fine-tuning and redistribution with no royalties and almost no conditions beyond keeping the copyright notice, so you can build products on the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download zai-org/GLM-5.1
from transformers import AutoModel
model = AutoModel.from_pretrained("zai-org/GLM-5.1")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

GLM-5.1 is an open-weight large language model from zai-org (Z.AI), built as a 753.9B-parameter Mixture-of-Experts with about 40B parameters active per token. It is tuned for agentic coding and long-horizon reasoning, and it scores 58.4 on SWE-Bench Pro. The weights are public on Hugging Face under the MIT license, so it can run locally in apps like Atomic Chat.

The full FP8 checkpoint needs roughly 860GB of memory, and full precision exceeds 1.5TB, which means a multi-GPU rig or cluster. Quantized GGUF builds cut this down a lot: 4-bit is around 476GB and a 2-bit dynamic quant fits in about 240GB, runnable on a 256GB Mac Studio or a multi-card workstation. On consumer hardware expect a few tokens per second.

Yes. GLM-5.1 is released under the MIT license, so you can download, fine-tune, and deploy it at no cost. The license also permits commercial use with no royalties or usage restrictions. The only practical cost is the hardware needed to run a model of this size.

Yes, once the weights are downloaded the model runs entirely on your own machine with no internet connection. Loading it through Atomic Chat keeps every prompt and response on-device, so nothing is sent to an external server. The main constraint is having enough memory for the model or a sufficiently quantized version.

It is strongest at agentic coding: working across a repository, calling tools, running experiments, and solving multi-step problems, which is reflected in its SWE-Bench Pro score. Native tool calling and a 198K context window also make it well suited to long-context reasoning over large codebases or document sets. Thinking mode is enabled by default for these tasks.