What is GLM-4.7-Flash?
GLM-4.7-Flash is a mixture of experts model from Z.ai. It carries 31.2B total parameters but activates only about 3B per token, which is what the 30B-A3B label means. Z.ai calls it the strongest model in the 30B class and presents it as a new option for lightweight deployment that balances performance and efficiency. The weights went up on Hugging Face on January 19, 2026 under the MIT license, and the card names vLLM and SGLang for local serving, plus a transformers example for running the weights directly.
| Specification | GLM-4.7-Flash |
|---|---|
| Total parameters | 31.2B (31,221,488,576 exactly) |
| Active parameters | About 3B per token (30B-A3B) |
| Architecture | Mixture of experts (MoE) |
| Task | Text generation, text in and out |
| Languages | English and Chinese |
| Reasoning | Thinking mode; Preserved Thinking recommended for multi-turn agent tasks |
| Speculative decoding | MTP (vLLM) and EAGLE (SGLang) in the vendor serve configs |
| Vendor sampling defaults | Temperature 1.0, top-p 0.95 |
| Max new tokens in vendor evals | 131,072 |
| Serving stacks | vLLM and SGLang (main branches only), plus Transformers |
| Parsers | glm47 tool-call parser, glm45 reasoning parser |
| Release date | January 19, 2026 |
| License | MIT |
Z.ai's own serve commands ship with speculative decoding switched on: MTP under vLLM with one speculative token, and EAGLE under SGLang with three speculative steps and four draft tokens. That is the mechanism covered in our speculative decoding guide. Both commands also name the same two parsers, glm47 for tool calls and glm45 for reasoning, and the vLLM one turns on automatic tool choice as well. For multi-turn agentic tasks, which Z.ai lists as τ²-Bench and Terminal Bench 2, the card says to turn on Preserved Thinking mode.
GLM-4.7-Flash benchmarks
Z.ai's launch numbers, from the model card, compare GLM-4.7-Flash with Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B:
| Benchmark | GLM-4.7-Flash | Qwen3-30B-A3B-Thinking-2507 | GPT-OSS-20B |
|---|---|---|---|
AIME 25 Competition math | 91.6 | 85.0 | 91.7 |
GPQA Expert science | 75.2 | 73.4 | 71.5 |
LCB v6 Competitive coding | 64.0 | 66.0 | 61.0 |
HLE Expert questions | 14.4 | 9.8 | 10.9 |
SWE-bench Verified Software engineering | 59.2 | 22.0 | 34.0 |
τ²-Bench Agentic tools | 79.5 | 49.0 | 47.7 |
BrowseComp Web browsing | 42.8 | 2.29 | 28.3 |
GLM-4.7-Flash takes five of the seven rows, with the largest margins on SWE-bench Verified, τ²-Bench and BrowseComp. It gives up AIME 25 to GPT-OSS-20B by a tenth of a point and LCB v6 to the Qwen model by two.
Z.ai did not run every row at the same settings. Terminal Bench and SWE-bench Verified were scored at temperature 0.7, top-p 1.0 and a 16,384-token generation cap, and τ²-Bench at temperature 0 with the same cap, rather than at the defaults in the table above. For τ²-Bench it also added a prompt to the Retail and Telecom user interaction to avoid failures caused by users ending the interaction incorrectly, and applied the Airline domain fixes from the Claude Opus 4.5 release report. Z.ai also says to run that benchmark and Terminal Bench 2 with Preserved Thinking on, so the agentic scores describe a specific setup, not a plain chat.
GLM-4.7-Flash hardware requirements
The system requirement to check is memory. The GGUF builds below come from the community repo unsloth/GLM-4.7-Flash-GGUF, and the sizes are the real file sizes from that listing.
| Memory | Build to pick | File size |
|---|---|---|
| 10 GB | UD-TQ1_0 | 8.33 GB |
| 12 GB | UD-IQ2_M | 10.99 GB |
| 16 GB | UD-Q3_K_XL | 13.78 GB |
| 20 GB | IQ4_XS | 16.27 GB |
| 24 GB | Q4_K_M | 18.31 GB |
| 32 GB | Q6_K | 24.69 GB |
| 48 GB and up | Q8_0 | 31.84 GB |
Neighbouring files differ by a gigabyte or two, so when two builds both fit, take the larger one. The listing goes further down than the table does, to a 9.25 GB UD-IQ1_S and a 9.81 GB UD-IQ1_M; reach for those 1-bit builds only when nothing else fits. The unquantized BF16 weights are also there, 59.91 GB across two files. If the format is new to you, start with what GGUF is.
Serving the full weights on GPUs is the other path: Z.ai's reference commands run at tensor parallel size 4 under both vLLM and SGLang, and both frameworks support the model only on their main branches. On Blackwell cards the SGLang command also needs the triton attention backends.
How to run GLM-4.7-Flash in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for GLM-4.7-Flash in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
Leave the sampler at Z.ai's defaults for most tasks, temperature 1.0 with top-p 0.95, and give thinking room: the vendor evals allow up to 131,072 new tokens per answer. For the rest of the family, see every GLM model you can run locally.
GLM-4.7-Flash license
GLM-4.7-Flash is released under the MIT license. That permits commercial use, modification and redistribution, provided the copyright and license notice travel with the software, which makes it one of the most permissive licenses an open-weight model can carry.
