What is Gemma 4 26B A4B?
Gemma 4 26B A4B is the Mixture-of-Experts model in Google DeepMind's Gemma 4 family, published in instruction-tuned form as gemma-4-26B-A4B-it. It carries 25.2B total parameters, but only 3.8B are active for any given token, so Google positions it as the fast option next to the dense 31B: close to that model's scores, at the decode speed of a model near 4B. The weights went up on Hugging Face on March 11, 2026 under Apache 2.0.
| Specification | Gemma 4 26B A4B |
|---|---|
| Total parameters | 25.2B |
| Active parameters | 3.8B |
| Experts | 8 active of 128, plus 1 shared |
| Layers | 30 |
| Attention | Local sliding window (1024 tokens) interleaved with full global layers |
| Context window | 256K tokens |
| Modalities | Text and image input, text output |
| Vision encoder | ~550M parameters |
| Languages | 35+ out of the box, pre-trained on 140+ |
| Reasoning | Thinking mode, toggled with a system prompt token |
| Function calling | Native support |
| Release date | March 11, 2026 |
| License | Apache 2.0 |
The attention layout is what keeps memory in check at long context: most layers only look back over a 1024-token sliding window, full global layers are interleaved between them with the final layer always global, and the global layers use unified keys and values. Gemma 4 also adds native support for the system role, and the whole family is trained to think step by step, a mode you can switch on or off per request. Two limits worth knowing at this size: audio input is reserved for the E2B, E4B and 12B models, and output is text only.
Gemma 4 26B A4B benchmarks
Google's launch numbers, from the model card, compare the 26B A4B with its dense siblings Gemma 4 31B and Gemma 4 12B Unified, and with the previous-generation Gemma 3 27B run without thinking:
| Benchmark | Gemma 4 26B A4B | Gemma 4 31B | Gemma 4 12B Unified | Gemma 3 27B (no think) |
|---|---|---|---|---|
MMLU Pro Academic knowledge | 82.6% | 85.2% | 77.2% | 67.6% |
AIME 2026 (no tools) Competition math | 88.3% | 89.2% | 77.5% | 20.8% |
LiveCodeBench v6 Competitive coding | 77.1% | 80.0% | 72.0% | 29.1% |
GPQA Diamond Expert science | 82.3% | 84.3% | 78.8% | 42.4% |
Tau2 (average over 3) Tool use | 68.2% | 76.9% | 69.0% | 16.2% |
MMMU Pro Multimodal reasoning | 73.8% | 76.9% | 69.1% | 49.7% |
The dense 31B leads every row, but on knowledge, math, code and science the 26B A4B lands within one to three points of it while activating a fraction of the parameters, and it clears Gemma 3 27B by a wide margin everywhere. The one soft spot is Tau2, where even the 12B Unified comes out slightly ahead.
Gemma 4 26B A4B hardware requirements
The system requirement to check is memory: the model file has to fit in your RAM or VRAM with room left over for context. The builds below are Unsloth's dynamic quants from unsloth/gemma-4-26B-A4B-it-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 12 GB | UD-Q2_K_XL | 10.55 GB |
| 16 GB | UD-IQ4_NL | 13.61 GB |
| 24 GB | UD-Q4_K_XL | 17.01 GB |
| 32 GB | UD-Q5_K_XL | 21.22 GB |
| 48 GB and up | UD-Q8_K_XL | 27.64 GB |
When two builds both fit, take the larger one. Since only 3.8B parameters are active per token, even the biggest file here decodes fast for its size. If the format is new to you, start with what GGUF is.
How to run Gemma 4 26B A4B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Gemma 4 26B A4B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Gemma model you can run locally.
Gemma 4 26B A4B license
Gemma 4 26B A4B is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.
