gemma-4-26B-A4B-it

Updated
05.10.2026
Thinking
Embedding
Vision
Audio
Reasoning
Code
Multilingual

Google’s Gemma 4 MoE: 25.2B parameters with 3.8B active, 256K context and image input, Apache 2.0. Runs at the speed of a model near 4B.

At a glance

  • License: Apache 2.0
  • Parameters: 25.2B total, 3.8B active
  • Context length: 256K tokens
  • Modalities: Text and image input, text output
  • Minimum hardware: 12 GB of RAM or VRAM (smallest GGUF is 9.92 GB)

What is Gemma 4 26B A4B?

Gemma 4 26B A4B is the Mixture-of-Experts model in Google DeepMind's Gemma 4 family, published in instruction-tuned form as gemma-4-26B-A4B-it. It carries 25.2B total parameters, but only 3.8B are active for any given token, so Google positions it as the fast option next to the dense 31B: close to that model's scores, at the decode speed of a model near 4B. The weights went up on Hugging Face on March 11, 2026 under Apache 2.0.

SpecificationGemma 4 26B A4B
Total parameters25.2B
Active parameters3.8B
Experts8 active of 128, plus 1 shared
Layers30
AttentionLocal sliding window (1024 tokens) interleaved with full global layers
Context window256K tokens
ModalitiesText and image input, text output
Vision encoder~550M parameters
Languages35+ out of the box, pre-trained on 140+
ReasoningThinking mode, toggled with a system prompt token
Function callingNative support
Release dateMarch 11, 2026
LicenseApache 2.0

The attention layout is what keeps memory in check at long context: most layers only look back over a 1024-token sliding window, full global layers are interleaved between them with the final layer always global, and the global layers use unified keys and values. Gemma 4 also adds native support for the system role, and the whole family is trained to think step by step, a mode you can switch on or off per request. Two limits worth knowing at this size: audio input is reserved for the E2B, E4B and 12B models, and output is text only.

Gemma 4 26B A4B benchmarks

Google's launch numbers, from the model card, compare the 26B A4B with its dense siblings Gemma 4 31B and Gemma 4 12B Unified, and with the previous-generation Gemma 3 27B run without thinking:

BenchmarkGemma 4 26B A4BGemma 4 31BGemma 4 12B UnifiedGemma 3 27B (no think)
MMLU Pro
Academic knowledge
82.6%85.2%77.2%67.6%
AIME 2026 (no tools)
Competition math
88.3%89.2%77.5%20.8%
LiveCodeBench v6
Competitive coding
77.1%80.0%72.0%29.1%
GPQA Diamond
Expert science
82.3%84.3%78.8%42.4%
Tau2 (average over 3)
Tool use
68.2%76.9%69.0%16.2%
MMMU Pro
Multimodal reasoning
73.8%76.9%69.1%49.7%

The dense 31B leads every row, but on knowledge, math, code and science the 26B A4B lands within one to three points of it while activating a fraction of the parameters, and it clears Gemma 3 27B by a wide margin everywhere. The one soft spot is Tau2, where even the 12B Unified comes out slightly ahead.

Gemma 4 26B A4B hardware requirements

The system requirement to check is memory: the model file has to fit in your RAM or VRAM with room left over for context. The builds below are Unsloth's dynamic quants from unsloth/gemma-4-26B-A4B-it-GGUF.

MemoryBuild to pickFile size
12 GBUD-Q2_K_XL10.55 GB
16 GBUD-IQ4_NL13.61 GB
24 GBUD-Q4_K_XL17.01 GB
32 GBUD-Q5_K_XL21.22 GB
48 GB and upUD-Q8_K_XL27.64 GB

When two builds both fit, take the larger one. Since only 3.8B parameters are active per token, even the biggest file here decodes fast for its size. If the format is new to you, start with what GGUF is.

How to run Gemma 4 26B A4B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Gemma 4 26B A4B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Gemma model you can run locally.

Gemma 4 26B A4B license

Gemma 4 26B A4B is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download google/gemma-4-26B-A4B-it
from transformers import AutoModel
model = AutoModel.from_pretrained("google/gemma-4-26B-A4B-it")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is an instruction-tuned Mixture-of-Experts model from Google in the Gemma 4 family. It has 26.5B total parameters but activates only about 4B per token through top-8 routing over 128 experts, so it runs roughly as fast as a 4B model. It is multimodal, supports a 262144-token context, and handles code, reasoning, and vision.

A 4-bit (Q4_K_M) build of gemma-4-26B-A4B-it needs around 18GB of VRAM, which a single 24GB consumer GPU can hold. Apple Silicon Macs with enough unified memory can run it too. Leave extra headroom because the long context window expands the KV cache during longer chats.

Yes. The weights are published under the apache-2.0 license, which allows free personal and commercial use, modification, and redistribution. Atomic Chat, the app that runs it on atomic.chat, is also free and open-source.

Yes. After you download the weights once, gemma-4-26B-A4B-it runs entirely on your machine with no internet connection required. Prompts, documents, and images never leave your device, which is the point of running it locally in Atomic Chat.

It is strongest at coding, multi-step reasoning, and agentic tool-use, and it has a configurable thinking mode that reasons before answering. Its vision support lets it analyze images and short video, and the 262144-token context makes it useful for long documents and large codebases.