What is diffusiongemma-26B-A4B-it?
diffusiongemma-26B-A4B-it, DiffusionGemma for short, is an open-weights model from Google DeepMind that generates text with discrete diffusion instead of one-token-at-a-time autoregression. It is built on the 26B A4B Mixture-of-Experts Gemma 4 architecture, takes text, image and video input, and denoises whole 256-token blocks in parallel: 15-20 tokens per forward pass, which Google measured at over 1100 tokens per second per user on an H100 in FP8 at low batch size. With 3.8B active parameters out of 25.2B total, and tuning aimed at low-latency, small-batch inference on a single accelerator, it is built for exactly the conditions of a local machine. Google published the weights on Hugging Face in June 2026 under Apache 2.0.
| Specification | diffusiongemma-26B-A4B-it |
|---|---|
| Total parameters | 25.2B |
| Active parameters | 3.8B |
| Architecture | Gemma 4 MoE, encoder-decoder, discrete diffusion decoding |
| Experts | 8 active of 128, plus 1 shared |
| Layers | 30 |
| Context window | Up to 256K tokens |
| Canvas length | 256 tokens |
| Modalities | Text, image and video input, text output |
| Thinking | Configurable, enabled with the <|think|> control token |
| Release date | June 9, 2026 |
| License | Apache 2.0 |
The decode loop is the unusual part. An autoregressive encoder processes the prompt and fills the KV cache, then a decoder with bidirectional attention denoises a 256-token canvas over up to 48 steps, keeping the confident tokens and renoising the rest. Once a canvas settles, it is appended to the cache and the next one begins. It attacks the same sequential bottleneck that speculative decoding targets, and the step count adapts to the task: structured output like code needs fewer denoising steps, so simpler prompts stream faster.
diffusiongemma-26B-A4B-it benchmarks
Google's launch numbers, from the model card, compare the diffusion model with the autoregressive Gemma 4 26B A4B it is built on, both instruction-tuned, with the recommended Entropy Bound sampler:
| Benchmark | DiffusionGemma 26B A4B | Gemma 4 26B A4B |
|---|---|---|
MMLU Pro Academic knowledge | 77.6 | 82.6 |
AIME 2026 (no tools) Competition math | 69.1 | 88.3 |
LiveCodeBench v6 Competitive coding | 69.1 | 77.1 |
GPQA Diamond Expert science | 73.2 | 82.3 |
HLE (no tools) Expert questions | 11.0 | 8.7 |
MMMU Pro Multimodal knowledge | 54.3 | 73.8 |
Google's table is candid: Gemma 4 26B A4B stays ahead on every other row too, vision and long context included, and DiffusionGemma's one win is Humanity's Last Exam without tools. You are trading benchmark points for decode speed, and the speed is the point.
diffusiongemma-26B-A4B-it hardware requirements
The system requirement to check is memory. The builds below are quantized from the original weights and published as unsloth/diffusiongemma-26B-A4B-it-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 24 GB | Q4_K_M | 16.81 GB |
| 32 GB | Q5_K_M | 19.15 GB |
| 48 GB and up | Q8_0 | 26.88 GB |
The repo also holds a Q6_K at 22.65 GB and the full BF16 at 50.54 GB. When two builds both fit, take the larger one. Only 3.8B of the 25.2B parameters are active per token, which is what keeps a 26B-class model practical on this kind of hardware. If the GGUF format is new to you, start with what GGUF is.
How to run diffusiongemma-26B-A4B-it in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for diffusiongemma-26B-A4B-it in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Gemma model you can run locally.
diffusiongemma-26B-A4B-it license
diffusiongemma-26B-A4B-it is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.
