diffusiongemma-26B-A4B-it

Updated
05.10.2026
Thinking
Vision
Audio
Reasoning
Code
Multilingual

DiffusionGemma 26B A4B is Google DeepMind’s diffusion text model: a Gemma 4 MoE that denoises 256-token blocks in parallel, 15-20 tokens per pass.

At a glance

  • License: Apache 2.0
  • Parameters: 25.2B total, 3.8B active
  • Context length: Up to 256K tokens
  • Modalities: Text, image and video input, text output
  • Minimum hardware: 24 GB memory for the Q4_K_M build (16.81 GB)

What is diffusiongemma-26B-A4B-it?

diffusiongemma-26B-A4B-it, DiffusionGemma for short, is an open-weights model from Google DeepMind that generates text with discrete diffusion instead of one-token-at-a-time autoregression. It is built on the 26B A4B Mixture-of-Experts Gemma 4 architecture, takes text, image and video input, and denoises whole 256-token blocks in parallel: 15-20 tokens per forward pass, which Google measured at over 1100 tokens per second per user on an H100 in FP8 at low batch size. With 3.8B active parameters out of 25.2B total, and tuning aimed at low-latency, small-batch inference on a single accelerator, it is built for exactly the conditions of a local machine. Google published the weights on Hugging Face in June 2026 under Apache 2.0.

Specificationdiffusiongemma-26B-A4B-it
Total parameters25.2B
Active parameters3.8B
ArchitectureGemma 4 MoE, encoder-decoder, discrete diffusion decoding
Experts8 active of 128, plus 1 shared
Layers30
Context windowUp to 256K tokens
Canvas length256 tokens
ModalitiesText, image and video input, text output
ThinkingConfigurable, enabled with the <|think|> control token
Release dateJune 9, 2026
LicenseApache 2.0

The decode loop is the unusual part. An autoregressive encoder processes the prompt and fills the KV cache, then a decoder with bidirectional attention denoises a 256-token canvas over up to 48 steps, keeping the confident tokens and renoising the rest. Once a canvas settles, it is appended to the cache and the next one begins. It attacks the same sequential bottleneck that speculative decoding targets, and the step count adapts to the task: structured output like code needs fewer denoising steps, so simpler prompts stream faster.

diffusiongemma-26B-A4B-it benchmarks

Google's launch numbers, from the model card, compare the diffusion model with the autoregressive Gemma 4 26B A4B it is built on, both instruction-tuned, with the recommended Entropy Bound sampler:

BenchmarkDiffusionGemma 26B A4BGemma 4 26B A4B
MMLU Pro
Academic knowledge
77.682.6
AIME 2026 (no tools)
Competition math
69.188.3
LiveCodeBench v6
Competitive coding
69.177.1
GPQA Diamond
Expert science
73.282.3
HLE (no tools)
Expert questions
11.08.7
MMMU Pro
Multimodal knowledge
54.373.8

Google's table is candid: Gemma 4 26B A4B stays ahead on every other row too, vision and long context included, and DiffusionGemma's one win is Humanity's Last Exam without tools. You are trading benchmark points for decode speed, and the speed is the point.

diffusiongemma-26B-A4B-it hardware requirements

The system requirement to check is memory. The builds below are quantized from the original weights and published as unsloth/diffusiongemma-26B-A4B-it-GGUF.

MemoryBuild to pickFile size
24 GBQ4_K_M16.81 GB
32 GBQ5_K_M19.15 GB
48 GB and upQ8_026.88 GB

The repo also holds a Q6_K at 22.65 GB and the full BF16 at 50.54 GB. When two builds both fit, take the larger one. Only 3.8B of the 25.2B parameters are active per token, which is what keeps a 26B-class model practical on this kind of hardware. If the GGUF format is new to you, start with what GGUF is.

How to run diffusiongemma-26B-A4B-it in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for diffusiongemma-26B-A4B-it in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Gemma model you can run locally.

diffusiongemma-26B-A4B-it license

diffusiongemma-26B-A4B-it is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download google/diffusiongemma-26B-A4B-it
from transformers import AutoModel
model = AutoModel.from_pretrained("google/diffusiongemma-26B-A4B-it")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is an open-weight text diffusion model from Google DeepMind, based on the Gemma 4 architecture. Rather than writing one token at a time, it denoises a block of up to 256 tokens in parallel, which makes it well suited to code infilling and inline editing. It is a Mixture-of-Experts model with 25.8B total parameters and about 3.8B active per pass.

A 4-bit quant of the 26B-A4B class needs roughly 18GB of VRAM, so a 24GB consumer GPU like an RTX 4090 or 5090 can run it comfortably. The MoE design activates only ~3.8B of its 25.8B parameters per pass, which keeps the memory footprint low for its size. Leave extra room for the KV cache, which grows with longer context.

Yes. It is released under the apache-2.0 license, which allows commercial use, modification, and redistribution at no cost. You can download the weights, run them on your own hardware, and fine-tune the model without paying a usage fee.

Yes. Once you download the weights, it runs fully on your own device with no internet connection. Inside Atomic Chat every prompt and response stays local, so nothing is sent to a server and the model works on a plane or behind a firewall.

Its parallel, bidirectional denoising gives it an edge on tasks that need awareness of surrounding context, such as filling a gap in the middle of a code file or making a quick inline edit to existing text. It also handles structured output like brackets and tags well. On raw benchmark quality it trails standard Gemma 4, so its real advantage is speed and these specific editing tasks.