What is gemma-4-31B-it?
gemma-4-31B-it is the instruction-tuned 31B dense model from Google DeepMind's Gemma 4 family: it takes text and images as input and writes text out. The release spans five sizes (E2B, E4B, 12B, 26B A4B and 31B); Google targets E2B and E4B at mobile and edge devices, and 12B, 26B A4B and 31B at consumer GPUs and workstations. The weights went up on March 11, 2026 under Apache 2.0, with a 256K-token context window.
| Specification | gemma-4-31B-it |
|---|---|
| Total parameters | 30.7B, plus a ~550M vision encoder |
| Architecture | Dense, hybrid attention (sliding-window + global layers) |
| Layers | 60 |
| Sliding window | 1,024 tokens |
| Context window | 256K tokens |
| Vocabulary | 262K tokens |
| Modalities | Text and image input, text output (no audio on this size) |
| Video input | Frames only, up to 60 seconds at 1 fps |
| Image input | Variable aspect ratio, visual token budget 70 to 1,120 |
| Languages | 35+ out of the box, 140+ in pre-training |
| Training data cutoff | January 2025 |
| Reasoning | Thinking mode, toggled with the <|think|> system token |
| Function calling | Native |
| Release date | March 11, 2026 |
| License | Apache 2.0 |
The attention stack interleaves local sliding-window layers (a 1,024 token window) with full global layers, and the final layer is always global. Global layers share unified keys and values and use Proportional RoPE, which is what keeps the memory cost of a 256K-token context in check. Thinking is switchable: put the <|think|> token at the start of the system prompt to enable it, remove it to turn it off. With thinking off the model still emits the thought tags, only with an empty block inside. Pre-training ran on web documents, code, mathematics and images with a cutoff of January 2025 and more than 140 languages.
gemma-4-31B-it benchmarks
Google's launch numbers from the model card compare the 31B with its smaller siblings Gemma 4 26B A4B and Gemma 4 12B, and with last generation's Gemma 3 27B without thinking:
| Benchmark | gemma-4-31B-it | Gemma 4 26B A4B | Gemma 4 12B | Gemma 3 27B (no think) |
|---|---|---|---|---|
MMLU Pro Academic knowledge | 85.2 | 82.6 | 77.2 | 67.6 |
AIME 2026 (no tools) Competition math | 89.2 | 88.3 | 77.5 | 20.8 |
LiveCodeBench v6 Competitive coding | 80.0 | 77.1 | 72.0 | 29.1 |
GPQA Diamond Expert science | 84.3 | 82.3 | 78.8 | 42.4 |
Tau2 Tool use | 76.9 | 68.2 | 69.0 | 16.2 |
MMMU Pro Multimodal understanding | 76.9 | 73.8 | 69.1 | 49.7 |
MRCR v2 (128k) Long context | 66.4 | 44.1 | 43.4 | 13.5 |
The 31B posts the top score in every row of the vendor table. Against the Gemma 3 27B column the row deltas are 17.6 points on MMLU Pro, 27.2 on MMMU Pro, 41.9 on GPQA Diamond, 50.9 on LiveCodeBench v6, 52.9 on MRCR v2 128k, 60.7 on Tau2 and 68.4 on AIME 2026. Codeforces ELO moves 110 to 2150 over the same generation.
The card fills in what those rows do not cover. On images Google claims object detection, document and PDF parsing, screen and UI understanding, chart comprehension, multilingual OCR, handwriting recognition and pointing. The vision rows agree: MATH-Vision 85.6, MedXPertQA MM 61.3, and OmniDocBench 1.5 at 0.131 average edit distance against 0.365 for Gemma 3 27B, where lower is better. Multilingual text holds up too: MMMLU 88.4, BigBench Extra Hard 74.4. The ceiling shows on Humanity's Last Exam: 19.5 without tools, 26.5 with search.
gemma-4-31B-it hardware requirements
The system requirement to check is memory. Quantized builds with real file sizes are published in the unsloth/gemma-4-31B-it-GGUF repository on Hugging Face:
| Memory | Build to pick | File size |
|---|---|---|
| 12 GB | UD-IQ2_M | 10.75 GB |
| 14 GB | UD-IQ3_XXS | 11.84 GB |
| 16 GB | Q3_K_S | 13.21 GB |
| 20 GB | UD-Q3_K_XL | 15.38 GB |
| 24 GB | Q4_K_M | 18.32 GB |
| 32 GB | Q5_K_M | 21.66 GB |
| 40 GB | Q6_K | 25.20 GB |
| 48 GB and up | Q8_0 | 32.64 GB |
Neighbouring builds differ by a gigabyte or two, so when two builds both fit, take the larger one. That matters most at the bottom of the ladder: UD-IQ2_M is 10.75 GB against 13.21 GB for Q3_K_S, so dropping to 2-bit buys back only 2.46 GB, and the smallest build in the repo, UD-IQ2_XXS, is 8.53 GB. Image input needs the mmproj file from the same repo on top of the weights, another 1.20 GB. If the format is new to you, start with what GGUF is.
Google's own serving path is Transformers, via AutoProcessor and AutoModelForMultimodalLM, with thinking turned on by enable_thinking=True; Transformers and llama.cpp both apply the chat template for you. The card asks for one sampling config in every use case: temperature 1.0, top_p 0.95, top_k 64. Put images before the text, and in multi-turn chats keep old thinking blocks out of the history, tool call turns excepted.
How to run gemma-4-31B-it in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for gemma-4-31B-it in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the lineup, see every Gemma model you can run locally.
gemma-4-31B-it license
gemma-4-31B-it is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.
