gemma-4-31B-it

Updated
05.10.2026
Thinking
Embedding
Vision
Audio
Reasoning
Code
Multilingual

Gemma 4 31B is Google DeepMind’s dense 30.7B open model: text and image input, a 256K-token context window, and an Apache 2.0 license.

At a glance

  • License: Apache 2.0
  • Parameters: 30.7B dense
  • Context length: 256K tokens
  • Modalities: Text and image input, text output
  • Minimum hardware: 12 GB of memory (2-bit GGUF, 10.75 GB)

What is gemma-4-31B-it?

gemma-4-31B-it is the instruction-tuned 31B dense model from Google DeepMind's Gemma 4 family: it takes text and images as input and writes text out. The release spans five sizes (E2B, E4B, 12B, 26B A4B and 31B); Google targets E2B and E4B at mobile and edge devices, and 12B, 26B A4B and 31B at consumer GPUs and workstations. The weights went up on March 11, 2026 under Apache 2.0, with a 256K-token context window.

Specificationgemma-4-31B-it
Total parameters30.7B, plus a ~550M vision encoder
ArchitectureDense, hybrid attention (sliding-window + global layers)
Layers60
Sliding window1,024 tokens
Context window256K tokens
Vocabulary262K tokens
ModalitiesText and image input, text output (no audio on this size)
Video inputFrames only, up to 60 seconds at 1 fps
Image inputVariable aspect ratio, visual token budget 70 to 1,120
Languages35+ out of the box, 140+ in pre-training
Training data cutoffJanuary 2025
ReasoningThinking mode, toggled with the <|think|> system token
Function callingNative
Release dateMarch 11, 2026
LicenseApache 2.0

The attention stack interleaves local sliding-window layers (a 1,024 token window) with full global layers, and the final layer is always global. Global layers share unified keys and values and use Proportional RoPE, which is what keeps the memory cost of a 256K-token context in check. Thinking is switchable: put the <|think|> token at the start of the system prompt to enable it, remove it to turn it off. With thinking off the model still emits the thought tags, only with an empty block inside. Pre-training ran on web documents, code, mathematics and images with a cutoff of January 2025 and more than 140 languages.

gemma-4-31B-it benchmarks

Google's launch numbers from the model card compare the 31B with its smaller siblings Gemma 4 26B A4B and Gemma 4 12B, and with last generation's Gemma 3 27B without thinking:

Benchmarkgemma-4-31B-itGemma 4 26B A4BGemma 4 12BGemma 3 27B (no think)
MMLU Pro
Academic knowledge
85.282.677.267.6
AIME 2026 (no tools)
Competition math
89.288.377.520.8
LiveCodeBench v6
Competitive coding
80.077.172.029.1
GPQA Diamond
Expert science
84.382.378.842.4
Tau2
Tool use
76.968.269.016.2
MMMU Pro
Multimodal understanding
76.973.869.149.7
MRCR v2 (128k)
Long context
66.444.143.413.5

The 31B posts the top score in every row of the vendor table. Against the Gemma 3 27B column the row deltas are 17.6 points on MMLU Pro, 27.2 on MMMU Pro, 41.9 on GPQA Diamond, 50.9 on LiveCodeBench v6, 52.9 on MRCR v2 128k, 60.7 on Tau2 and 68.4 on AIME 2026. Codeforces ELO moves 110 to 2150 over the same generation.

The card fills in what those rows do not cover. On images Google claims object detection, document and PDF parsing, screen and UI understanding, chart comprehension, multilingual OCR, handwriting recognition and pointing. The vision rows agree: MATH-Vision 85.6, MedXPertQA MM 61.3, and OmniDocBench 1.5 at 0.131 average edit distance against 0.365 for Gemma 3 27B, where lower is better. Multilingual text holds up too: MMMLU 88.4, BigBench Extra Hard 74.4. The ceiling shows on Humanity's Last Exam: 19.5 without tools, 26.5 with search.

gemma-4-31B-it hardware requirements

The system requirement to check is memory. Quantized builds with real file sizes are published in the unsloth/gemma-4-31B-it-GGUF repository on Hugging Face:

MemoryBuild to pickFile size
12 GBUD-IQ2_M10.75 GB
14 GBUD-IQ3_XXS11.84 GB
16 GBQ3_K_S13.21 GB
20 GBUD-Q3_K_XL15.38 GB
24 GBQ4_K_M18.32 GB
32 GBQ5_K_M21.66 GB
40 GBQ6_K25.20 GB
48 GB and upQ8_032.64 GB

Neighbouring builds differ by a gigabyte or two, so when two builds both fit, take the larger one. That matters most at the bottom of the ladder: UD-IQ2_M is 10.75 GB against 13.21 GB for Q3_K_S, so dropping to 2-bit buys back only 2.46 GB, and the smallest build in the repo, UD-IQ2_XXS, is 8.53 GB. Image input needs the mmproj file from the same repo on top of the weights, another 1.20 GB. If the format is new to you, start with what GGUF is.

Google's own serving path is Transformers, via AutoProcessor and AutoModelForMultimodalLM, with thinking turned on by enable_thinking=True; Transformers and llama.cpp both apply the chat template for you. The card asks for one sampling config in every use case: temperature 1.0, top_p 0.95, top_k 64. Put images before the text, and in multi-turn chats keep old thinking blocks out of the history, tool call turns excepted.

How to run gemma-4-31B-it in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for gemma-4-31B-it in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every Gemma model you can run locally.

gemma-4-31B-it license

gemma-4-31B-it is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download google/gemma-4-31B-it
from transformers import AutoModel
model = AutoModel.from_pretrained("google/gemma-4-31B-it")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is an instruction-tuned open-weight model from Google's Gemma 4 family, built by Google DeepMind, with 32.7B parameters and a 262,144-token context window. The HuggingFace tags list it as an image-text-to-text model, so it handles both images and text, and it adds reasoning, code, and multilingual support. In Atomic Chat it runs entirely on your own hardware, offline and private.

A 4-bit (Q4) quantization needs roughly 18-19GB of VRAM, and a 24GB card such as an RTX 3090 or RTX 4090 is a comfortable target for daily use. An 8-bit (Q8) build is around 32.6GB. Keep in mind that long contexts add a large KV cache: near the full 262K window it can grow by about 22GB on top of the weights.

Yes. The model is released under the Apache 2.0 license, which allows free personal and commercial use with no licensing fee and no agreement with Google. You only need to include the license text when you redistribute it and note any modifications you make.

Yes. Once the weights are downloaded, gemma-4-31B-it runs fully on your machine with no network connection. Atomic Chat loads it locally, so prompts, documents, and images stay on your device and nothing is sent to a remote server.

It is strong at multimodal understanding, meaning reading documents and PDFs, OCR and handwriting, chart and screenshot analysis, alongside step-by-step reasoning, code, and structured tool use for agentic tasks. Its 140-plus language support and 262,144-token context make it useful for long, multilingual work that you want to keep on local hardware.