gemma-4-12B-it

Updated
05.10.2026
Thinking
Embedding
Vision
Audio
Reasoning
Code
Multilingual

Gemma 4 12B is Google DeepMind’s encoder-free multimodal model: text, image, audio and video input, a 256K context window, Apache 2.0.

At a glance

  • License: Apache 2.0
  • Parameters: 11.95B, dense
  • Context length: 256K tokens
  • Modalities: Text, image, audio and video input
  • Minimum hardware: 5 GB of memory (UD-Q2_K_XL build, 4.66 GB)

What is Gemma 4 12B?

Gemma 4 12B is an 11.95B dense multimodal model from Google DeepMind, the mid-size member of a five-model family that also covers E2B, E4B, 26B A4B and 31B. It is the family's "Unified" model: other Gemma 4 models use dedicated encoders to process multimodal data before passing it to the LLM, while the 12B projects raw image patches and audio waveforms straight into the decoder through lightweight linear layers. One checkpoint reads text, images, audio and video, which is what you want on a consumer machine. Google DeepMind describes the result as a deployment size suited to consumer devices and streamlined local execution.

SpecificationGemma 4 12B
Total parameters11.95B
ArchitectureDense decoder-only, encoder-free multimodal
Layers48
AttentionLocal sliding-window layers (1,024 tokens) interleaved with global layers
Context window256K tokens
Vocabulary262K tokens
ModalitiesText, image, audio and video input; text output
ReasoningThinking mode, toggled from the system prompt
Function callingNative tool-use support
Languages35+ out of the box, pre-trained on 140+
Training data cutoffJanuary 2025
Base modelgoogle/gemma-4-12B
Sampling defaultstemperature 1.0, top_p 0.95, top_k 64
LicenseApache 2.0

Global layers share unified keys and values and use Proportional RoPE, which is how a 12B model affords a 256K context without the KV cache eating your memory. The hybrid attention interleaves local sliding-window attention with full global attention and always ends on a global layer. Google's stated payoff for dropping the encoders: lower multimodal latency, and fine-tuning the whole model in one pass. Thinking is switched on by putting the <|think|> token at the start of the system prompt and off by removing it; with thinking off the model still emits the thought tags around an empty block. Gemma 4 also adds native support for the system role. Images arrive at variable aspect ratios and resolutions with a visual token budget of 70, 140, 280, 560 or 1,120 tokens per image: low for classification, captioning and video frames, high for OCR, document parsing and small text. Audio clips run up to 30 seconds, video up to 60 seconds at one frame per second.

Gemma 4 12B benchmarks

Google's launch numbers compare the instruction-tuned 12B with its siblings Gemma 4 26B A4B and Gemma 4 31B, the on-device E4B, and last generation's Gemma 3 27B:

BenchmarkGemma 4 12BGemma 4 26B A4BGemma 4 31BGemma 4 E4BGemma 3 27B (no think)
MMLU Pro
Academic knowledge
77.282.685.269.467.6
AIME 2026
Competition math
77.588.389.242.520.8
LiveCodeBench v6
Competitive coding
72.077.180.052.029.1
Codeforces ELO
Contest rating
165917182150940110
GPQA Diamond
Expert science
78.882.384.358.642.4
Tau2
Agentic tools
69.068.276.942.216.2
MMMU Pro
Multimodal reasoning
69.173.876.952.649.7

The 31B takes every row, as the largest model should. The 12B's results are the surprising ones: it edges out the MoE 26B A4B on Tau2 tool use and beats Gemma 3 27B on all seven rows at under half the size, though Google measured the Gemma 3 column without thinking.

The model card also lists what the 12B is for beyond the leaderboard. On vision: object detection, document and PDF parsing, screen and UI understanding, chart comprehension, multilingual OCR, handwriting recognition and pointing, with text and images interleaved in any order. It scores 79.7% on MATH-Vision, 48.7% on MedXPertQA MM and 0.164 average edit distance on OmniDocBench 1.5, where lower is better. On audio: speech recognition and speech-to-translated-text, scored at 38.5 on CoVoST and 0.069 on FLEURS (lower is better), both excluding Chinese. Audio runs only on E2B, E4B and 12B. Long context is scored separately: 43.4% on MRCR v2 8-needle at 128K, against 13.5% for Gemma 3 27B.

Gemma 4 12B hardware requirements

The system requirement to check is memory. The builds and file sizes below come from the unsloth/gemma-4-12B-it-GGUF repository, which also carries mmproj files at 0.18 GB in F16 and 0.21 GB in F32, an MTP head at 0.47 GB in Q8_0, and a two-bit UD-IQ2_M at 4.21 GB below the smallest row in the table.

MemoryBuild to pickFile size
5 GBUD-Q2_K_XL4.66 GB
6 GBQ3_K_S5.14 GB
8 GBQ4_K_M7.12 GB
10 GBQ5_K_M8.41 GB
12 GBQ6_K9.79 GB
16 GBQ8_012.67 GB
32 GB and upBF1623.83 GB

When two builds both fit, take the larger one.

How to run Gemma 4 12B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Gemma 4 12B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Gemma model you can run locally.

Gemma 4 12B license

Gemma 4 12B is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties: fine-tune the weights, ship products on them, and run the model on your own hardware without a usage fee. Google DeepMind still asks developers to add content-safety safeguards matching their own product policies and use cases.

Get the weights from Hugging Face

huggingface-cli download google/gemma-4-12B-it
from transformers import AutoModel
model = AutoModel.from_pretrained("google/gemma-4-12B-it")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

gemma-4-12B-it is a 12B-parameter, instruction-tuned model in Google's Gemma 4 family. It uses an encoder-free unified architecture that takes text, images, and audio in the same prompt, and it ships with a 128,000-token context window. The "it" marks it as the instruction-tuned build, fine-tuned on top of the Gemma 4 12B base for chat and task following.

A 4-bit quantized build weighs about 6.7 GB, so it fits on an 8 GB GPU for short prompts. For comfortable use with longer context, aim for 16 GB of VRAM or unified memory, because the KV cache grows as the prompt gets longer. Apple M-series Macs with 16-32 GB of unified memory are a good match since the full memory pool is usable.

Yes. It is released under the apache-2.0 license, so the weights are free to download, run, modify, and use commercially. The only cost is the hardware or cloud you run it on, and in Atomic Chat it runs locally for free.

Yes. Once the weights are downloaded, the model runs entirely on your own CPU or GPU with no internet connection. In Atomic Chat every prompt and response stays on your device, which suits private, confidential, or air-gapped work.

It handles general Q&A, code generation and debugging, summarization, and image or audio analysis from a single local model. The built-in thinking mode helps with step-by-step reasoning and math, and support for 140+ languages covers translation and multilingual writing. For most everyday tasks it is fast, private, and free to run locally.