What is Gemma 4 12B?
Gemma 4 12B is an 11.95B dense multimodal model from Google DeepMind, the mid-size member of a five-model family that also covers E2B, E4B, 26B A4B and 31B. It is the family's "Unified" model: other Gemma 4 models use dedicated encoders to process multimodal data before passing it to the LLM, while the 12B projects raw image patches and audio waveforms straight into the decoder through lightweight linear layers. One checkpoint reads text, images, audio and video, which is what you want on a consumer machine. Google DeepMind describes the result as a deployment size suited to consumer devices and streamlined local execution.
| Specification | Gemma 4 12B |
|---|---|
| Total parameters | 11.95B |
| Architecture | Dense decoder-only, encoder-free multimodal |
| Layers | 48 |
| Attention | Local sliding-window layers (1,024 tokens) interleaved with global layers |
| Context window | 256K tokens |
| Vocabulary | 262K tokens |
| Modalities | Text, image, audio and video input; text output |
| Reasoning | Thinking mode, toggled from the system prompt |
| Function calling | Native tool-use support |
| Languages | 35+ out of the box, pre-trained on 140+ |
| Training data cutoff | January 2025 |
| Base model | google/gemma-4-12B |
| Sampling defaults | temperature 1.0, top_p 0.95, top_k 64 |
| License | Apache 2.0 |
Global layers share unified keys and values and use Proportional RoPE, which is how a 12B model affords a 256K context without the KV cache eating your memory. The hybrid attention interleaves local sliding-window attention with full global attention and always ends on a global layer. Google's stated payoff for dropping the encoders: lower multimodal latency, and fine-tuning the whole model in one pass. Thinking is switched on by putting the <|think|> token at the start of the system prompt and off by removing it; with thinking off the model still emits the thought tags around an empty block. Gemma 4 also adds native support for the system role. Images arrive at variable aspect ratios and resolutions with a visual token budget of 70, 140, 280, 560 or 1,120 tokens per image: low for classification, captioning and video frames, high for OCR, document parsing and small text. Audio clips run up to 30 seconds, video up to 60 seconds at one frame per second.
Gemma 4 12B benchmarks
Google's launch numbers compare the instruction-tuned 12B with its siblings Gemma 4 26B A4B and Gemma 4 31B, the on-device E4B, and last generation's Gemma 3 27B:
| Benchmark | Gemma 4 12B | Gemma 4 26B A4B | Gemma 4 31B | Gemma 4 E4B | Gemma 3 27B (no think) |
|---|---|---|---|---|---|
MMLU Pro Academic knowledge | 77.2 | 82.6 | 85.2 | 69.4 | 67.6 |
AIME 2026 Competition math | 77.5 | 88.3 | 89.2 | 42.5 | 20.8 |
LiveCodeBench v6 Competitive coding | 72.0 | 77.1 | 80.0 | 52.0 | 29.1 |
Codeforces ELO Contest rating | 1659 | 1718 | 2150 | 940 | 110 |
GPQA Diamond Expert science | 78.8 | 82.3 | 84.3 | 58.6 | 42.4 |
Tau2 Agentic tools | 69.0 | 68.2 | 76.9 | 42.2 | 16.2 |
MMMU Pro Multimodal reasoning | 69.1 | 73.8 | 76.9 | 52.6 | 49.7 |
The 31B takes every row, as the largest model should. The 12B's results are the surprising ones: it edges out the MoE 26B A4B on Tau2 tool use and beats Gemma 3 27B on all seven rows at under half the size, though Google measured the Gemma 3 column without thinking.
The model card also lists what the 12B is for beyond the leaderboard. On vision: object detection, document and PDF parsing, screen and UI understanding, chart comprehension, multilingual OCR, handwriting recognition and pointing, with text and images interleaved in any order. It scores 79.7% on MATH-Vision, 48.7% on MedXPertQA MM and 0.164 average edit distance on OmniDocBench 1.5, where lower is better. On audio: speech recognition and speech-to-translated-text, scored at 38.5 on CoVoST and 0.069 on FLEURS (lower is better), both excluding Chinese. Audio runs only on E2B, E4B and 12B. Long context is scored separately: 43.4% on MRCR v2 8-needle at 128K, against 13.5% for Gemma 3 27B.
Gemma 4 12B hardware requirements
The system requirement to check is memory. The builds and file sizes below come from the unsloth/gemma-4-12B-it-GGUF repository, which also carries mmproj files at 0.18 GB in F16 and 0.21 GB in F32, an MTP head at 0.47 GB in Q8_0, and a two-bit UD-IQ2_M at 4.21 GB below the smallest row in the table.
| Memory | Build to pick | File size |
|---|---|---|
| 5 GB | UD-Q2_K_XL | 4.66 GB |
| 6 GB | Q3_K_S | 5.14 GB |
| 8 GB | Q4_K_M | 7.12 GB |
| 10 GB | Q5_K_M | 8.41 GB |
| 12 GB | Q6_K | 9.79 GB |
| 16 GB | Q8_0 | 12.67 GB |
| 32 GB and up | BF16 | 23.83 GB |
When two builds both fit, take the larger one.
How to run Gemma 4 12B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Gemma 4 12B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Gemma model you can run locally.
Gemma 4 12B license
Gemma 4 12B is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties: fine-tune the weights, ship products on them, and run the model on your own hardware without a usage fee. Google DeepMind still asks developers to add content-safety safeguards matching their own product policies and use cases.
