What is gemma-4-E2B-it?
gemma-4-E2B-it is the smallest instruction-tuned model in Google DeepMind's Gemma 4 family, published on Hugging Face in March 2026 under Apache 2.0. The E stands for effective parameters: the model computes with 2.3B parameters but stores 5.1B in total, because Per-Layer Embeddings give every decoder layer its own small per-token embedding table. Those tables are big but only used for quick lookups, which is how the model keeps phone-level compute cost. It is also one of the three Gemma 4 models with native audio input, so it can transcribe and translate speech fully offline.
| Specification | gemma-4-E2B-it |
|---|---|
| Effective parameters | 2.3B (5.1B total with embeddings) |
| Architecture | Dense, hybrid attention: local sliding window plus full global layers |
| Layers | 35 |
| Sliding window | 512 tokens |
| Context window | 128K tokens |
| Vocabulary | 262K |
| Modalities | Text, image and audio input, text output |
| Encoders | ~150M vision, ~300M audio |
| Reasoning | Thinking mode, enabled via a system prompt token |
| Release date | March 2, 2026 |
| License | Apache 2.0 |
The attention stack interleaves local sliding window layers, 512 tokens wide on the E2B, with full global layers, and the final layer is always global. Global layers use unified Keys and Values plus Proportional RoPE, which keeps long-context memory in check. On the multimodal side, images come in at variable aspect ratios and resolutions with a configurable visual token budget from 70 to 1120 tokens, audio clips run up to 30 seconds, and video is processed as frames, up to 60 seconds at one frame per second. Google states out-of-the-box support for 35+ languages, with pretraining across more than 140.
gemma-4-E2B-it benchmarks
Google DeepMind published these numbers on the Gemma 4 model card; the columns next to the E2B are its bigger sibling Gemma 4 E4B and the previous generation's Gemma 3 27B, run without thinking:
| Benchmark | gemma-4-E2B-it | Gemma 4 E4B | Gemma 3 27B (no think) |
|---|---|---|---|
MMLU Pro Academic knowledge | 60.0% | 69.4% | 67.6% |
AIME 2026 Competition math | 37.5% | 42.5% | 20.8% |
LiveCodeBench v6 Competitive coding | 44.0% | 52.0% | 29.1% |
GPQA Diamond Expert science | 43.4% | 58.6% | 42.4% |
MMMU Pro Multimodal understanding | 44.2% | 52.6% | 49.7% |
MMMLU Multilingual knowledge | 67.4% | 76.6% | 70.7% |
CoVoST Speech translation | 33.47 | 35.54 | - |
The E4B takes every row, so if your machine has the memory for it, run the E4B instead. The E2B's case is against the previous generation: it beats Gemma 3 27B on competition math, coding and expert science at a fraction of the size, and it is the smallest Gemma 4 that accepts audio.
gemma-4-E2B-it hardware requirements
The system requirement to check is memory. The builds below come from unsloth/gemma-4-E2B-it-GGUF, with the real file sizes from the repo listing.
| Memory | Build to pick | File size |
|---|---|---|
| 4 GB | UD-Q2_K_XL | 2.40 GB |
| 6 GB | Q4_K_M | 3.11 GB |
| 8 GB | Q6_K | 4.50 GB |
| 12 GB | Q8_0 | 5.05 GB |
| 16 GB and up | BF16 | 9.31 GB |
Neighbouring builds differ by half a gigabyte or so, so when two builds both fit, take the larger one. For image and audio input, add the mmproj file from the same repo, an extra 0.99 GB next to the main build. If the format is new to you, start with what GGUF is.
How to run gemma-4-E2B-it in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for gemma-4-E2B-it in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Gemma model you can run locally, or step up to the fast MoE sibling Gemma 4 26B A4B.
gemma-4-E2B-it license
gemma-4-E2B-it is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can ship it inside your own products and run it on your own hardware without a usage fee.
