gemma-4-E2B-it

Updated
05.10.2026
Thinking
Embedding
Vision
Audio
Reasoning
Code
Multilingual

Gemma 4 E2B is Google DeepMind’s smallest Gemma 4 model: 2.3B effective parameters, text, image and audio input, and a 128K context window.

At a glance

  • License: Apache 2.0
  • Parameters: 2.3B effective, 5.1B total
  • Context length: 128K tokens
  • Modalities: Text, image and audio input
  • Minimum hardware: 4 GB of RAM

What is gemma-4-E2B-it?

gemma-4-E2B-it is the smallest instruction-tuned model in Google DeepMind's Gemma 4 family, published on Hugging Face in March 2026 under Apache 2.0. The E stands for effective parameters: the model computes with 2.3B parameters but stores 5.1B in total, because Per-Layer Embeddings give every decoder layer its own small per-token embedding table. Those tables are big but only used for quick lookups, which is how the model keeps phone-level compute cost. It is also one of the three Gemma 4 models with native audio input, so it can transcribe and translate speech fully offline.

Specificationgemma-4-E2B-it
Effective parameters2.3B (5.1B total with embeddings)
ArchitectureDense, hybrid attention: local sliding window plus full global layers
Layers35
Sliding window512 tokens
Context window128K tokens
Vocabulary262K
ModalitiesText, image and audio input, text output
Encoders~150M vision, ~300M audio
ReasoningThinking mode, enabled via a system prompt token
Release dateMarch 2, 2026
LicenseApache 2.0

The attention stack interleaves local sliding window layers, 512 tokens wide on the E2B, with full global layers, and the final layer is always global. Global layers use unified Keys and Values plus Proportional RoPE, which keeps long-context memory in check. On the multimodal side, images come in at variable aspect ratios and resolutions with a configurable visual token budget from 70 to 1120 tokens, audio clips run up to 30 seconds, and video is processed as frames, up to 60 seconds at one frame per second. Google states out-of-the-box support for 35+ languages, with pretraining across more than 140.

gemma-4-E2B-it benchmarks

Google DeepMind published these numbers on the Gemma 4 model card; the columns next to the E2B are its bigger sibling Gemma 4 E4B and the previous generation's Gemma 3 27B, run without thinking:

Benchmarkgemma-4-E2B-itGemma 4 E4BGemma 3 27B (no think)
MMLU Pro
Academic knowledge
60.0%69.4%67.6%
AIME 2026
Competition math
37.5%42.5%20.8%
LiveCodeBench v6
Competitive coding
44.0%52.0%29.1%
GPQA Diamond
Expert science
43.4%58.6%42.4%
MMMU Pro
Multimodal understanding
44.2%52.6%49.7%
MMMLU
Multilingual knowledge
67.4%76.6%70.7%
CoVoST
Speech translation
33.4735.54-

The E4B takes every row, so if your machine has the memory for it, run the E4B instead. The E2B's case is against the previous generation: it beats Gemma 3 27B on competition math, coding and expert science at a fraction of the size, and it is the smallest Gemma 4 that accepts audio.

gemma-4-E2B-it hardware requirements

The system requirement to check is memory. The builds below come from unsloth/gemma-4-E2B-it-GGUF, with the real file sizes from the repo listing.

MemoryBuild to pickFile size
4 GBUD-Q2_K_XL2.40 GB
6 GBQ4_K_M3.11 GB
8 GBQ6_K4.50 GB
12 GBQ8_05.05 GB
16 GB and upBF169.31 GB

Neighbouring builds differ by half a gigabyte or so, so when two builds both fit, take the larger one. For image and audio input, add the mmproj file from the same repo, an extra 0.99 GB next to the main build. If the format is new to you, start with what GGUF is.

How to run gemma-4-E2B-it in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for gemma-4-E2B-it in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Gemma model you can run locally, or step up to the fast MoE sibling Gemma 4 26B A4B.

gemma-4-E2B-it license

gemma-4-E2B-it is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can ship it inside your own products and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download google/gemma-4-E2B-it
from transformers import AutoModel
model = AutoModel.from_pretrained("google/gemma-4-E2B-it")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is an instruction-tuned model from Google in the Gemma 4 family. The E2B variant is a compact multimodal model with about 2.3B effective parameters (5.1B including embeddings) and a 125K-token context window. It handles text, images, and audio, and is built to run on edge and consumer hardware.

A 4-bit quantized build of gemma-4-E2B-it uses around 5GB of RAM, and Google lists a minimum of roughly 4GB of free RAM to load it. It runs on a CPU, so a dedicated GPU is not required, though a GPU speeds up generation. The model was engineered for offline mobile and IoT devices, including boards like the Jetson Orin Nano.

Yes. The weights are released under the apache-2.0 license and are free to download and run locally, with no subscription or per-token charge. The license also allows commercial use, modification, and redistribution.

Yes. Once you download the weights, gemma-4-E2B-it runs fully offline with no network connection. Running it through Atomic Chat keeps every prompt and file on your own device, so your data is never sent to a server.

It is a good fit for private, on-device tasks: reading images and audio, step-by-step reasoning through its built-in thinking mode, writing and explaining code, and multilingual chat. Its small footprint makes it practical for laptops, phones, and small edge devices where larger models will not fit.