gemma-3-270m

Updated
24.08.2026
Reasoning
Multilingual

Gemma 3 270M is the smallest Gemma 3 model: 268M parameters, 32K context, trained on 6 trillion tokens. The Q8_0 GGUF is a 0.29 GB file.

At a glance

  • License: Gemma
  • Parameters: 268M (0.27B)
  • Context length: 32K tokens
  • Modalities: Text in, text out
  • Minimum hardware: the F16 GGUF is 0.54 GB, the smallest build 0.18 GB

What is Gemma 3 270M?

Gemma 3 270M is the smallest model in Google's Gemma 3 family: a text-only model with 268M parameters from Google DeepMind, built from the same research and technology behind the Gemini models. The unusual part is the training budget. Google trained it on 6 trillion tokens, three times the 2 trillion the 1B sibling got and more than the 4B's 4 trillion. That much training behind so few parameters is why it matters locally: it runs on effectively any machine you own.

SpecificationGemma 3 270M
Total parameters268M (0.27B)
ArchitectureGemma 3, text-only variant
Context window32K tokens
ModalitiesText in, text out
Training data6 trillion tokens, knowledge cutoff August 2024
LanguagesTrained on content in over 140 languages
Release dateAugust 5, 2025
LicenseGemma

The training mix is web documents, code and mathematics, with content in over 140 languages and a knowledge cutoff of August 2024. Unlike the 4B, 12B and 27B members of the family, the 270M takes no image input, and its context window is 32K tokens rather than the larger models' 128K. Google pitches the family at environments with limited resources, laptops, desktops or your own cloud infrastructure, and the 270M is the far end of that scale: the unquantized F16 GGUF of the instruction-tuned model is a 0.54 GB file.

Gemma 3 270M benchmarks

Google's published numbers, from the Gemma 3 model card, cover the pre-trained (PT) and instruction-tuned (IT) checkpoints of the 270M. The vendor table compares it against no other models, so read these as a capability floor for the size, not a leaderboard:

BenchmarkGemma 3 270M PTGemma 3 270M IT
HellaSwag
Commonsense completion
40.937.7
BoolQ
Binary questions
61.4-
PIQA
Physical commonsense
67.766.2
TriviaQA
Trivia recall
15.4-
ARC-c
Science questions
29.028.2
WinoGrande
Commonsense reasoning
52.052.3
BIG-Bench Hard
Hard reasoning
-26.7
IF Eval
Instruction following
-51.2

The scores land where a 268M-parameter model should: reasonable physical commonsense (PIQA 67.7), weak hard reasoning (BIG-Bench Hard 26.7), and almost no stored world knowledge (TriviaQA 15.4). The useful number is IF Eval at 51.2: the instruction-tuned checkpoint holds up on instruction following, which matches Google's own guidance that these models do best on tasks framed with clear prompts and instructions rather than open-ended conversation.

Gemma 3 270M hardware requirements

The system requirement to check is memory, and here it barely registers. The builds come from community repos: ggml-org/gemma-3-270m-GGUF carries a single Q8_0 and unsloth/gemma-3-270m-it-GGUF carries the full ladder. The sizes below are the unsloth files of the instruction-tuned model.

MemoryBuild to pickFile size
2 GBQ4_K_M0.25 GB
4 GBQ8_00.29 GB
8 GB and upF160.54 GB

The entire ladder spans 0.18 GB to 0.54 GB, so a low-bit quant saves a couple hundred megabytes at most: take Q8_0 or the F16 file and keep the quality. If the format is new to you, start with what GGUF is.

How to run Gemma 3 270M in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Gemma 3 270M in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Gemma model you can run locally, or the next size up, Gemma 3 1B.

Gemma 3 270M license

Gemma 3 270M ships under Google's Gemma license, not Apache 2.0. The weights are open for both the pre-trained and instruction-tuned variants, with use governed by Google's Terms of Use and the Gemma Prohibited Use Policy. One practical note: the Hugging Face repo is gated, so you log in and accept Google's usage license, and access is granted immediately.

Get the weights from Hugging Face

huggingface-cli download google/gemma-3-270m
from transformers import AutoModel
model = AutoModel.from_pretrained("google/gemma-3-270m")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

gemma-3-270m is the smallest model in Google's Gemma 3 family, a text-only transformer with 0.27B parameters and a 32K context window. It is a fine-tune base built for narrow, efficient tasks like extraction, classification, and routing rather than broad chat. You can run it fully on-device through Atomic Chat.

Very little. Quantized to INT4, gemma-3-270m uses under 200 MB of memory and runs without a dedicated GPU, so a standard laptop, desktop, or even a phone can handle it. For local fine-tuning rather than inference, you'd want an NVIDIA GPU with around 8 GB of VRAM, but plain inference in Atomic Chat works on modest CPUs.

Yes. The weights are open and free to download from Hugging Face under the Gemma license, which permits responsible commercial use including fine-tuning and deploying your own version. Running it locally in Atomic Chat costs nothing beyond your own hardware and electricity.

Yes. Once you download the weights, gemma-3-270m runs entirely on your machine with no network connection. In Atomic Chat your prompts and outputs stay on-device, which keeps the model usable on a plane, in the field, or anywhere private data shouldn't leave your computer.

It is strongest on focused, high-volume tasks: entity extraction, sentiment analysis, query routing, and turning unstructured text into structured output. As a fine-tune base it is cheap to specialize for one job and then run that version locally. For long open-ended conversation or deep reasoning, a larger model is a better fit.