Llama-3.2-3B-Instruct

Updated
05.10.2026
Tools
Reasoning
Code
Multilingual

Llama 3.2 3B Instruct is Meta’s 3.21B multilingual chat model with a 128K context. The 4-bit GGUF build is a 2 GB file.

At a glance

  • License: Llama 3.2 Community License
  • Parameters: 3.21B
  • Context length: 128K tokens
  • Modalities: Multilingual text in, text and code out
  • Minimum hardware: 4 GB memory (2.02 GB Q4_K_M build)

What is Llama-3.2-3B-Instruct?

Llama-3.2-3B-Instruct is a 3.21B instruction-tuned text model from Meta, the larger of the two text-only models in the Llama 3.2 release of September 25, 2024. Meta expects the 1B and 3B to run in constrained environments such as mobile devices, and lists assistant-like chat, knowledge retrieval, summarization, mobile writing assistants and query rewriting as the intended jobs. That design goal is what makes it useful locally: the 4-bit GGUF file is about 2 GB, small enough for machines where bigger models do not fit.

SpecificationLlama-3.2-3B-Instruct
Parameters3.21B (3,212,749,824) in BF16
ArchitectureAuto-regressive optimized transformer, text in, text out
AttentionGrouped-Query Attention with shared embeddings
Context window128K tokens, 8K on Meta's own quantized builds
Input modalitiesMultilingual text
Output modalitiesMultilingual text and code
Supported languagesEnglish, German, French, Italian, Portuguese, Hindi, Spanish, Thai
Training dataPublic online data, up to 9T tokens
Knowledge cutoffDecember 2023
AlignmentSupervised fine-tuning and RLHF
Training compute460k GPU hours on H100-80GB
Release dateSeptember 25, 2024
LicenseLlama 3.2 Community License

For the 1B and 3B, Meta fed logits from the Llama 3.1 8B and 70B models into pretraining as token-level targets, and says knowledge distillation was used after pruning to recover performance. Post-training follows a recipe similar to Llama 3.1: several rounds of supervised fine-tuning, rejection sampling and direct preference optimization. Every version in the family uses Grouped-Query Attention, which Meta credits for improved inference scalability, and both text models are listed with shared embeddings.

Llama-3.2-3B-Instruct benchmarks

Meta's published numbers, from the model card, put the instruction-tuned 3B next to its 1B sibling and Llama 3.1 8B:

BenchmarkLlama-3.2-3B-InstructLlama 3.2 1BLlama 3.1 8B
MMLU
Academic knowledge
63.449.369.4
IFEval
Instruction following
77.459.580.4
GSM8K (CoT)
Grade-school math
77.744.484.5
MATH (CoT)
Competition math
48.030.651.9
ARC-C
Science reasoning
78.659.483.4
BFCL V2
Function calling
67.025.767.1
TLDR9+
Text summarization
19.016.817.2

Llama 3.1 8B leads almost every row, as it should at more than twice the parameters. The 3B stays within a point of it on function calling, 67.0 against 67.1 on BFCL V2, and beats it outright on TLDR9+ summarization. Meta also reports MMLU in the seven supported languages beyond English, where the 3B runs from 43.3 in Hindi to 55.1 in Spanish.

Meta claims these instruction-tuned text models "outperform many of the available open source and closed chat models" on common industry benchmarks, and the rows outside the table test that. Multilingual math holds up at 58.2 on MGSM chain-of-thought against 24.5 for the 1B, while Nexus tool use is thinner at 34.3 against 38.5 for the 8B. Long context is the weakest spot: 84.7 against 98.8 on NIH multi-needle recall, 19.8 against 27.3 on InfiniteBench En.QA, and 63.3 against 72.2 on En.MC. Treat the 128K window as room to hold a long document, not a promise of perfect recall inside it.

Llama-3.2-3B-Instruct hardware requirements

The system requirement to check is memory. The sizes below are real file sizes from the community repo unsloth/Llama-3.2-3B-Instruct-GGUF.

MemoryBuild to pickFile size
2 GBUD-IQ2_M1.26 GB
3 GBQ3_K_M1.69 GB
4 GBQ4_K_M2.02 GB
5 GBQ5_K_M2.32 GB
6 GBQ6_K2.64 GB
8 GBQ8_03.42 GB
16 GB and upBF166.43 GB

Even the full-precision BF16 file is only 6.43 GB, so when two builds both fit, take the larger one. The repo goes lower still, down to a 0.91 GB UD-IQ1_S. Meta's own figures show what its released quantized 3B builds give up: on BFCL V2, 60.1 for SpinQuant and 63.5 for QLoRA against 67.0 at BF16, while GSM8K barely moves, 75.7 and 77.9 against 77.7. Keep the IQ1 and IQ2 files as a last resort, not a default.

Meta names two ways to run the original weights: the Transformers pipeline from version 4.43.0 onward, and its own llama codebase. Its quantized builds target PyTorch ExecuTorch with an Arm CPU backend, where the BF16 3B decodes at 7.6 tokens per second on an Android OnePlus 12 and the SpinQuant build at 19.7. If GGUF is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits alongside it.

How to run Llama-3.2-3B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Llama-3.2-3B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Llama model you can run locally, or the bigger sibling Llama-3.1-8B-Instruct from the table above.

Llama-3.2-3B-Instruct license

Use of Llama-3.2-3B-Instruct is governed by the Llama 3.2 Community License, which Meta describes as a custom, commercial license agreement. Meta states the model is intended for commercial and research use in multiple languages, and that developers may fine-tune it for languages beyond the eight officially supported ones, provided they comply with the license and the Acceptable Use Policy. Out of scope, in Meta's wording, is any use that violates applicable laws or regulations, including trade compliance laws.

Get the weights from Hugging Face

huggingface-cli download meta-llama/Llama-3.2-3B-Instruct
from transformers import AutoModel
model = AutoModel.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Llama-3.2-3B-Instruct is a 3.2-billion-parameter instruction-tuned language model from Meta, built on a dense transformer architecture with a 128K-token context window. Meta tuned it for following instructions, multilingual dialogue, summarization, and tool use. Its small size makes it a good fit for running on a laptop or single consumer GPU through Atomic Chat.

In full FP16 precision the model needs about 7 GB of VRAM. A 4-bit quantized build cuts that to roughly 1.8-2 GB, so it runs on a 6 GB GPU and even on modern laptops without a dedicated GPU. Using the full 128K context window adds memory on top of those figures.

Yes. The weights are released under the Llama 3.2 Community License, which allows free commercial and research use. Running it locally in Atomic Chat means no API fees or usage limits. Products with more than 700 million monthly active users need a separate license from Meta.

Yes. Once you download the weights, the model runs fully on your own machine with no internet connection. Prompts and files never leave the device, which suits privacy-sensitive work. Atomic Chat loads the model on-device so chats stay local and offline.

It works well for chat assistants, summarizing long documents, and answering questions grounded in text you provide, all within its 128K context. It also supports tool calling and multilingual conversation, so you can connect it to local scripts or use it across several languages. For heavier reasoning or large codebases, a bigger model will perform better.