What is Llama-3.2-3B-Instruct?
Llama-3.2-3B-Instruct is a 3.21B instruction-tuned text model from Meta, the larger of the two text-only models in the Llama 3.2 release of September 25, 2024. Meta expects the 1B and 3B to run in constrained environments such as mobile devices, and lists assistant-like chat, knowledge retrieval, summarization, mobile writing assistants and query rewriting as the intended jobs. That design goal is what makes it useful locally: the 4-bit GGUF file is about 2 GB, small enough for machines where bigger models do not fit.
| Specification | Llama-3.2-3B-Instruct |
|---|---|
| Parameters | 3.21B (3,212,749,824) in BF16 |
| Architecture | Auto-regressive optimized transformer, text in, text out |
| Attention | Grouped-Query Attention with shared embeddings |
| Context window | 128K tokens, 8K on Meta's own quantized builds |
| Input modalities | Multilingual text |
| Output modalities | Multilingual text and code |
| Supported languages | English, German, French, Italian, Portuguese, Hindi, Spanish, Thai |
| Training data | Public online data, up to 9T tokens |
| Knowledge cutoff | December 2023 |
| Alignment | Supervised fine-tuning and RLHF |
| Training compute | 460k GPU hours on H100-80GB |
| Release date | September 25, 2024 |
| License | Llama 3.2 Community License |
For the 1B and 3B, Meta fed logits from the Llama 3.1 8B and 70B models into pretraining as token-level targets, and says knowledge distillation was used after pruning to recover performance. Post-training follows a recipe similar to Llama 3.1: several rounds of supervised fine-tuning, rejection sampling and direct preference optimization. Every version in the family uses Grouped-Query Attention, which Meta credits for improved inference scalability, and both text models are listed with shared embeddings.
Llama-3.2-3B-Instruct benchmarks
Meta's published numbers, from the model card, put the instruction-tuned 3B next to its 1B sibling and Llama 3.1 8B:
| Benchmark | Llama-3.2-3B-Instruct | Llama 3.2 1B | Llama 3.1 8B |
|---|---|---|---|
MMLU Academic knowledge | 63.4 | 49.3 | 69.4 |
IFEval Instruction following | 77.4 | 59.5 | 80.4 |
GSM8K (CoT) Grade-school math | 77.7 | 44.4 | 84.5 |
MATH (CoT) Competition math | 48.0 | 30.6 | 51.9 |
ARC-C Science reasoning | 78.6 | 59.4 | 83.4 |
BFCL V2 Function calling | 67.0 | 25.7 | 67.1 |
TLDR9+ Text summarization | 19.0 | 16.8 | 17.2 |
Llama 3.1 8B leads almost every row, as it should at more than twice the parameters. The 3B stays within a point of it on function calling, 67.0 against 67.1 on BFCL V2, and beats it outright on TLDR9+ summarization. Meta also reports MMLU in the seven supported languages beyond English, where the 3B runs from 43.3 in Hindi to 55.1 in Spanish.
Meta claims these instruction-tuned text models "outperform many of the available open source and closed chat models" on common industry benchmarks, and the rows outside the table test that. Multilingual math holds up at 58.2 on MGSM chain-of-thought against 24.5 for the 1B, while Nexus tool use is thinner at 34.3 against 38.5 for the 8B. Long context is the weakest spot: 84.7 against 98.8 on NIH multi-needle recall, 19.8 against 27.3 on InfiniteBench En.QA, and 63.3 against 72.2 on En.MC. Treat the 128K window as room to hold a long document, not a promise of perfect recall inside it.
Llama-3.2-3B-Instruct hardware requirements
The system requirement to check is memory. The sizes below are real file sizes from the community repo unsloth/Llama-3.2-3B-Instruct-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 2 GB | UD-IQ2_M | 1.26 GB |
| 3 GB | Q3_K_M | 1.69 GB |
| 4 GB | Q4_K_M | 2.02 GB |
| 5 GB | Q5_K_M | 2.32 GB |
| 6 GB | Q6_K | 2.64 GB |
| 8 GB | Q8_0 | 3.42 GB |
| 16 GB and up | BF16 | 6.43 GB |
Even the full-precision BF16 file is only 6.43 GB, so when two builds both fit, take the larger one. The repo goes lower still, down to a 0.91 GB UD-IQ1_S. Meta's own figures show what its released quantized 3B builds give up: on BFCL V2, 60.1 for SpinQuant and 63.5 for QLoRA against 67.0 at BF16, while GSM8K barely moves, 75.7 and 77.9 against 77.7. Keep the IQ1 and IQ2 files as a last resort, not a default.
Meta names two ways to run the original weights: the Transformers pipeline from version 4.43.0 onward, and its own llama codebase. Its quantized builds target PyTorch ExecuTorch with an Arm CPU backend, where the BF16 3B decodes at 7.6 tokens per second on an Android OnePlus 12 and the SpinQuant build at 19.7. If GGUF is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits alongside it.
How to run Llama-3.2-3B-Instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Llama-3.2-3B-Instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Llama model you can run locally, or the bigger sibling Llama-3.1-8B-Instruct from the table above.
Llama-3.2-3B-Instruct license
Use of Llama-3.2-3B-Instruct is governed by the Llama 3.2 Community License, which Meta describes as a custom, commercial license agreement. Meta states the model is intended for commercial and research use in multiple languages, and that developers may fine-tune it for languages beyond the eight officially supported ones, provided they comply with the license and the Acceptable Use Policy. Out of scope, in Meta's wording, is any use that violates applicable laws or regulations, including trade compliance laws.
