Llama-3.1-8B-Instruct

Updated
05.10.2026
Tools
Reasoning
Code
Multilingual

Llama-3.1-8B-Instruct is Meta’s 8B multilingual chat model with a 128K context and tool calling. The Q4_K_M build is 4.92 GB.

At a glance

  • License: Llama 3.1 Community License
  • Parameters: 8B
  • Context length: 128K tokens
  • Modalities: Multilingual text in, text and code out
  • Minimum hardware: 4 GB of memory (UD-Q2_K_XL build, 3.39 GB)

What is Llama-3.1-8B-Instruct?

Llama-3.1-8B-Instruct is the smallest of the three models in Meta's Llama 3.1 collection (8B, 70B and 405B): an 8 billion parameter, instruction tuned, text only model optimized for multilingual dialogue. Meta aligned it with supervised fine-tuning and reinforcement learning with human feedback and released it on July 23, 2024. The size is what makes it a local staple: the Q4_K_M build is a 4.92 GB file, so the model fits an 8 GB GPU with room left for context.

SpecificationLlama-3.1-8B-Instruct
Parameters8B (8,030,261,248 in the BF16 weights)
ArchitectureAuto-regressive, optimized transformer
AttentionGrouped-Query Attention (GQA)
Context window128K tokens
ModalitiesMultilingual text in, multilingual text and code out
Supported languagesEnglish, German, French, Italian, Portuguese, Hindi, Spanish and Thai
Pretraining dataPublicly available online data, 15T+ tokens
Knowledge cutoffDecember 2023
AlignmentSupervised fine-tuning and RLHF
Fine-tuning dataPublic instruction datasets plus over 25M synthetic examples
Training compute1.46M GPU hours on H100-80GB
Release dateJuly 23, 2024
LicenseLlama 3.1 Community License

Meta ships the repository in two versions, one for Transformers 4.43.0 and later and one for the original llama codebase, and points at the huggingface-llama-recipes collection for local runs with torch.compile(), assisted generation and quantised weights. Llama 3.1 also supports multiple tool use formats, and tool calling is wired into the Transformers chat template: you pass Python functions to apply_chat_template and append each result back with the tool role. That is enough to drive a small agent loop on your own machine. Meta calls this a static model trained on an offline dataset.

Llama-3.1-8B-Instruct benchmarks

Meta's published numbers, from the model card, compare the instruction tuned 8B with its predecessor Llama 3 8B Instruct and the two larger Llama 3.1 models:

BenchmarkLlama 3.1 8B InstructLlama 3 8B InstructLlama 3.1 70B InstructLlama 3.1 405B Instruct
MMLU
General knowledge
69.468.583.687.3
IFEval
Instruction following
80.476.887.588.6
GPQA
Expert science
30.434.646.750.7
HumanEval
Code generation
72.660.480.589.0
MATH (CoT)
Harder math
51.929.168.073.8
API-Bank
Tool APIs
82.648.390.092.0
BFCL
Function calling
76.160.384.888.5
Multilingual MGSM
Multilingual math
68.9-86.991.6

Inside the family the 405B takes every row, as you would expect. The column worth reading is Llama 3 8B Instruct: at the same size, the 3.1 update gains 34.3 points on API-Bank, 22.8 on MATH, 15.8 on BFCL and 12.2 on HumanEval, and gives back 4.2 points on GPQA.

Meta's own framing is narrow. The instruction tuned text only models are built for assistant-like chat in the eight supported languages, and Meta claims they "outperform many of the available open source and closed chat models" on common industry benchmarks. The multilingual half of that claim comes with its own table: on per-language MMLU the 8B scores 62.45 in Spanish, 62.34 in French, 62.12 in Portuguese, 61.63 in Italian and 60.59 in German, while Hindi lands at 50.88 and Thai at 50.32, about ten points below the European languages. Meta notes the model saw more languages in pretraining than the eight it supports, and discourages conversing in the others without fine-tuning and system controls of your own.

Llama-3.1-8B-Instruct hardware requirements

The system requirement to check is memory. The builds below, with their real file sizes, come from unsloth/Llama-3.1-8B-Instruct-GGUF.

MemoryBuild to pickFile size
4 GBUD-Q2_K_XL3.39 GB
6 GBQ3_K_M4.02 GB
8 GBQ4_K_M4.92 GB
10 GBQ5_K_M5.73 GB
12 GBQ6_K6.60 GB
16 GBUD-Q8_K_XL10.58 GB
24 GB and upBF1616.07 GB

Neighbouring builds are often only a few hundred megabytes apart, so when two both fit, take the larger one. That matters most at the bottom of the table, where quality falls fastest: the repo goes down to 2.16 GB with UD-IQ1_S, but files that small are a last resort, not a default. Leave some headroom if you plan to use long inputs, the 128K window needs memory of its own. If the format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits alongside it.

How to run Llama-3.1-8B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Llama-3.1-8B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Llama model you can run locally, or the smaller sibling Llama-3.2-3B-Instruct.

Llama-3.1-8B-Instruct license

Llama-3.1-8B-Instruct ships under the Llama 3.1 Community License, a custom commercial license from Meta. It permits commercial and research use in the supported languages, and it explicitly allows using the model's outputs to improve other models, including synthetic data generation and distillation. Use has to stay inside applicable law, Meta's Acceptable Use Policy and the eight supported languages, so read the license text before shipping a product on top of it.

Get the weights from Hugging Face

huggingface-cli download meta-llama/Llama-3.1-8B-Instruct
from transformers import AutoModel
model = AutoModel.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Llama-3.1-8B-Instruct is an 8-billion-parameter open-weight language model from Meta. It is the instruction-tuned variant of Llama 3.1 8B, refined with supervised fine-tuning and RLHF so it follows prompts and chats naturally. It handles chat, writing, summarizing, and coding, and supports a 128K-token context window.

A 4-bit quantized build (Q4_K_M) is about 5 GB and runs on a GPU with roughly 6 GB of VRAM, with an RTX 3060 12GB being a comfortable target. Full 16-bit weights need closer to 16 GB. With no dedicated GPU you can still run it on CPU with 16 GB or more of system RAM at a few tokens per second.

Yes. The weights are open and free to download under Meta's llama3.1 community license, which allows commercial and research use. In Atomic Chat there is no per-token cost because the model runs on your own machine. You only pay for the hardware and electricity you already own.

Yes. After the weights download once, the model runs entirely on-device with no internet connection needed. Atomic Chat loads it locally, so prompts and responses stay on your machine. This makes it usable on a plane, in an air-gapped setup, or anywhere you want to keep data private.

It is a strong general-purpose model for its size, good at conversation, summarizing, content generation, and coding help. Its capability tags cover tools, reasoning, code, and multilingual text, and the 128K context lets it work over long documents in one pass. The 8B size keeps it fast enough for real-time use on a laptop or a mid-range GPU.