Mistral-7B-Instruct-v0.2

Updated
02.10.2026
Reasoning
Code

Mistral-7B-Instruct-v0.2 is a 7B instruction-tuned model with a 32K context window. Explore its hardware needs, GGUF builds and local setup.

At a glance

  • License: Apache 2.0
  • Parameters: 7.24B
  • Context length: 32k tokens
  • Modalities: Text
  • Minimum hardware: 4 GB of memory (Q3_K_S GGUF, 3.16 GB file)

What is Mistral-7B-Instruct-v0.2?

Mistral-7B-Instruct-v0.2 is a 7.24B parameter language model from Mistral AI, an instruct fine-tune of the Mistral-7B-v0.2 base trained to follow chat instructions. Mistral published the weights on Hugging Face on December 11, 2023 under Apache 2.0, and the repo has passed 1.1 million downloads there. For local use the appeal is size: the whole model fits in a 4.37 GB file at 4-bit, so it runs comfortably on an ordinary laptop.

SpecificationMistral-7B-Instruct-v0.2
Total parameters7.24B
ArchitectureTransformer, no sliding-window attention
Context window32k tokens (8k in v0.1)
Rope theta1e6
ModalitiesText
Prompt format[INST] ... [/INST], shipped as a chat template
Release dateDecember 11, 2023
LicenseApache 2.0
SuccessorMistral-7B-Instruct-v0.3

The v0.2 base made three changes against the original Mistral 7B: the context window grew from 8k to 32k tokens, rope-theta moved to 1e6, and sliding-window attention was removed. The instruct tune expects prompts wrapped in [INST] and [/INST] tokens, with a begin-of-sentence token before the first instruction only. You rarely type that by hand: the format ships as a chat template, so the tokenizer builds it for you. Mistral treats its own mistral_common tokenizer as the reference implementation and notes on the model card that the transformers tokenizer does not yet match it one to one.

What Mistral-7B-Instruct-v0.2 is good at

Mistral publishes no benchmark numbers on the v0.2 model card, so the honest summary comes from what the card does say. The model is an instruction follower first: general chat, multi-turn conversation in the [INST] format, and the everyday tasks a tuned 7B gets asked to do. Mistral itself frames the release modestly, calling it a quick demonstration that the base model can be easily fine-tuned to compelling performance. The 32k context window, four times the 8k of v0.1, is the practical headline: it takes long documents and long conversations in one pass.

Two limitations come from the same card. The model has no moderation mechanisms, so anything that needs filtered output has to add its own guardrails on top. And it is a 2023 model with a newer successor: Mistral points to Mistral-7B-Instruct-v0.3 as the current version of the line. If you want something recent in the same weight class, Llama 3.1 8B Instruct runs on the same class of hardware.

Mistral-7B-Instruct-v0.2 hardware requirements

The system requirement to check is memory. The standard GGUF builds for this model are published in TheBloke/Mistral-7B-Instruct-v0.2-GGUF:

MemoryBuild to pickFile size
4 GBQ3_K_S3.16 GB
6 GBQ4_K_M4.37 GB
8 GBQ5_K_M5.13 GB
12 GBQ6_K5.94 GB
16 GB and upQ8_07.70 GB

When two builds both fit, take the larger one: at 7B the quality gap between neighbouring quants costs you a gigabyte or so of memory at most. If the GGUF format is new to you, start with what GGUF is.

How to run Mistral-7B-Instruct-v0.2 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Mistral-7B-Instruct-v0.2 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the newer models in the family, see every Mistral model you can run locally.

Mistral-7B-Instruct-v0.2 license

Mistral-7B-Instruct-v0.2 is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can ship it inside a product or run it on your own hardware without a usage fee. The one caveat comes from Mistral itself: the model ships without moderation mechanisms, so a production deployment is expected to bring its own output filtering.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download mistralai/Mistral-7B-Instruct-v0.2
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.2",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2")
messages = [{"role": "user", "content": "What is your favourite condiment?"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=256)
print(tokenizer.decode(out[0]))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "mistralai/Mistral-7B-Instruct-v0.2",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Mistral-7B-Instruct-v0.2 is a 7-billion-parameter instruction-tuned large language model from Mistral AI. It is a fine-tuned version of the Mistral-7B-v0.2 base model, built for chat and instruction-following tasks. Compared to v0.1, it adds a 32K context window, sets rope-theta to 1e6, and drops sliding-window attention.

At full FP16 precision the model needs roughly 15 GB of VRAM, so a single 16 GB or 24 GB GPU handles it comfortably. With 4-bit quantization (GGUF via llama.cpp or Ollama), it runs on about 6 GB of VRAM and can even run on CPU with enough system RAM, at lower speed.

Yes. Mistral-7B-Instruct-v0.2 is released under the Apache 2.0 license, which permits free commercial use, modification, and redistribution without paying royalties. The weights are downloadable from Hugging Face, and the model can be self-hosted or accessed through providers like Ollama, Cloudflare Workers AI, and Fireworks.

Mistral-7B-Instruct-v0.2 supports a 32K-token context window. This is a fourfold increase over the 8K window in v0.1, and it was achieved by raising rope-theta to 1e6 and removing sliding-window attention, so the full window uses standard dense attention.

Mistral-7B-Instruct-v0.3 is the newer release that supersedes v0.2. The main additions in v0.3 are an extended vocabulary of 32,768 tokens and support for the v3 tokenizer and function calling. The v0.2 model keeps the same 7B size and 32K context but lacks the native tool-calling support and updated vocabulary of v0.3.