Qwen2.5-32B-Instruct

Updated
05.10.2026
Reasoning
Code
Multilingual
Tools

A 32.5B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math, and 29+ language support.

At a glance

  • License: Apache 2.0
  • Parameters: 32.5B (31.0B non-embedding)
  • Context length: 131,072 tokens, 8,192-token generation
  • Modalities: Text input and output
  • Minimum hardware: 16 GB memory with the Q2_K build (12.3 GB)

What is Qwen2.5-32B-Instruct?

Qwen2.5-32B-Instruct is the 32.5B instruction-tuned model in Alibaba Cloud's Qwen2.5 series, a family of base and instruct models spanning 0.5B to 72B parameters, released in September 2024 under Apache 2.0. It sits at the size where a serious general model still fits on a single 24 GB GPU: the official Q4_K_M build totals 19.9 GB. The Hugging Face repo counts over 2.3 million downloads.

SpecificationQwen2.5-32B-Instruct
Total parameters32.5B (31.0B non-embedding)
Exact parameter count32,763,876,352 in the safetensors weights
Model typeCausal language model
Training stagesPretraining and post-training
ArchitectureTransformer with RoPE, SwiGLU, RMSNorm and attention QKV bias
Layers64
Attention heads (GQA)40 for Q, 8 for KV
Context window131,072 tokens
Context set in config.json32,768 tokens, lifted to the full window by YaRN at factor 4.0
Max generation8,192 tokens
ModalitiesText input and output
LanguagesOver 29, including Chinese, English, French, Spanish, German, Russian, Japanese, Korean and Arabic
Chat templateShipped with the tokenizer, applied through apply_chat_template
Framework minimumtransformers 4.37.0
Deployment stack Qwen recommendsvLLM
Base modelQwen2.5-32B
Release dateSeptember 2024, repo created September 17
LicenseApache 2.0

One note on context: the shipped config.json is set to 32,768 tokens. The full 131,072 needs YaRN rope scaling enabled in the config, a rope_scaling block with a factor of 4.0 over an original 32,768-token window, and Qwen advises adding it only when you actually process long inputs, since static YaRN scaling can cost some quality on short prompts. For serving at long context the team recommends vLLM, the one deployment stack the card names. On the Python side the weights need transformers 4.37.0 or newer; older versions fail to load them with a KeyError on qwen2.

What Qwen2.5-32B-Instruct is good at

Qwen's model card publishes no per-benchmark scores for the 32B; the detailed evaluation results live in the Qwen2.5 blog, and GPU memory and throughput figures live in a separate speed benchmark page in the Qwen documentation. What the card does claim: compared with Qwen2, this generation has significantly more knowledge and much stronger coding and mathematics, trained with specialized expert models in both domains.

The rest of the claim list is specific enough to check against your own prompts. Qwen states significant improvements in instruction following, in generating long texts past 8K tokens, in understanding structured data such as tables, and in generating structured output, JSON in particular. It also calls the model "more resilient to the diversity of system prompts", and ties that to role-play implementation and condition-setting for chatbots, which is the part that matters if you run a fixed persona or a strict output contract in the system slot. Every one of these gains is stated against Qwen2, the previous generation, not against models from other vendors.

Language coverage is broad: over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai and Arabic. That makes it a practical pick when one local model has to cover multilingual chat, code and structured extraction at the same time.

Qwen2.5-32B-Instruct hardware requirements

The system requirement to check is memory. Qwen publishes official GGUF conversions as Qwen/Qwen2.5-32B-Instruct-GGUF; every quant ships as split files capped at about 4 GB, and the sizes below are the totals per build.

MemoryBuild to pickFile size
16 GBQ2_K12.3 GB
20 GBQ3_K_M15.9 GB
24 GBQ4_K_M19.9 GB
32 GBQ5_K_M23.3 GB
36 GBQ6_K26.9 GB
48 GBQ8_034.8 GB
80 GB and upfp1665.5 GB

Neighbouring builds differ by a few gigabytes, so when two of them fit your memory, take the larger one. That matters most at the bottom of the table, where quality falls fastest: Q2_K at 12.3 GB is the smallest build in the repo, and the step up to Q3_K_M buys back the most for 3.6 GB more on disk. The repo also carries the legacy formats, Q4_0 at 18.6 GB and Q5_0 at 22.6 GB, if a runtime asks for them, and the unquantized fp16 conversion totals 65.5 GB, so full precision means multiple GPUs. If the format is new to you, start with what GGUF is.

How to run Qwen2.5-32B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen2.5-32B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally, the coding sibling Qwen2.5-Coder-32B-Instruct, or Qwen2.5-14B-Instruct if 20 GB of weights is more than your machine can hold.

Qwen2.5-32B-Instruct license

Qwen2.5-32B-Instruct is released under Apache 2.0, with the license file included in the repo and linked from the model card itself. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download Qwen/Qwen2.5-32B-Instruct
# or via Ollama
ollama run qwen2.5:32b-instruct
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-32B-Instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-32B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [{"role": "user", "content": "Give me a short intro to LLMs."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.batch_decode(out, skip_special_tokens=True)[0])
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "sk-noauth" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen2.5-32B-Instruct",
  messages: [{ role: "user", content: "Hello!" }]
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen2.5-32B-Instruct is a 32.5-billion-parameter instruction-tuned large language model from Alibaba Cloud's Qwen team, released in September 2024. It is part of the Qwen2.5 series and is fine-tuned from the Qwen2.5-32B base model for chat, instruction following, coding, and mathematics. It supports a 128K-token context window and over 29 languages.

At 4-bit quantization (Q4_K_M / GGUF or AWQ), Qwen2.5-32B-Instruct needs roughly 20 GB of memory and runs on a single 24 GB GPU such as an RTX 3090 or RTX 4090, or an Apple Silicon Mac with 32 GB of unified memory. Full FP16 precision requires about 64 GB, which means multiple GPUs or a 64 GB+ Mac. Reducing the context length lowers memory use further.

Yes. Qwen2.5-32B-Instruct is released under the Apache 2.0 license, which permits both research and commercial use with no licensing fees. You can download the weights from Hugging Face, run them locally, fine-tune them, and deploy them in commercial products as long as you comply with the standard Apache 2.0 terms.

Qwen2.5-32B-Instruct supports a full context length of 131,072 tokens (128K) and can generate up to 8,192 tokens in a single response. The shipped config.json defaults to 32,768 tokens; to use the full 128K window you enable YaRN rope scaling, which the Qwen team recommends turning on only when you actually need long-context processing.

The 72B model scores higher across most benchmarks, including MMLU, GSM8K, and HumanEval. But Qwen2.5-32B-Instruct gives a much better performance-per-GPU ratio: it runs on a single 24 GB card at 4-bit, while the 72B model needs far more memory. For many chat, reasoning, and coding tasks the 32B version is the practical choice for local deployment, and it outperforms the older Qwen2-72B.