Qwen2.5-72B-Instruct

Updated
05.10.2026
Reasoning
Code
Multilingual
Tools

A 72.7B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math and multilingual ability across 29+ languages.

At a glance

  • License: Qwen license (custom, full text in the model repo)
  • Parameters: 72.7B total, 70.0B non-embedding
  • Context length: 131,072 tokens, generation up to 8,192
  • Modalities: Text input, text output
  • Minimum hardware: 32 GB memory (Q2_K GGUF, 27.3 GB)

What is Qwen2.5-72B-Instruct?

Qwen2.5-72B-Instruct is the 72.7B instruction-tuned flagship of Alibaba's Qwen2.5 series, published on Hugging Face in September 2024. The series spans base and instruct models from 0.5 to 72 billion parameters, and this is the largest of the instruct line: a causal language model with 80 layers, grouped-query attention and a 131,072-token context window. It is still a practical local model, because Qwen publishes official GGUF builds that start at 27.3 GB.

SpecificationQwen2.5-72B-Instruct
Total parameters72.7B (70.0B non-embedding)
ArchitectureTransformers with RoPE, SwiGLU, RMSNorm and Attention QKV bias
Layers80
Attention heads (GQA)64 for Q, 8 for KV
Context window131,072 tokens, generation up to 8,192
LanguagesOver 29, including Chinese, English, French, German, Russian, Japanese and Arabic
Release dateSeptember 2024
LicenseQwen license

One detail to know before you rely on the long context: the shipped config.json is set to 32,768 tokens. The full 131,072-token window needs a YaRN rope_scaling block added to the config, and Qwen advises adding it only when you actually process long inputs, because vLLM, the deployment stack Qwen recommends, presently supports only static YaRN: the scaling factor stays constant regardless of input length, which can affect performance on shorter texts.

What Qwen2.5-72B-Instruct is good at

Qwen keeps its detailed evaluation results in the Qwen2.5 blog rather than on the model card, so what follows is the team's own summary of the release. Compared with Qwen2, the model has significantly more knowledge and greatly improved capabilities in coding and mathematics, which the team credits to its specialized expert models in those domains. Instruction following got better, along with two abilities that matter for pipelines: understanding structured data such as tables, and generating structured output, especially JSON.

The model also generates long texts past 8K tokens and is more resilient to diverse system prompts, which the team calls out as useful for role-play and condition-setting in chatbots. Multilingual coverage spans more than 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai and Arabic.

Qwen2.5-72B-Instruct hardware requirements

The system requirement to check is memory. Qwen publishes official quantized builds at Qwen/Qwen2.5-72B-Instruct-GGUF. Each build ships as split parts; the sizes below are the totals across all parts of a build.

MemoryBuild to pickFile size
32 GBQ2_K27.3 GB
48 GBQ4_K_M44.0 GB
64 GBQ5_K_M51.7 GB
80 GBQ6_K59.9 GB
96 GB and upQ8_077.5 GB

When two builds both fit, take the larger one. The repo also carries Q3_K_M at 35.5 GB and Q4_0 at 41.3 GB if your memory lands between the tiers above, and the unquantized fp16 at 145.8 GB. If the format is new to you, start with what GGUF is. If you serve the original weights through Hugging Face transformers instead, Qwen asks for version 4.37.0 or newer; older versions stop with KeyError: 'qwen2'.

How to run Qwen2.5-72B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen2.5-72B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally, or step down to the sibling Qwen2.5-32B-Instruct if 72B is more than your machine holds.

Qwen2.5-72B-Instruct license

Qwen2.5-72B-Instruct is released under the Qwen license, Alibaba's own terms rather than a standard permissive license like Apache 2.0. The full text ships as the LICENSE file in the model repo, so read it there before building a commercial product on the model.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download Qwen/Qwen2.5-72B-Instruct
# Or serve with vLLM:
vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 2
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-72B-Instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-72B-Instruct", torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-72B-Instruct")
messages = [{"role": "user", "content": "Give me a short intro to LLMs."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(out[0], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "EMPTY" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen2.5-72B-Instruct",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen2.5-72B-Instruct is a 72.7-billion-parameter instruction-tuned large language model from the Qwen team at Alibaba Cloud, released in September 2024. It is a causal (decoder-only) transformer post-trained for chat, coding, math, and structured-output tasks, and it supports a context window of up to 128K tokens.

At full BF16/FP16 precision the 72B weights need roughly 145 GB of memory, so running unquantized usually means multiple GPUs (for example 2x A100 80GB). With 4-bit quantization the model fits in about 45-48 GB of VRAM, which makes a single 48 GB card or two 24 GB cards workable. For long 128K-token contexts you also need extra memory for the KV cache.

The weights are openly available to download from Hugging Face at no cost. The 72B model is released under the Qwen License rather than a standard permissive license such as Apache 2.0. The Qwen License allows commercial use, but products or services with more than 100 million monthly active users must request a separate license from Alibaba Cloud.

Qwen2.5-72B-Instruct supports more than 29 languages. These include Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic. This broad multilingual coverage makes it suitable for translation, cross-language chat, and content generation outside English.

The model supports a full context length of 131,072 tokens (about 128K) and can generate up to 8,192 tokens in a single response. The shipped config.json defaults to 32,768 tokens; to use the full 128K window you enable YaRN rope scaling, which Qwen recommends only when long-context input is actually needed since it can reduce quality on short prompts.