Qwen2.5-Coder-32B-Instruct

Updated
05.10.2026
Code
Reasoning
Tools
Multilingual

A 32.5B code-specialized LLM from Alibaba’s Qwen2.5-Coder series with open-model state-of-the-art coding ability and 128K context.

At a glance

  • License: Apache 2.0
  • Parameters: 32.5B (31.0B non-embedding)
  • Context length: 131,072 tokens (32,768 default, YaRN beyond)
  • Modalities: Text only
  • Minimum hardware: 16 GB memory (Q2_K GGUF, 12.31 GB)

What is Qwen2.5-Coder-32B-Instruct?

Qwen2.5-Coder-32B-Instruct is the largest model in Alibaba's Qwen2.5-Coder series, the line of code-specific Qwen models formerly known as CodeQwen. Qwen trained the series on 5.5 trillion tokens of source code, text-code grounding and synthetic data, and calls the 32B the state-of-the-art open-source code LLM at release, with coding ability it reports as matching GPT-4o. The model is dense, 32.5B parameters, so the vendor's own Q4_K_M build fits on a single 24 GB GPU. Alibaba published the weights on November 6, 2024 under Apache 2.0.

SpecificationQwen2.5-Coder-32B-Instruct
Total parameters32.5B (31.0B non-embedding)
ArchitectureDense transformer with RoPE, SwiGLU, RMSNorm and attention QKV bias
Layers64
Attention headsGQA, 40 for Q and 8 for KV
Context window131,072 tokens (32,768 without YaRN)
ModalitiesText only
Training data5.5 trillion tokens: source code, text-code grounding, synthetic data
Release dateNovember 6, 2024
LicenseApache 2.0

One spec needs a closer look. The config.json ships with the context window set to 32,768 tokens, and the full 131,072 comes from YaRN, a rope-scaling method you switch on in the config with a factor of 4.0. Qwen advises adding it only when your inputs actually run long: the serving stack Qwen recommends, vLLM, supports only static YaRN, where the scaling factor stays constant regardless of input length and can cost quality on short texts. For everyday coding sessions the default 32K window is the better setting.

What Qwen2.5-Coder-32B-Instruct is good at

The model card names three areas where the series moved past CodeQwen1.5: code generation, code reasoning and code fixing. The 32B-Instruct is the strongest of the six sizes in the series (0.5, 1.5, 3, 7, 14 and 32 billion parameters) and the one carrying the GPT-4o comparison. Qwen did not print score tables in the model card itself, so there is no benchmark table here; the detailed evaluation results live in the Qwen2.5-Coder family blog post.

Qwen also positions the model as a foundation for code agents: the training kept the mathematics and general competencies of the Qwen2.5 base instead of trading them away for code scores, which matters when an agent has to reason about a task and not just emit a diff.

Qwen2.5-Coder-32B-Instruct hardware requirements

The system requirement to check is memory. Qwen publishes official GGUF conversions as Qwen/Qwen2.5-Coder-32B-Instruct-GGUF, and the file sizes below come from that repo.

MemoryBuild to pickFile size
16 GBQ2_K12.31 GB
24 GBQ4_K_M19.85 GB
32 GBQ5_K_M23.26 GB
48 GBQ6_K26.89 GB
64 GB and upQ8_034.82 GB

When two builds both fit, take the larger one: the step up in quality costs nothing but disk space. If the GGUF format is new to you, start with what GGUF is.

How to run Qwen2.5-Coder-32B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen2.5-Coder-32B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

If the 32B is more than your machine holds, Qwen2.5-Coder-7B-Instruct is the same recipe at a fraction of the size, and every Qwen model you can run locally is on the family page.

Qwen2.5-Coder-32B-Instruct license

Qwen2.5-Coder-32B-Instruct is released under Apache 2.0. That permits commercial use, modification, and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download Qwen/Qwen2.5-Coder-32B-Instruct
# or run quantized locally:
ollama run qwen2.5-coder:32b
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-Coder-32B-Instruct",
    "messages": [{"role": "user", "content": "Refactor this function for readability."}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-Coder-32B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [{"role": "user", "content": "Write a quicksort in Python."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "not-needed" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen2.5-Coder-32B-Instruct",
  messages: [{ role: "user", content: "Write a binary search in TypeScript." }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen2.5-Coder-32B-Instruct is a 32.5B-parameter, code-specialized large language model from Alibaba's Qwen team. It is the instruction-tuned variant of Qwen2.5-Coder-32B, trained on 5.5 trillion tokens of source code, text-code grounding data, and synthetic data. It targets code generation, code reasoning, and code fixing, and supports a 128K-token context window.

Yes. Qwen2.5-Coder-32B-Instruct is released under the Apache 2.0 license, which permits free commercial and private use, modification, and redistribution with attribution. The weights are published on Hugging Face and can be downloaded and run locally at no cost.

At a 4-bit quantization the 32B model fits in roughly 24 GB of VRAM, so a single 24 GB GPU such as an RTX 3090 or 4090 can run it. Running in full BF16 precision needs about 65 GB and typically two high-memory GPUs. The model also runs on Apple Silicon Macs with 32 GB or more of unified memory using llama.cpp or Ollama with quantized GGUF builds.

It was the strongest open-source code model at release, with coding ability the Qwen team reports as matching GPT-4o. It leads open models on EvalPlus, LiveCodeBench, and BigCodeBench, scores 73.7 on the Aider code-repair benchmark, and reaches 65.9 on McEval across more than 40 programming languages.

The model supports a full context length of 131,072 tokens (128K). The default config.json caps context at 32,768 tokens; to process longer inputs you enable YaRN rope scaling in the config, which Qwen recommends only when long-context handling is actually needed since static YaRN can affect performance on shorter inputs.