Qwen2.5-Coder-7B-Instruct

Updated
05.10.2026
Code
Reasoning
Tools

A 7.6B code-specialized LLM from Alibaba’s Qwen2.5-Coder series, tuned for code generation, reasoning, and fixing.

At a glance

  • License: Apache 2.0
  • Parameters: 7.61B (6.53B non-embedding)
  • Context length: 131,072 tokens (32,768 by default, YaRN beyond that)
  • Modalities: Text in, text out
  • Minimum hardware: 6 GB of free memory (Q3_K_M GGUF, 3.81 GB)

What is Qwen2.5-Coder-7B-Instruct?

Qwen2.5-Coder-7B-Instruct is the instruction-tuned 7.61B parameter member of Alibaba's Qwen2.5-Coder family, the code-specific line of Qwen models formerly known as CodeQwen. The series ships in six sizes, from 0.5B to 32B, and the 7B sits in the middle of that range, small enough that a quantized build fits an 8 GB laptop. The weights went up on Hugging Face in September 2024 under Apache 2.0.

SpecificationQwen2.5-Coder-7B-Instruct
Total parameters7.61B (6.53B non-embedding)
ArchitectureCausal transformer with RoPE, SwiGLU, RMSNorm and attention QKV bias
Layers28
Attention headsGQA, 28 for queries and 4 for KV
Context window131,072 tokens (32,768 by default, YaRN beyond that)
TrainingPretraining and post-training, 5.5 trillion tokens across the series
Release dateSeptember 2024
LicenseApache 2.0

The config as shipped caps context at 32,768 tokens. To go past that, Qwen uses YaRN rope scaling with a factor of 4.0, which unlocks the full 131,072 tokens. The team advises adding the rope_scaling block only when you actually feed long inputs: the supported implementation is static, so the scaling applies at every length and can cost some quality on short prompts. Loading the model needs transformers 4.37 or newer, and for serving Qwen recommends vLLM.

What Qwen2.5-Coder-7B-Instruct is good at

Qwen positions the Coder series around three jobs: code generation, code reasoning and code fixing, all of which it says improved significantly over CodeQwen1.5. The gain comes from scale: training data grew to 5.5 trillion tokens, mixing source code, text-code grounding and synthetic data on top of the Qwen2.5 base. The team also names code agents as a target use case, and states the models keep the mathematics and general competence of Qwen2.5 rather than trading everything for code.

The model card publishes no benchmark table; Qwen keeps the detailed evaluation results in its blog post on the Coder family. The one comparative claim on the card is about the 32B flagship, which Qwen calls the state-of-the-art open-source code LLM, matching the coding ability of GPT-4o. The 7B gives you the same training recipe in a size that runs locally; if your machine has the memory for the bigger one, see Qwen2.5-Coder-32B-Instruct.

Qwen2.5-Coder-7B-Instruct hardware requirements

The system requirement to check is memory: the model file has to fit in your RAM or VRAM with room left over for context. Qwen publishes official GGUF builds in Qwen/Qwen2.5-Coder-7B-Instruct-GGUF; these are the real file sizes:

MemoryBuild to pickFile size
6 GBQ3_K_M3.81 GB
8 GBQ4_K_M4.68 GB
12 GBQ6_K6.25 GB
16 GBQ8_08.10 GB
24 GB and upFP1615.24 GB

When two builds both fit, take the larger one. The same repo also lists split copies of several builds, fp16 in four numbered parts and q8_0 in three, alongside the single files in the table above. If the format is new to you, start with what GGUF is.

How to run Qwen2.5-Coder-7B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen2.5-Coder-7B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally.

Qwen2.5-Coder-7B-Instruct license

Qwen2.5-Coder-7B-Instruct is released under Apache 2.0, with the license file included in the repo. That permits commercial use, modification and redistribution with no royalties, so you can ship it inside a product, fine-tune it, or run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download Qwen/Qwen2.5-Coder-7B-Instruct
# or run quantized via Ollama:
ollama run qwen2.5-coder:7b-instruct
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-Coder-7B-Instruct",
    "messages": [{"role": "user", "content": "Write a quicksort in Python."}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-Coder-7B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [{"role": "user", "content": "Write a quicksort in Python."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.batch_decode(out, skip_special_tokens=True)[0])
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "sk-local" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen2.5-Coder-7B-Instruct",
  messages: [{ role: "user", content: "Write a quicksort in Python." }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen2.5-Coder-7B-Instruct is a 7.61B-parameter, code-specialized large language model from Alibaba's Qwen team, released in September 2024. It is the instruction-tuned variant of Qwen2.5-Coder-7B, fine-tuned for code generation, code reasoning, and bug fixing. The Qwen2.5-Coder series was trained on 5.5 trillion tokens including source code and text-code grounding data.

At full BF16 precision the 7.61B-parameter model needs roughly 16 GB of VRAM. With 4-bit quantization (GGUF or GPTQ) it fits comfortably in about 8 GB, making it runnable on a single consumer GPU such as an RTX 3060/4060 or even on Apple Silicon via llama.cpp. Quantized GGUF builds are available for use with Ollama and llama.cpp.

Yes. Qwen2.5-Coder-7B-Instruct is released under the Apache 2.0 license, which permits free commercial use, modification, and redistribution with attribution. The weights are openly available on Hugging Face, so you can download and self-host the model at no cost.

The model supports a context window of up to 131,072 tokens (128K). The default config.json is set to 32,768 tokens; to handle inputs beyond that, you enable YaRN rope scaling in the config. For long-context deployment the Qwen team recommends vLLM, noting that static YaRN can slightly affect performance on shorter inputs.

Qwen2.5-Coder-32B-Instruct is the flagship of the series and scores higher on coding benchmarks, with the Qwen team describing its coding ability as matching GPT-4o. The 7B model trades some accuracy for far lower hardware demands: it runs on a single consumer GPU and is much faster, which makes it a practical choice for local coding assistants and IDE integration when a 32B model is too heavy.