Qwen2.5-14B-Instruct

Updated
05.10.2026
Reasoning
Code
Multilingual
Tools

A 14.7B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math, and 29+ language support.

At a glance

  • License: Apache 2.0
  • Parameters: 14.7B (13.1B non-embedding)
  • Context length: 131,072 tokens, up to 8,192 generated
  • Modalities: Text input and output
  • Minimum hardware: 8 GB memory with the 5.8 GB Q2_K GGUF build

What is Qwen2.5-14B-Instruct?

Qwen2.5-14B-Instruct is a dense 14.7B instruction-tuned model from Alibaba Cloud's Qwen team, released on September 16, 2024 as the mid-size chat model of the Qwen2.5 series, which runs from 0.5B to 72B parameters. It is a practical local size: the official Q4_K_M GGUF build totals about 9 GB, so the model fits a 12 GB GPU or a 16 GB Mac with room left for context. The weights are published under Apache 2.0.

SpecificationQwen2.5-14B-Instruct
Total parameters14.7B (13.1B non-embedding)
Base modelQwen2.5-14B
Training stagePretraining and post-training
ArchitectureDense transformer with RoPE, SwiGLU, RMSNorm and attention QKV bias
Layers48
Attention heads (GQA)40 for queries, 8 for keys and values
Context window131,072 tokens
Max generation8,192 tokens
ModalitiesText input and output
Release dateSeptember 16, 2024
LicenseApache 2.0

One detail to know before a long-document job: the shipped config.json caps context at 32,768 tokens. The full 131,072-token window is enabled by adding a YaRN rope_scaling block to the config, and Qwen advises adding it only when you actually need long inputs, because static YaRN can cost some quality on short texts. For serving outside a local app, the Qwen team recommends vLLM.

What Qwen2.5-14B-Instruct is good at

Qwen does not print a benchmark table on this model card; detailed evaluation results live in the Qwen2.5 blog post. The card is specific about what changed over Qwen2, though. It claims significantly more knowledge and greatly improved capabilities in coding and mathematics, which the team credits to the specialized expert models it trained in those two domains.

The card lists a second group of improvements alongside those: better instruction following, generating long texts over 8K tokens, understanding structured data such as tables, and generating structured output, especially JSON. It also calls the model more resilient to varied system prompts, which matters for role-play setups and chatbots with fixed conditions. It supports more than 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai and Arabic.

Qwen2.5-14B-Instruct hardware requirements

The system requirement to check is memory. Qwen publishes official GGUF builds in Qwen/Qwen2.5-14B-Instruct-GGUF; each build ships split into files of up to 4 GB, and the sizes below are the totals of those parts.

MemoryBuild to pickFile size
8 GBQ2_K5.8 GB
12 GBQ4_K_M9.0 GB
16 GBQ6_K12.1 GB
24 GBQ8_015.7 GB
32 GB and upFP1629.6 GB

When two builds both fit, take the larger one. Q3_K_M (7.3 GB) and Q5_K_M (10.5 GB) sit between the rows above when you want an intermediate step. If the format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else runs in that footprint.

How to run Qwen2.5-14B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen2.5-14B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally, or step up to Qwen2.5-32B-Instruct if you have the memory for it.

Qwen2.5-14B-Instruct license

Qwen2.5-14B-Instruct is released under Apache 2.0, with the license text linked from the repository. That permits commercial use, modification and redistribution with no royalties: you can build a product on the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download Qwen/Qwen2.5-14B-Instruct
# or serve with vLLM:
vllm serve Qwen/Qwen2.5-14B-Instruct
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-14B-Instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-14B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [{"role": "user", "content": "Give me a short intro to LLMs."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "not-needed" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen2.5-14B-Instruct",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen2.5-14B-Instruct is an instruction-tuned large language model from Alibaba Cloud's Qwen team, part of the Qwen2.5 series released in September 2024. It has 14.7 billion parameters and is built on a causal Transformer architecture with RoPE, SwiGLU, RMSNorm, and GQA attention. It is tuned for chat, instruction following, coding, and math.

At 4-bit quantization (Q4_K_M) the 14.7B model needs roughly 9 GB of memory, so a 12 GB GPU is the practical minimum and 16 GB runs it comfortably with room for context. Higher quants need more: about 11 GB at Q5_K_M, 15 GB at Q8_0, and around 28 GB at full FP16. Apple Silicon Macs with 18 GB or more unified memory can also run it. Long contexts add several GB of KV cache on top.

Qwen2.5-14B-Instruct supports a context window of up to 131,072 tokens (128K) and can generate up to 8,192 tokens. By default the shipped config.json caps context at 32,768 tokens; to use the full 128K window you enable YaRN rope scaling, which the Qwen team recommends only when long inputs are actually needed.

Yes. Qwen2.5-14B-Instruct is released under the Apache 2.0 license, which permits free commercial and private use, modification, and redistribution. The weights are openly available to download from Hugging Face, with no per-token fees when you self-host.

Qwen2.5-14B-Instruct supports more than 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic. The Qwen2.5 series also improved structured-data understanding and JSON output, which helps across multilingual tool-use and chat tasks.