Qwen3-8B

Updated
05.10.2026
Thinking
Tools
Reasoning
Code
Multilingual

An 8.2B dense LLM from Alibaba’s Qwen3 series with switchable thinking mode, strong reasoning, coding, and 100+ language support.

At a glance

  • License: Apache 2.0
  • Parameters: 8.2B (6.95B non-embedding)
  • Context length: 32,768 native, 131,072 with YaRN
  • Modalities: Text input and output
  • Minimum hardware: 8 GB of memory (Q4_K_M, 5.03 GB)

What is Qwen3-8B?

Qwen3-8B is an 8.2B parameter causal language model from Alibaba's Qwen team, published in April 2025 as part of the Qwen3 generation. Its defining feature is a built-in switch between thinking mode, where the model reasons step by step before answering, and non-thinking mode for fast general dialogue, both inside a single model. The Q4_K_M build is 5.03 GB, so it fits on an 8 GB machine.

SpecificationQwen3-8B
Total parameters8.2B (6.95B non-embedding)
Layers36
AttentionGQA, 32 heads for queries and 8 for KV
Context window32,768 tokens native, 131,072 with YaRN
ModalitiesText input and output
ReasoningThinking mode on by default, switchable per turn
Languages100+ languages and dialects
Release dateApril 2025
LicenseApache 2.0

Thinking mode is on by default and wraps the reasoning in a think block before the final answer. You can disable it entirely with a hard switch, or flip it per turn by adding /think or /no_think to a prompt; the model follows the most recent instruction in a conversation. Qwen recommends different sampling for each mode, temperature 0.6 with top-p 0.95 while thinking and 0.7 with top-p 0.8 without, and warns against greedy decoding, which can cause endless repetitions. Top-k 20 and min-p 0 apply in both modes, and Qwen recommends an output length of 32,768 tokens for most queries.

What Qwen3-8B is good at

Qwen ships no benchmark table on the model card itself, but its claims are specific. In thinking mode the team reports Qwen3-8B surpasses the earlier QwQ on math, code generation and commonsense logical reasoning; in non-thinking mode it surpasses the Qwen2.5 instruct models on the same tasks. Tool calling is a stated focus: the card recommends the Qwen-Agent framework, tools can be defined through an MCP configuration file, and Qwen describes the model's performance on complex agent tasks as leading among open-source models.

The model supports over 100 languages and dialects with multilingual instruction following and translation, and post-training also targets creative writing, role-play and multi-turn dialogue. Native context is 32,768 tokens; Qwen has validated up to 131,072 tokens with YaRN scaling, and advises enabling it only when you actually need long inputs, since static YaRN can degrade quality on short texts. If you have memory to spare, the next size up in the family is Qwen3-14B.

Qwen3-8B hardware requirements

The system requirement to check is memory. Qwen publishes its own official GGUF builds as Qwen/Qwen3-8B-GGUF, and the file sizes below are the real sizes from that repo.

MemoryBuild to pickFile size
8 GBQ4_K_M5.03 GB
12 GBQ6_K6.73 GB
16 GB and upQ8_08.71 GB

Q5_0 (5.72 GB) and Q5_K_M (5.85 GB) sit in between, so when two builds both fit, take the larger one. If the GGUF format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits in the same memory.

How to run Qwen3-8B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3-8B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally.

Qwen3-8B license

Qwen3-8B is released under Apache 2.0, with the license file included in the repo. That permits commercial use, modification and redistribution with no royalties, so you can ship products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U "transformers>=4.51.0"
huggingface-cli download Qwen/Qwen3-8B
# or with Ollama:
ollama run qwen3:8b
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-8B",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen3-8B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Give me a short intro to LLMs."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=2048)
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "sk-local" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen3-8B",
  messages: [{ role: "user", content: "Hello!" }]
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen3-8B is an 8.2-billion-parameter dense language model from Alibaba's Qwen team, released as part of the Qwen3 generation. It is a causal language model post-trained for chat, reasoning, and agentic tool use. A defining feature is its dual-mode design: it can switch between a thinking mode for math, coding, and logic, and a non-thinking mode for fast general dialogue, all within the same checkpoint.

At 4-bit quantization (Q4_K_M) the model weights take roughly 5 GB, so a GPU with 8 GB of VRAM can run it comfortably at standard context lengths. Q8 needs about 9 GB. CPU-only inference works at Q4_K_M with 16 GB of system RAM, though throughput drops to a few tokens per second. Extending context toward 128K with YaRN adds several GB of KV-cache memory on top of these figures.

Qwen3-8B has a native context window of 32,768 tokens. Using YaRN rope scaling it can be extended to 131,072 tokens (128K) for long-document and long-context tasks. The Qwen team recommends enabling YaRN only when you actually need the longer window, because it can slightly reduce quality on inputs shorter than 32K and increases KV-cache memory use.

Yes. Qwen3-8B is released under the Apache 2.0 license, which permits free use, modification, redistribution, and commercial deployment without paying royalties. The weights are openly available on Hugging Face, so you can download and run the model on your own hardware. Apache 2.0 requires you to preserve the license and copyright notices in redistributions.

Qwen3-8B supports over 100 languages and dialects, with strong multilingual instruction-following and translation ability. This wide coverage spans major European, Asian, and Middle Eastern languages, making it usable for cross-lingual chat and translation tasks well beyond English and Chinese.