What is Qwen3-8B?
Qwen3-8B is an 8.2B parameter causal language model from Alibaba's Qwen team, published in April 2025 as part of the Qwen3 generation. Its defining feature is a built-in switch between thinking mode, where the model reasons step by step before answering, and non-thinking mode for fast general dialogue, both inside a single model. The Q4_K_M build is 5.03 GB, so it fits on an 8 GB machine.
| Specification | Qwen3-8B |
|---|---|
| Total parameters | 8.2B (6.95B non-embedding) |
| Layers | 36 |
| Attention | GQA, 32 heads for queries and 8 for KV |
| Context window | 32,768 tokens native, 131,072 with YaRN |
| Modalities | Text input and output |
| Reasoning | Thinking mode on by default, switchable per turn |
| Languages | 100+ languages and dialects |
| Release date | April 2025 |
| License | Apache 2.0 |
Thinking mode is on by default and wraps the reasoning in a think block before the final answer. You can disable it entirely with a hard switch, or flip it per turn by adding /think or /no_think to a prompt; the model follows the most recent instruction in a conversation. Qwen recommends different sampling for each mode, temperature 0.6 with top-p 0.95 while thinking and 0.7 with top-p 0.8 without, and warns against greedy decoding, which can cause endless repetitions. Top-k 20 and min-p 0 apply in both modes, and Qwen recommends an output length of 32,768 tokens for most queries.
What Qwen3-8B is good at
Qwen ships no benchmark table on the model card itself, but its claims are specific. In thinking mode the team reports Qwen3-8B surpasses the earlier QwQ on math, code generation and commonsense logical reasoning; in non-thinking mode it surpasses the Qwen2.5 instruct models on the same tasks. Tool calling is a stated focus: the card recommends the Qwen-Agent framework, tools can be defined through an MCP configuration file, and Qwen describes the model's performance on complex agent tasks as leading among open-source models.
The model supports over 100 languages and dialects with multilingual instruction following and translation, and post-training also targets creative writing, role-play and multi-turn dialogue. Native context is 32,768 tokens; Qwen has validated up to 131,072 tokens with YaRN scaling, and advises enabling it only when you actually need long inputs, since static YaRN can degrade quality on short texts. If you have memory to spare, the next size up in the family is Qwen3-14B.
Qwen3-8B hardware requirements
The system requirement to check is memory. Qwen publishes its own official GGUF builds as Qwen/Qwen3-8B-GGUF, and the file sizes below are the real sizes from that repo.
| Memory | Build to pick | File size |
|---|---|---|
| 8 GB | Q4_K_M | 5.03 GB |
| 12 GB | Q6_K | 6.73 GB |
| 16 GB and up | Q8_0 | 8.71 GB |
Q5_0 (5.72 GB) and Q5_K_M (5.85 GB) sit in between, so when two builds both fit, take the larger one. If the GGUF format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits in the same memory.
How to run Qwen3-8B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3-8B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Qwen model you can run locally.
Qwen3-8B license
Qwen3-8B is released under Apache 2.0, with the license file included in the repo. That permits commercial use, modification and redistribution with no royalties, so you can ship products on top of the model and run it on your own hardware without a usage fee.
