What is Qwen3-30B-A3B?
Qwen3-30B-A3B is a mixture-of-experts model from Alibaba's Qwen team, part of the Qwen3 generation. It holds 30.5B parameters in total but activates only 3.3B per token: each forward pass routes through 8 of its 128 experts. That split is what shapes a local run: all 30.5B parameters have to be resident in memory, while the compute spent on each token is that of the 3.3B active parameters. Qwen published the weights on April 27, 2025 under Apache 2.0.
| Specification | Qwen3-30B-A3B |
|---|---|
| Total parameters | 30.5B (29.9B non-embedding) |
| Active parameters | 3.3B per token |
| Experts | 128 total, 8 active |
| Layers | 48 |
| Attention | GQA, 32 heads for Q, 4 for KV |
| Context window | 32,768 tokens native, 131,072 with YaRN |
| Modalities | Text in, text out |
| Reasoning | Thinking on by default, hard and soft switches to turn it off |
| Release date | April 27, 2025 |
| License | Apache 2.0 |
The distinctive feature is the thinking switch: Qwen3 puts reasoning and plain chat in one model. A hard switch in the chat template, enable_thinking, turns reasoning fully on or off. A soft switch steers a live conversation: append /think or /no_think to any message and the model follows the most recent instruction. The two modes want different sampling. Qwen recommends Temperature 0.6 with TopP 0.95 while thinking, 0.7 with 0.8 without, and warns against greedy decoding in thinking mode, which can push the model into endless repetition.
What Qwen3-30B-A3B is good at
Qwen positions the Qwen3 line around reasoning, agents and language coverage. In thinking mode the model surpasses the earlier QwQ on math, code generation and commonsense logical reasoning; with thinking off it beats the Qwen2.5 instruct models it replaces. Qwen also claims leading performance among open-source models on complex agent tasks: the model is trained for precise tool calling in both modes, and the vendor's Qwen-Agent framework connects it to MCP servers through a config file instead of hand-written tool parsers.
The other stated strength is breadth: support for 100+ languages and dialects with multilingual instruction following and translation, plus post-training for human preference that Qwen ties to creative writing, role-play and multi-turn dialogue. The model card publishes no per-benchmark scores, so treat these as vendor claims rather than measured numbers.
Qwen3-30B-A3B hardware requirements
The system requirement to check is memory: all 30.5B parameters stay loaded even though 3.3B are active per token. Qwen publishes official GGUF builds at Qwen/Qwen3-30B-A3B-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 24 GB | Q4_K_M | 18.56 GB |
| 32 GB | Q5_K_M | 21.73 GB |
| 48 GB | Q6_K | 25.09 GB |
| 64 GB and up | Q8_0 | 32.48 GB |
The official repo starts at Q4_K_M, so 24 GB of VRAM or unified memory is the practical floor. When two builds both fit, take the larger one: Q5_0 at 21.08 GB and Q5_K_M at 21.73 GB sit less than a gigabyte apart, and the K_M variant is the one to keep. Out of the box a GGUF runs the native 32,768-token window; to reach 131,072, llama.cpp takes Qwen's YaRN flags, though the vendor advises leaving scaling off unless you actually feed the model long inputs. If the format is new to you, start with what GGUF is.
How to run Qwen3-30B-A3B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3-30B-A3B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
Newer spins of this checkpoint have their own pages, Qwen3-30B-A3B-Instruct-2507 and Qwen3-Coder-30B-A3B, and the full lineup is on our Qwen family page.
Qwen3-30B-A3B license
Qwen3-30B-A3B is released under Apache 2.0, with the license file shipped in the repo. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model, fine-tune it, and run it on your own hardware without a usage fee.
