Qwen3-30B-A3B

Updated
24.08.2026
Tools
Thinking
Reasoning
Code
Multilingual

Qwen3-30B-A3B is a 30.5B MoE model that activates 3.3B params per token, with a switchable thinking mode and 100+ languages.

At a glance

  • License: Apache 2.0
  • Parameters: 30.5B total, 3.3B active per token
  • Context length: 32,768 native, 131,072 with YaRN
  • Modalities: Text in, text out
  • Minimum hardware: 24 GB memory (Q4_K_M, 18.56 GB)

What is Qwen3-30B-A3B?

Qwen3-30B-A3B is a mixture-of-experts model from Alibaba's Qwen team, part of the Qwen3 generation. It holds 30.5B parameters in total but activates only 3.3B per token: each forward pass routes through 8 of its 128 experts. That split is what shapes a local run: all 30.5B parameters have to be resident in memory, while the compute spent on each token is that of the 3.3B active parameters. Qwen published the weights on April 27, 2025 under Apache 2.0.

SpecificationQwen3-30B-A3B
Total parameters30.5B (29.9B non-embedding)
Active parameters3.3B per token
Experts128 total, 8 active
Layers48
AttentionGQA, 32 heads for Q, 4 for KV
Context window32,768 tokens native, 131,072 with YaRN
ModalitiesText in, text out
ReasoningThinking on by default, hard and soft switches to turn it off
Release dateApril 27, 2025
LicenseApache 2.0

The distinctive feature is the thinking switch: Qwen3 puts reasoning and plain chat in one model. A hard switch in the chat template, enable_thinking, turns reasoning fully on or off. A soft switch steers a live conversation: append /think or /no_think to any message and the model follows the most recent instruction. The two modes want different sampling. Qwen recommends Temperature 0.6 with TopP 0.95 while thinking, 0.7 with 0.8 without, and warns against greedy decoding in thinking mode, which can push the model into endless repetition.

What Qwen3-30B-A3B is good at

Qwen positions the Qwen3 line around reasoning, agents and language coverage. In thinking mode the model surpasses the earlier QwQ on math, code generation and commonsense logical reasoning; with thinking off it beats the Qwen2.5 instruct models it replaces. Qwen also claims leading performance among open-source models on complex agent tasks: the model is trained for precise tool calling in both modes, and the vendor's Qwen-Agent framework connects it to MCP servers through a config file instead of hand-written tool parsers.

The other stated strength is breadth: support for 100+ languages and dialects with multilingual instruction following and translation, plus post-training for human preference that Qwen ties to creative writing, role-play and multi-turn dialogue. The model card publishes no per-benchmark scores, so treat these as vendor claims rather than measured numbers.

Qwen3-30B-A3B hardware requirements

The system requirement to check is memory: all 30.5B parameters stay loaded even though 3.3B are active per token. Qwen publishes official GGUF builds at Qwen/Qwen3-30B-A3B-GGUF.

MemoryBuild to pickFile size
24 GBQ4_K_M18.56 GB
32 GBQ5_K_M21.73 GB
48 GBQ6_K25.09 GB
64 GB and upQ8_032.48 GB

The official repo starts at Q4_K_M, so 24 GB of VRAM or unified memory is the practical floor. When two builds both fit, take the larger one: Q5_0 at 21.08 GB and Q5_K_M at 21.73 GB sit less than a gigabyte apart, and the K_M variant is the one to keep. Out of the box a GGUF runs the native 32,768-token window; to reach 131,072, llama.cpp takes Qwen's YaRN flags, though the vendor advises leaving scaling off unless you actually feed the model long inputs. If the format is new to you, start with what GGUF is.

How to run Qwen3-30B-A3B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3-30B-A3B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

Newer spins of this checkpoint have their own pages, Qwen3-30B-A3B-Instruct-2507 and Qwen3-Coder-30B-A3B, and the full lineup is on our Qwen family page.

Qwen3-30B-A3B license

Qwen3-30B-A3B is released under Apache 2.0, with the license file shipped in the repo. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model, fine-tune it, and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download Qwen/Qwen3-30B-A3B
from transformers import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3-30B-A3B")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen3-30B-A3B is an open-weight Mixture-of-Experts model from Qwen with 30.5B total parameters and roughly 3B active per token. It supports a 128K context window, tool calling, multilingual text, and a switchable thinking mode for step-by-step reasoning. The A3B in the name refers to the 3B active parameters that give it small-model speed.

Because it is a MoE model, all 30.5B parameters load into memory even though only 3B run per token. A 4-bit quant (Q4_K_M) needs around 17 GB, which fits a 24 GB GPU like the RTX 4090, and many people run it with at least 32 GB of system or unified memory on Apple Silicon. Higher precision needs more.

Yes. It is released under the Apache 2.0 license, so the weights are free to download from Hugging Face and free to run. The license also allows commercial use and modification, so there is no per-token or API fee when you run it yourself.

It runs fully offline once the weights are downloaded. In Atomic Chat the model loads on your own hardware and answers with no internet connection, so your prompts and files stay on-device. That makes it a fit for private or air-gapped work.

Qwen3-30B-A3B ships with thinking mode enabled by default. You can switch it per turn by adding /think or /no_think to your prompt, or set the enable_thinking parameter to False through the chat template or API. Turning it off gives faster, shorter replies for simple chat, while leaving it on helps with reasoning, math, and code.