What is Qwen3-30B-A3B-Instruct-2507?
Qwen3-30B-A3B-Instruct-2507 is the updated version of the Qwen3-30B-A3B non-thinking mode, released by the Qwen team on July 28, 2025. It is a causal language model with 128 experts, 30.5B parameters in total and 3.3B activated. Qwen reports significant improvements over the original in instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage, substantial gains in long-tail knowledge across multiple languages, better alignment on subjective and open-ended tasks, and enhanced 256K long-context understanding.
| Specification | Qwen3-30B-A3B-Instruct-2507 |
|---|---|
| Total parameters | 30.5B |
| Non-embedding parameters | 29.9B |
| Activated parameters | 3.3B per token |
| Architecture | Mixture of experts, causal language model |
| Experts | 128 total, 8 activated |
| Layers | 48 |
| Attention | GQA, 32 heads for Q and 4 for KV |
| Context window | 262,144 tokens natively, 1M with Dual Chunk Attention |
| Training stage | Pretraining and post-training |
| Mode | Non-thinking only, no <think> blocks |
| Modalities | Text input, text output |
| Release date | July 28, 2025 |
| License | Apache 2.0 |
Eight of the model's 128 experts are activated, across 48 layers, with grouped-query attention using 32 heads for Q and 4 for KV. Qwen also documents a 1M-token configuration built on Dual Chunk Attention plus MInference sparse attention, which it measures at up to 3x faster than standard attention on sequences approaching 1M tokens; that configuration needs roughly 240 GB of total GPU memory for weights, KV cache and peak activations. For sampling, Qwen recommends temperature 0.7, top-p 0.8, top-k 20 and min-p 0, an output length of 16,384 tokens, and a presence penalty between 0 and 2 to reduce endless repetitions.
Qwen3-30B-A3B-Instruct-2507 benchmarks
The numbers below are from the Qwen team's model card, which compares the 2507 update with the original Qwen3-30B-A3B, the larger Qwen3-235B-A22B, DeepSeek-V3-0324, GPT-4o-0327 and Gemini-2.5-Flash, all in non-thinking mode:
| Benchmark | Qwen3-30B-A3B-Instruct-2507 | Qwen3-30B-A3B | Qwen3-235B-A22B | DeepSeek-V3-0324 | GPT-4o-0327 | Gemini-2.5-Flash |
|---|---|---|---|---|---|---|
MMLU-Pro Academic knowledge | 78.4 | 69.1 | 75.2 | 81.2 | 79.8 | 81.1 |
GPQA Expert science | 70.4 | 54.8 | 62.9 | 68.4 | 66.9 | 78.3 |
AIME25 Competition math | 61.3 | 21.6 | 24.7 | 46.6 | 26.7 | 61.6 |
ZebraLogic Logic puzzles | 90.0 | 33.2 | 37.7 | 83.4 | 52.6 | 57.9 |
LiveCodeBench v6 Competitive coding | 43.2 | 29.0 | 32.9 | 45.2 | 35.8 | 40.1 |
MultiPL-E Multilingual coding | 83.8 | 74.6 | 79.3 | 82.2 | 82.7 | 77.7 |
Arena-Hard v2 Human preference | 69.0 | 24.8 | 52.0 | 45.6 | 61.9 | 58.3 |
Creative Writing v3 Creative writing | 86.0 | 68.1 | 80.4 | 81.6 | 84.9 | 84.6 |
The 2507 update takes the top score in this field on ZebraLogic, MultiPL-E, Arena-Hard v2 and Creative Writing v3, and nearly triples its predecessor on AIME25 math. DeepSeek-V3-0324 still leads on MMLU-Pro and LiveCodeBench v6, and Gemini-2.5-Flash on GPQA and AIME25.
Elsewhere in the same table, Qwen puts 2507 first on IFEval at 84.7, WritingBench at 85.5 and PolyMATH at 43.1, and reports 43.0 on HMMT25 against 12.0 for the original. The losses are specific: Aider-Polyglot lands at 35.6 against 59.6 for Qwen3-235B-A22B, TAU2-Telecom at 12.3 is the lowest score in its row, and INCLUDE at 71.9 trails Gemini-2.5-Flash at 83.8. Long context moves most: on the 1M version of RULER, Qwen reports 86.8 average accuracy against 72.0 for the original, holding 89.1 at 128k, 82.5 at 256k and 72.8 at 1000k.
Qwen3-30B-A3B-Instruct-2507 hardware requirements
The system requirement to check is memory, and the file sizes below are real GGUF builds from the unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF repo.
| Memory | Build to pick | File size |
|---|---|---|
| 10 GB | UD-IQ1_M | 9.69 GB |
| 12 GB | UD-IQ2_M | 10.85 GB |
| 16 GB | UD-Q3_K_XL | 13.83 GB |
| 24 GB | Q4_K_M | 18.56 GB |
| 32 GB | Q5_K_M | 21.73 GB |
| 40 GB | Q6_K | 25.09 GB |
| 48 GB and up | Q8_0 | 32.48 GB |
When two builds both fit, take the larger one, and drop to the 9.69 GB UD-IQ1_M only when nothing above it fits.
How to run Qwen3-30B-A3B-Instruct-2507 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3-30B-A3B-Instruct-2507 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Qwen model you can run locally, the original Qwen3-30B-A3B, or the coding sibling Qwen3-Coder-30B-A3B-Instruct.
Qwen3-30B-A3B-Instruct-2507 license
Qwen3-30B-A3B-Instruct-2507 is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model and run it on your own hardware without a usage fee.
