Qwen3-30B-A3B-Instruct-2507

Updated
05.10.2026
Reasoning
Code
Multilingual
Tools

A 30.5B-parameter (3.3B active) MoE instruct model from Alibaba’s Qwen3 series with 256K context and strong reasoning, coding, and tool use.

At a glance

  • License: Apache 2.0
  • Parameters: 30.5B total, 3.3B activated (128 experts, 8 activated)
  • Context length: 262,144 tokens natively, 1M with Dual Chunk Attention
  • Modalities: Text input and output
  • Minimum hardware: 10 GB of memory (UD-IQ1_M GGUF, 9.69 GB)

What is Qwen3-30B-A3B-Instruct-2507?

Qwen3-30B-A3B-Instruct-2507 is the updated version of the Qwen3-30B-A3B non-thinking mode, released by the Qwen team on July 28, 2025. It is a causal language model with 128 experts, 30.5B parameters in total and 3.3B activated. Qwen reports significant improvements over the original in instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage, substantial gains in long-tail knowledge across multiple languages, better alignment on subjective and open-ended tasks, and enhanced 256K long-context understanding.

SpecificationQwen3-30B-A3B-Instruct-2507
Total parameters30.5B
Non-embedding parameters29.9B
Activated parameters3.3B per token
ArchitectureMixture of experts, causal language model
Experts128 total, 8 activated
Layers48
AttentionGQA, 32 heads for Q and 4 for KV
Context window262,144 tokens natively, 1M with Dual Chunk Attention
Training stagePretraining and post-training
ModeNon-thinking only, no <think> blocks
ModalitiesText input, text output
Release dateJuly 28, 2025
LicenseApache 2.0

Eight of the model's 128 experts are activated, across 48 layers, with grouped-query attention using 32 heads for Q and 4 for KV. Qwen also documents a 1M-token configuration built on Dual Chunk Attention plus MInference sparse attention, which it measures at up to 3x faster than standard attention on sequences approaching 1M tokens; that configuration needs roughly 240 GB of total GPU memory for weights, KV cache and peak activations. For sampling, Qwen recommends temperature 0.7, top-p 0.8, top-k 20 and min-p 0, an output length of 16,384 tokens, and a presence penalty between 0 and 2 to reduce endless repetitions.

Qwen3-30B-A3B-Instruct-2507 benchmarks

The numbers below are from the Qwen team's model card, which compares the 2507 update with the original Qwen3-30B-A3B, the larger Qwen3-235B-A22B, DeepSeek-V3-0324, GPT-4o-0327 and Gemini-2.5-Flash, all in non-thinking mode:

BenchmarkQwen3-30B-A3B-Instruct-2507Qwen3-30B-A3BQwen3-235B-A22BDeepSeek-V3-0324GPT-4o-0327Gemini-2.5-Flash
MMLU-Pro
Academic knowledge
78.469.175.281.279.881.1
GPQA
Expert science
70.454.862.968.466.978.3
AIME25
Competition math
61.321.624.746.626.761.6
ZebraLogic
Logic puzzles
90.033.237.783.452.657.9
LiveCodeBench v6
Competitive coding
43.229.032.945.235.840.1
MultiPL-E
Multilingual coding
83.874.679.382.282.777.7
Arena-Hard v2
Human preference
69.024.852.045.661.958.3
Creative Writing v3
Creative writing
86.068.180.481.684.984.6

The 2507 update takes the top score in this field on ZebraLogic, MultiPL-E, Arena-Hard v2 and Creative Writing v3, and nearly triples its predecessor on AIME25 math. DeepSeek-V3-0324 still leads on MMLU-Pro and LiveCodeBench v6, and Gemini-2.5-Flash on GPQA and AIME25.

Elsewhere in the same table, Qwen puts 2507 first on IFEval at 84.7, WritingBench at 85.5 and PolyMATH at 43.1, and reports 43.0 on HMMT25 against 12.0 for the original. The losses are specific: Aider-Polyglot lands at 35.6 against 59.6 for Qwen3-235B-A22B, TAU2-Telecom at 12.3 is the lowest score in its row, and INCLUDE at 71.9 trails Gemini-2.5-Flash at 83.8. Long context moves most: on the 1M version of RULER, Qwen reports 86.8 average accuracy against 72.0 for the original, holding 89.1 at 128k, 82.5 at 256k and 72.8 at 1000k.

Qwen3-30B-A3B-Instruct-2507 hardware requirements

The system requirement to check is memory, and the file sizes below are real GGUF builds from the unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF repo.

MemoryBuild to pickFile size
10 GBUD-IQ1_M9.69 GB
12 GBUD-IQ2_M10.85 GB
16 GBUD-Q3_K_XL13.83 GB
24 GBQ4_K_M18.56 GB
32 GBQ5_K_M21.73 GB
40 GBQ6_K25.09 GB
48 GB and upQ8_032.48 GB

When two builds both fit, take the larger one, and drop to the 9.69 GB UD-IQ1_M only when nothing above it fits.

How to run Qwen3-30B-A3B-Instruct-2507 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3-30B-A3B-Instruct-2507 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally, the original Qwen3-30B-A3B, or the coding sibling Qwen3-Coder-30B-A3B-Instruct.

Qwen3-30B-A3B-Instruct-2507 license

Qwen3-30B-A3B-Instruct-2507 is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U "transformers>=4.51.0"
huggingface-cli download Qwen/Qwen3-30B-A3B-Instruct-2507
# Or serve an OpenAI-compatible endpoint with vLLM:
vllm serve Qwen/Qwen3-30B-A3B-Instruct-2507 --max-model-len 262144
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-30B-A3B-Instruct-2507",
    "messages": [{"role": "user", "content": "Explain mixture-of-experts in one paragraph."}],
    "temperature": 0.7,
    "top_p": 0.8
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen3-30B-A3B-Instruct-2507"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Give me a short intro to MoE models."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=16384)
print(tokenizer.decode(out[0][len(inputs.input_ids[0]):], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "EMPTY" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen3-30B-A3B-Instruct-2507",
  messages: [{ role: "user", content: "Explain mixture-of-experts in one paragraph." }],
  temperature: 0.7,
  top_p: 0.8,
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen3-30B-A3B-Instruct-2507 is an instruction-tuned language model from Alibaba's Qwen team, released in July 2025. It uses a mixture-of-experts (MoE) design with 30.5B total parameters but only 3.3B active per token (128 experts, 8 activated). It runs in non-thinking mode only, so it does not emit <think> reasoning blocks, and it improves on the original Qwen3-30B-A3B in instruction following, math, coding, and tool use.

Because only 3.3B of its 30.5B parameters are active per token, the model is lighter to run than a dense 30B. At a 4-bit quant the weights fit in roughly 18-20 GB, so a single 24 GB GPU (such as an RTX 4090) can run it with a moderate context window. Running at the full 256K context needs far more memory, and the 1M-token configuration requires about 240 GB of total GPU memory across multiple cards.

Yes. The model is released under the Apache 2.0 license, which permits commercial use, modification, and redistribution without paying royalties. The weights are published openly on Hugging Face, and the license requires only that you keep the copyright and license notices. You can self-host it or use it through API providers like OpenRouter and Fireworks AI.

Both share the same 30B-A3B MoE backbone, but they are tuned for different jobs. Qwen3-30B-A3B-Instruct-2507 is a general-purpose chat model strong across reasoning, math, writing, and multilingual tasks. The Coder variant is post-trained specifically for software engineering and agentic coding, so it tends to score higher on code benchmarks and tool-driven repo tasks while the Instruct model is the better all-rounder.

The model supports 262,144 tokens (256K) natively. Using Dual Chunk Attention and the MInference sparse-attention technique, it can be extended to roughly 1 million tokens with the provided config_1m.json, though that setup demands about 240 GB of GPU memory. On the RULER long-context benchmark it holds up well past 256K, scoring far better than the original Qwen3-30B-A3B at long ranges.