What is Qwen3-Coder-30B-A3B-Instruct?
Qwen3-Coder-30B-A3B-Instruct is the streamlined model in the Qwen3-Coder family from the Qwen team, built for agentic coding. Qwen3-Coder is available in multiple sizes, and Qwen says this one maintains performance and efficiency. It uses a Mixture of Experts design: 30.5B parameters in total and 3.3B activated per token, across 48 layers. For local use the number that decides whether you can run it is the file size: the 4-bit Q4_K_M build is 18.56 GB, which puts it in the 24 GB memory tier of the table below, and the license is Apache 2.0.
| Specification | Qwen3-Coder-30B-A3B-Instruct |
|---|---|
| Type | Causal language model |
| Training stage | Pretraining and post-training |
| Total parameters | 30.5B |
| Active parameters | 3.3B per token |
| Architecture | Mixture of Experts, 128 experts, 8 activated |
| Layers | 48 |
| Attention | Grouped Query Attention, 32 heads for Q, 4 for KV |
| Context window | 262,144 tokens native, up to 1M with Yarn |
| Modalities | Text |
| Reasoning | Non-thinking only, no think blocks in output |
| Release date | July 31, 2025 |
| License | Apache 2.0 |
The A3B in the name is the active parameter count. Each MoE layer holds 128 experts and 8 of them are activated per token, which is where the 3.3B figure comes from. The model runs in non-thinking mode only: it answers directly and does not generate think blocks in its output, so passing enable_thinking=False is no longer required.
What Qwen3-Coder-30B-A3B-Instruct is good at
Qwen's own model card positions it for agentic work: significant performance among open models on agentic coding, agentic browser-use and other foundational coding tasks. It ships a specially designed function call format for tool calling, and Qwen names Qwen Code and CLINE as supported platforms. The model is meant to sit inside an agent loop, calling the tools you define through an OpenAI-compatible endpoint, not just to complete snippets.
The other headline capability is context. The window is 262,144 tokens natively and extends to 1M tokens with Yarn, which Qwen optimized for repository-scale understanding: whole codebases rather than single files. For best results Qwen recommends temperature 0.7, top_p 0.8, top_k 20, repetition penalty 1.05, and an output budget of 65,536 tokens. For local use, Qwen lists Ollama, LM Studio, MLX-LM, llama.cpp and KTransformers as supported runtimes. If you go through transformers instead, use 4.51.0 or newer: older versions fail on this architecture with KeyError: qwen3_moe.
Qwen3-Coder-30B-A3B-Instruct hardware requirements
The system requirement to check is memory. Qwen publishes the original safetensors weights; the quantized builds below, with their real file sizes, come from the unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF repository on Hugging Face.
| Memory | Build to pick | File size |
|---|---|---|
| 12 GB | Q2_K | 11.26 GB |
| 16 GB | UD-Q3_K_XL | 13.81 GB |
| 24 GB | Q4_K_M | 18.56 GB |
| 32 GB | Q5_K_M | 21.73 GB |
| 48 GB and up | Q8_0 | 32.48 GB |
That repo goes lower and higher than the table: UD-TQ1_0 sits at 8.01 GB at the bottom end, and the BF16 weights split across two files of 49.66 GB and 11.44 GB at the top. In between sit middle steps such as IQ4_XS at 16.38 GB, Q5_K_S at 21.08 GB and Q6_K at 25.09 GB. Neighbouring builds differ by a gigabyte or two, so when two builds both fit, take the larger one. Leave room for context too: a long repository session needs memory beyond the file size, and Qwen suggests dropping the context to 32,768 tokens if you hit out-of-memory errors. If the format is new to you, start with what GGUF is.
How to run Qwen3-Coder-30B-A3B-Instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3-Coder-30B-A3B-Instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the lineup, see every Qwen model you can run locally, or the general-purpose sibling Qwen3-30B-A3B-Instruct-2507.
Qwen3-Coder-30B-A3B-Instruct license
Qwen3-Coder-30B-A3B-Instruct is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model, fine-tune it, and run it on your own hardware without a usage fee.
