Qwen3-Coder-30B-A3B-Instruct

Updated
05.10.2026
Tools
Reasoning
Code
Multilingual

Qwen3-Coder-30B-A3B-Instruct is Qwen’s agentic coding MoE: 30.5B parameters, 3.3B active per token, 262K native context. Apache 2.0.

At a glance

  • License: Apache 2.0
  • Parameters: 30.5B total, 3.3B active per token
  • Context length: 262,144 tokens native, up to 1M with Yarn
  • Modalities: Text
  • Minimum hardware: 12 GB of memory (Q2_K GGUF, 11.26 GB)

What is Qwen3-Coder-30B-A3B-Instruct?

Qwen3-Coder-30B-A3B-Instruct is the streamlined model in the Qwen3-Coder family from the Qwen team, built for agentic coding. Qwen3-Coder is available in multiple sizes, and Qwen says this one maintains performance and efficiency. It uses a Mixture of Experts design: 30.5B parameters in total and 3.3B activated per token, across 48 layers. For local use the number that decides whether you can run it is the file size: the 4-bit Q4_K_M build is 18.56 GB, which puts it in the 24 GB memory tier of the table below, and the license is Apache 2.0.

SpecificationQwen3-Coder-30B-A3B-Instruct
TypeCausal language model
Training stagePretraining and post-training
Total parameters30.5B
Active parameters3.3B per token
ArchitectureMixture of Experts, 128 experts, 8 activated
Layers48
AttentionGrouped Query Attention, 32 heads for Q, 4 for KV
Context window262,144 tokens native, up to 1M with Yarn
ModalitiesText
ReasoningNon-thinking only, no think blocks in output
Release dateJuly 31, 2025
LicenseApache 2.0

The A3B in the name is the active parameter count. Each MoE layer holds 128 experts and 8 of them are activated per token, which is where the 3.3B figure comes from. The model runs in non-thinking mode only: it answers directly and does not generate think blocks in its output, so passing enable_thinking=False is no longer required.

What Qwen3-Coder-30B-A3B-Instruct is good at

Qwen's own model card positions it for agentic work: significant performance among open models on agentic coding, agentic browser-use and other foundational coding tasks. It ships a specially designed function call format for tool calling, and Qwen names Qwen Code and CLINE as supported platforms. The model is meant to sit inside an agent loop, calling the tools you define through an OpenAI-compatible endpoint, not just to complete snippets.

The other headline capability is context. The window is 262,144 tokens natively and extends to 1M tokens with Yarn, which Qwen optimized for repository-scale understanding: whole codebases rather than single files. For best results Qwen recommends temperature 0.7, top_p 0.8, top_k 20, repetition penalty 1.05, and an output budget of 65,536 tokens. For local use, Qwen lists Ollama, LM Studio, MLX-LM, llama.cpp and KTransformers as supported runtimes. If you go through transformers instead, use 4.51.0 or newer: older versions fail on this architecture with KeyError: qwen3_moe.

Qwen3-Coder-30B-A3B-Instruct hardware requirements

The system requirement to check is memory. Qwen publishes the original safetensors weights; the quantized builds below, with their real file sizes, come from the unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF repository on Hugging Face.

MemoryBuild to pickFile size
12 GBQ2_K11.26 GB
16 GBUD-Q3_K_XL13.81 GB
24 GBQ4_K_M18.56 GB
32 GBQ5_K_M21.73 GB
48 GB and upQ8_032.48 GB

That repo goes lower and higher than the table: UD-TQ1_0 sits at 8.01 GB at the bottom end, and the BF16 weights split across two files of 49.66 GB and 11.44 GB at the top. In between sit middle steps such as IQ4_XS at 16.38 GB, Q5_K_S at 21.08 GB and Q6_K at 25.09 GB. Neighbouring builds differ by a gigabyte or two, so when two builds both fit, take the larger one. Leave room for context too: a long repository session needs memory beyond the file size, and Qwen suggests dropping the context to 32,768 tokens if you hit out-of-memory errors. If the format is new to you, start with what GGUF is.

How to run Qwen3-Coder-30B-A3B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3-Coder-30B-A3B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every Qwen model you can run locally, or the general-purpose sibling Qwen3-30B-A3B-Instruct-2507.

Qwen3-Coder-30B-A3B-Instruct license

Qwen3-Coder-30B-A3B-Instruct is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model, fine-tune it, and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download Qwen/Qwen3-Coder-30B-A3B-Instruct
from transformers import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3-Coder-30B-A3B-Instruct")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is a coding-focused large language model from Qwen (Alibaba) built on a Mixture-of-Experts design, with 30.5B total parameters and about 3.3B active per token. It is instruction-tuned for writing code, editing code, and running agentic coding workflows, and it supports a 256K-token context. In Atomic Chat it runs locally on your own hardware.

At a 4-bit quant (Q4_K_M) the model fits in roughly 18-22GB of VRAM, so a 24GB card like an RTX 4090 runs it comfortably. A Q6_K build wants about 27GB, and full FP16 weights need around 67GB. You can also run it on CPU with 32GB of system RAM, where users report 12-15 tokens per second at 4-bit.

Yes. The model is released under the Apache 2.0 license, which is free for personal and commercial use. You can download the weights from Hugging Face at no cost and run them locally in Atomic Chat without an API key or subscription.

Yes. Once you download the weights, the model runs entirely on your machine with no internet connection required. Prompts and code stay on the device, which is why it works well for private repositories and sensitive projects. Atomic Chat handles the download and then loads the model fully on-device.

It is built for agentic coding, so it is strong at multi-step coding tasks, tool calls, and code editing across many programming languages. The 256K context window lets it reason over large codebases rather than single snippets. Note that it runs in non-thinking mode and does not emit separate reasoning blocks in its output.