What is Qwen3.6-35B-A3B?
Qwen3.6-35B-A3B is a mixture-of-experts vision-language model from the Qwen team, and the first open-weight variant of Qwen3.6. It carries 35B parameters in total but activates only 3B per token, so it needs the memory of a 35B model while generating each token at the cost of a 3B one. The release is built around agentic coding: frontend workflows, repository-level reasoning, and a new option to keep reasoning context across turns. Qwen published the weights on April 15, 2026 under Apache 2.0.
| Specification | Qwen3.6-35B-A3B |
|---|---|
| Total parameters | 35B, 3B activated per token |
| Architecture | MoE with hybrid attention (Gated DeltaNet + Gated Attention) |
| Experts | 256, with 8 routed + 1 shared active |
| Layers | 40 |
| Context window | 262,144 tokens natively, extensible up to 1,010,000 |
| Modalities | Text, image and video input |
| Reasoning | Thinking on by default, can be switched off via API parameters |
| Multi-Token Prediction | Trained with multi-steps |
| Release date | April 15, 2026 |
| License | Apache 2.0 |
The hidden layout repeats ten times: three Gated DeltaNet blocks each feeding an MoE block, then one Gated Attention block feeding an MoE block. Three of every four layers therefore use Gated DeltaNet, and the Gated Attention layers run 16 query heads against 2 key/value heads. Every layer routes through the 256-expert pool, and just 8 routed experts plus 1 shared expert fire per token. The model is also trained for Multi-Token Prediction, the mechanism behind speculative decoding, so serving stacks that support it can draft several tokens per forward pass. New in 3.6 is thinking preservation: with the preserve_thinking option enabled the model keeps reasoning traces from earlier messages, which Qwen says improves decision consistency in agent loops and cuts redundant reasoning tokens.
Qwen3.6-35B-A3B benchmarks
Qwen's launch numbers, from the model card, compare the 35B-A3B with Qwen3.5-27B, its predecessor Qwen3.5-35B-A3B, Gemma4-31B and Gemma4-26B-A4B:
| Benchmark | Qwen3.6-35B-A3B | Qwen3.5-27B | Qwen3.5-35B-A3B | Gemma4-31B | Gemma4-26B-A4B |
|---|---|---|---|---|---|
SWE-bench Verified Software engineering | 73.4 | 75.0 | 70.0 | 52.0 | 17.4 |
SWE-bench Pro Harder engineering | 49.5 | 51.2 | 44.6 | 35.7 | 13.8 |
Terminal-Bench 2.0 Terminal agents | 51.5 | 41.6 | 40.5 | 42.9 | 34.2 |
Claw-Eval Avg Agentic coding | 68.7 | 64.3 | 65.4 | 48.5 | 58.8 |
NL2Repo Repo-level coding | 29.4 | 27.3 | 20.5 | 15.5 | 11.6 |
GPQA Expert science | 86.0 | 85.5 | 84.2 | 84.3 | 82.3 |
AIME26 Competition math | 92.7 | 92.6 | 91.0 | 89.2 | 88.3 |
The new model takes five of the seven rows shown: Terminal-Bench 2.0, Claw-Eval Avg, NL2Repo, GPQA and AIME26. On the two SWE-bench rows Qwen3.5-27B scores higher: 75.0 against 73.4 on Verified, and 51.2 against 49.5 on Pro. Elsewhere in Qwen's table Qwen3.5-27B also posts 69.3 on SWE-bench Multilingual against 67.2, 31.5 on Tool Decathlon against 26.9, 68.4 on MCP-Atlas against 62.8, and 66.4 on WideSearch against 60.1. The 35B-A3B leads on QwenWebBench with 1397 points against 1068, on MCPMark with 37.0 against 36.3, and on SkillsBench Avg5 with 28.7 against 27.2. Qwen also publishes a vision table: there the 35B-A3B reaches 85.3 on RealWorldQA, 89.9 on OmniDocBench1.5, 83.7 on VideoMMMU and 50.8 on ODInW13, against 84.1, 89.3, 80.4 and 42.6 for Qwen3.5-35B-A3B.
Qwen3.6-35B-A3B hardware requirements
The system requirement to check is memory. The GGUF builds come from the repo unsloth/Qwen3.6-35B-A3B-GGUF, and the sizes below are the actual files listed there.
| Memory | Build to pick | File size |
|---|---|---|
| 12 GB | UD-IQ2_M | 11.52 GB |
| 16 GB | UD-IQ3_S | 13.68 GB |
| 24 GB | UD-IQ4_XS | 17.73 GB |
| 32 GB | UD-Q4_K_XL | 22.36 GB |
| 48 GB and up | UD-Q6_K_XL | 31.84 GB |
Neighbouring files differ by a gigabyte or two, so when two builds both fit, take the larger one. Image and video input need the mmproj file from the same repo, about 0.9 GB on top of the weights. If the GGUF format is new to you, start with what GGUF is.
How to run Qwen3.6-35B-A3B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3.6-35B-A3B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the lineup, see every Qwen model you can run locally.
Qwen3.6-35B-A3B license
Qwen3.6-35B-A3B is released under Apache 2.0, with the license file included in the repo. That permits commercial use, modification, and redistribution with no royalties, so you can fine-tune the weights, ship them inside your own products, and run the model on your own hardware without a usage fee.
