Qwen3.6-35B-A3B

Updated
05.10.2026
Tools
Thinking
Embedding
Vision
Reasoning
Code
Multilingual

Qwen3.6-35B-A3B is Qwen’s open MoE vision model: 35B parameters, 3B active, 262K context, Apache 2.0. Low-bit GGUF builds fit in 12 GB.

At a glance

  • License: Apache 2.0
  • Parameters: 35B total, 3B active per token
  • Context length: 262,144 tokens natively, extensible up to 1,010,000
  • Modalities: Text, image and video input
  • Minimum hardware: 12 GB of memory (UD-IQ2_M GGUF, 11.52 GB)

What is Qwen3.6-35B-A3B?

Qwen3.6-35B-A3B is a mixture-of-experts vision-language model from the Qwen team, and the first open-weight variant of Qwen3.6. It carries 35B parameters in total but activates only 3B per token, so it needs the memory of a 35B model while generating each token at the cost of a 3B one. The release is built around agentic coding: frontend workflows, repository-level reasoning, and a new option to keep reasoning context across turns. Qwen published the weights on April 15, 2026 under Apache 2.0.

SpecificationQwen3.6-35B-A3B
Total parameters35B, 3B activated per token
ArchitectureMoE with hybrid attention (Gated DeltaNet + Gated Attention)
Experts256, with 8 routed + 1 shared active
Layers40
Context window262,144 tokens natively, extensible up to 1,010,000
ModalitiesText, image and video input
ReasoningThinking on by default, can be switched off via API parameters
Multi-Token PredictionTrained with multi-steps
Release dateApril 15, 2026
LicenseApache 2.0

The hidden layout repeats ten times: three Gated DeltaNet blocks each feeding an MoE block, then one Gated Attention block feeding an MoE block. Three of every four layers therefore use Gated DeltaNet, and the Gated Attention layers run 16 query heads against 2 key/value heads. Every layer routes through the 256-expert pool, and just 8 routed experts plus 1 shared expert fire per token. The model is also trained for Multi-Token Prediction, the mechanism behind speculative decoding, so serving stacks that support it can draft several tokens per forward pass. New in 3.6 is thinking preservation: with the preserve_thinking option enabled the model keeps reasoning traces from earlier messages, which Qwen says improves decision consistency in agent loops and cuts redundant reasoning tokens.

Qwen3.6-35B-A3B benchmarks

Qwen's launch numbers, from the model card, compare the 35B-A3B with Qwen3.5-27B, its predecessor Qwen3.5-35B-A3B, Gemma4-31B and Gemma4-26B-A4B:

BenchmarkQwen3.6-35B-A3BQwen3.5-27BQwen3.5-35B-A3BGemma4-31BGemma4-26B-A4B
SWE-bench Verified
Software engineering
73.475.070.052.017.4
SWE-bench Pro
Harder engineering
49.551.244.635.713.8
Terminal-Bench 2.0
Terminal agents
51.541.640.542.934.2
Claw-Eval Avg
Agentic coding
68.764.365.448.558.8
NL2Repo
Repo-level coding
29.427.320.515.511.6
GPQA
Expert science
86.085.584.284.382.3
AIME26
Competition math
92.792.691.089.288.3

The new model takes five of the seven rows shown: Terminal-Bench 2.0, Claw-Eval Avg, NL2Repo, GPQA and AIME26. On the two SWE-bench rows Qwen3.5-27B scores higher: 75.0 against 73.4 on Verified, and 51.2 against 49.5 on Pro. Elsewhere in Qwen's table Qwen3.5-27B also posts 69.3 on SWE-bench Multilingual against 67.2, 31.5 on Tool Decathlon against 26.9, 68.4 on MCP-Atlas against 62.8, and 66.4 on WideSearch against 60.1. The 35B-A3B leads on QwenWebBench with 1397 points against 1068, on MCPMark with 37.0 against 36.3, and on SkillsBench Avg5 with 28.7 against 27.2. Qwen also publishes a vision table: there the 35B-A3B reaches 85.3 on RealWorldQA, 89.9 on OmniDocBench1.5, 83.7 on VideoMMMU and 50.8 on ODInW13, against 84.1, 89.3, 80.4 and 42.6 for Qwen3.5-35B-A3B.

Qwen3.6-35B-A3B hardware requirements

The system requirement to check is memory. The GGUF builds come from the repo unsloth/Qwen3.6-35B-A3B-GGUF, and the sizes below are the actual files listed there.

MemoryBuild to pickFile size
12 GBUD-IQ2_M11.52 GB
16 GBUD-IQ3_S13.68 GB
24 GBUD-IQ4_XS17.73 GB
32 GBUD-Q4_K_XL22.36 GB
48 GB and upUD-Q6_K_XL31.84 GB

Neighbouring files differ by a gigabyte or two, so when two builds both fit, take the larger one. Image and video input need the mmproj file from the same repo, about 0.9 GB on top of the weights. If the GGUF format is new to you, start with what GGUF is.

How to run Qwen3.6-35B-A3B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3.6-35B-A3B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every Qwen model you can run locally.

Qwen3.6-35B-A3B license

Qwen3.6-35B-A3B is released under Apache 2.0, with the license file included in the repo. That permits commercial use, modification, and redistribution with no royalties, so you can fine-tune the weights, ship them inside your own products, and run the model on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download Qwen/Qwen3.6-35B-A3B
from transformers import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3.6-35B-A3B")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is an open-weight Mixture-of-Experts language model from Qwen (Alibaba's AI team), with 36B total parameters and about 3B active per token. It supports text, vision, tool calling, a thinking mode, and a 262,144-token context window, all under the Apache-2.0 license. In Atomic Chat it runs fully on your own hardware.

At Q4_K_M quantization the weights are about 21 GB, so a 24 GB GPU such as an RTX 3090, 4090, or 5090, or a Mac with 32 GB or more unified memory, runs it well. The Q8_0 build is closer to 37 GB and needs 48 GB-class hardware. 16 GB cards are not enough even at Q4 unless you use aggressive offloading.

Yes. The model is released under the Apache-2.0 license, which allows free use, modification, and commercial deployment. Running it locally through Atomic Chat means there are no API fees or per-token costs, only your own hardware and electricity.

Yes. Once the weights are downloaded, the model runs entirely on-device with no internet connection required. Every prompt and response stays on your machine, so nothing is sent to Alibaba or any other server, which is the main reason to run it in Atomic Chat.

The simplest route is Atomic Chat: open the app, find Qwen3.6-35B-A3B in the model list, and click to download and load a quantized build. If you prefer the command line, pull the weights with huggingface-cli download Qwen/Qwen3.6-35B-A3B and serve them through Transformers, vLLM, or llama.cpp.