What is Qwen3-235B-A22B?
Qwen3-235B-A22B is a mixture-of-experts model from Qwen3, which the Qwen team calls the latest generation of large language models in the Qwen series. It holds 235B total parameters but activates only 22B per token, so each response costs the compute of a 22B model while drawing on the weights of a much larger one. The same checkpoint switches between a thinking mode for complex logical reasoning, math and coding, and a non-thinking mode for efficient general-purpose dialogue. The weights went up on Hugging Face in April 2025 under Apache 2.0.
| Specification | Qwen3-235B-A22B |
|---|---|
| Total parameters | 235B, 22B activated per token |
| Architecture | Mixture-of-experts, 128 experts, 8 active |
| Layers | 94 |
| Attention | GQA, 64 query heads, 4 KV heads |
| Context window | 32,768 tokens native, 131,072 with YaRN |
| Reasoning | Thinking mode on by default, switchable per turn |
| Languages | 100+ languages and dialects |
| Release date | April 2025 |
| License | Apache 2.0 |
The thinking switch works at two levels. A hard enable_thinking flag in the chat template turns reasoning on or off for the whole session, and soft /think and /no_think tags dropped into any message flip it turn by turn. In thinking mode the model reasons inside a <think> block before it answers. Qwen recommends temperature 0.6 with thinking on and 0.7 with it off. For thinking mode it also says not to use greedy decoding, which it warns can degrade performance and produce endless repetitions.
What Qwen3-235B-A22B is good at
The model card ships no benchmark table, so what follows is what Qwen states. In thinking mode the model surpasses the earlier QwQ on mathematics, code generation and commonsense logical reasoning; with thinking off it beats the Qwen2.5 instruct models at the same tasks. Qwen also claims strong human preference alignment: creative writing, role-play, multi-turn dialogue and instruction following.
The other stated strength is agent work. Qwen reports leading performance among open-source models on complex agent tasks, with precise tool calling in both modes, and points to its Qwen-Agent framework, which wires the model to MCP servers and built-in tools through a config file. The model covers 100+ languages and dialects, with multilingual instruction following and translation called out as strengths.
Qwen3-235B-A22B hardware requirements
The system requirement to check is memory: all 235B parameters stay resident even though only 22B compute per token. Qwen publishes official GGUF builds in Qwen/Qwen3-235B-A22B-GGUF; the sizes below are the totals of the split files in that repo.
| Memory | Build to pick | File size |
|---|---|---|
| 160 GB | Q4_K_M | 142.2 GB |
| 192 GB | Q5_K_M | 166.8 GB |
| 256 GB | Q6_K | 193.0 GB |
| 512 GB and up | Q8_0 | 250.0 GB |
When two builds both fit, take the larger one. Q4_K_M is the floor: the repo has nothing smaller, so below roughly 160 GB of memory this checkpoint is out of reach. For a server endpoint Qwen names sglang 0.4.6.post1 or newer and vllm 0.8.5 or newer, and the SGLang command it prints runs with 8-way tensor parallelism; for local use the vendor lists llama.cpp, Ollama, LMStudio, MLX-LM and KTransformers. If the format is new to you, start with what GGUF is.
How to run Qwen3-235B-A22B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3-235B-A22B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
If 142 GB is more than your machine holds, the far smaller MoE sibling Qwen3-30B-A3B is the next stop, and the rest of the lineup is on our page of every Qwen model you can run locally.
Qwen3-235B-A22B license
Qwen3-235B-A22B is released under Apache 2.0, with the license file shipped in the repo. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model and serve it from your own hardware without a usage fee.
