Qwen3-235B-A22B

Updated
05.10.2026
Thinking
Reasoning
Code
Multilingual
Tools

A 235B mixture-of-experts LLM from Alibaba’s Qwen3 series that activates 22B parameters and switches between thinking and non-thinking modes.

At a glance

  • License: Apache 2.0
  • Parameters: 235B total, 22B active (128 experts, 8 active)
  • Context length: 32,768 native, 131,072 with YaRN
  • Modalities: Text in, text out
  • Minimum hardware: 160 GB memory (Q4_K_M GGUF, 142.2 GB)

What is Qwen3-235B-A22B?

Qwen3-235B-A22B is a mixture-of-experts model from Qwen3, which the Qwen team calls the latest generation of large language models in the Qwen series. It holds 235B total parameters but activates only 22B per token, so each response costs the compute of a 22B model while drawing on the weights of a much larger one. The same checkpoint switches between a thinking mode for complex logical reasoning, math and coding, and a non-thinking mode for efficient general-purpose dialogue. The weights went up on Hugging Face in April 2025 under Apache 2.0.

SpecificationQwen3-235B-A22B
Total parameters235B, 22B activated per token
ArchitectureMixture-of-experts, 128 experts, 8 active
Layers94
AttentionGQA, 64 query heads, 4 KV heads
Context window32,768 tokens native, 131,072 with YaRN
ReasoningThinking mode on by default, switchable per turn
Languages100+ languages and dialects
Release dateApril 2025
LicenseApache 2.0

The thinking switch works at two levels. A hard enable_thinking flag in the chat template turns reasoning on or off for the whole session, and soft /think and /no_think tags dropped into any message flip it turn by turn. In thinking mode the model reasons inside a <think> block before it answers. Qwen recommends temperature 0.6 with thinking on and 0.7 with it off. For thinking mode it also says not to use greedy decoding, which it warns can degrade performance and produce endless repetitions.

What Qwen3-235B-A22B is good at

The model card ships no benchmark table, so what follows is what Qwen states. In thinking mode the model surpasses the earlier QwQ on mathematics, code generation and commonsense logical reasoning; with thinking off it beats the Qwen2.5 instruct models at the same tasks. Qwen also claims strong human preference alignment: creative writing, role-play, multi-turn dialogue and instruction following.

The other stated strength is agent work. Qwen reports leading performance among open-source models on complex agent tasks, with precise tool calling in both modes, and points to its Qwen-Agent framework, which wires the model to MCP servers and built-in tools through a config file. The model covers 100+ languages and dialects, with multilingual instruction following and translation called out as strengths.

Qwen3-235B-A22B hardware requirements

The system requirement to check is memory: all 235B parameters stay resident even though only 22B compute per token. Qwen publishes official GGUF builds in Qwen/Qwen3-235B-A22B-GGUF; the sizes below are the totals of the split files in that repo.

MemoryBuild to pickFile size
160 GBQ4_K_M142.2 GB
192 GBQ5_K_M166.8 GB
256 GBQ6_K193.0 GB
512 GB and upQ8_0250.0 GB

When two builds both fit, take the larger one. Q4_K_M is the floor: the repo has nothing smaller, so below roughly 160 GB of memory this checkpoint is out of reach. For a server endpoint Qwen names sglang 0.4.6.post1 or newer and vllm 0.8.5 or newer, and the SGLang command it prints runs with 8-way tensor parallelism; for local use the vendor lists llama.cpp, Ollama, LMStudio, MLX-LM and KTransformers. If the format is new to you, start with what GGUF is.

How to run Qwen3-235B-A22B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3-235B-A22B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

If 142 GB is more than your machine holds, the far smaller MoE sibling Qwen3-30B-A3B is the next stop, and the rest of the lineup is on our page of every Qwen model you can run locally.

Qwen3-235B-A22B license

Qwen3-235B-A22B is released under Apache 2.0, with the license file shipped in the repo. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model and serve it from your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U "transformers>=4.51.0"
huggingface-cli download Qwen/Qwen3-235B-A22B
# serve with vLLM (8-way tensor parallel)
vllm serve Qwen/Qwen3-235B-A22B --enable-reasoning --reasoning-parser deepseek_r1 --tensor-parallel-size 8
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-235B-A22B",
    "messages": [{"role": "user", "content": "Explain mixture-of-experts in one paragraph."}],
    "temperature": 0.6,
    "top_p": 0.95
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen3-235B-A22B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Give me a short intro to LLMs."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=32768)
print(tokenizer.decode(out[0][len(inputs.input_ids[0]):], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "EMPTY" });
const res = await client.chat.completions.create({
  model: "Qwen/Qwen3-235B-A22B",
  messages: [{ role: "user", content: "Explain mixture-of-experts in one paragraph." }],
  temperature: 0.6,
  top_p: 0.95,
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen3-235B-A22B is a mixture-of-experts (MoE) large language model from Alibaba's Qwen team. It has 235 billion total parameters but activates only 22 billion per token, using 8 of its 128 experts on each forward pass. It is the flagship model of the Qwen3 series, released in 2025, and supports both a thinking mode for reasoning and a non-thinking mode for general dialogue.

The official Q4_K_M GGUF listed here totals about 142.2 GB across its shards. The 160 GB tier in the table is a planning estimate for weights plus runtime headroom, not a guarantee at long context. A GPU with 48 GB of VRAM cannot hold all of these weights; CPU offloading requires enough system RAM for the remaining weights and working memory. Requirements depend on the quantization, runtime and context length.

Yes. Qwen3-235B-A22B is released under the Apache 2.0 license, which permits free use, modification, and commercial deployment with no fees. The weights are available to download on Hugging Face. You can run it yourself on your own hardware or access it through hosted API providers such as OpenRouter and Alibaba Model Studio, which charge for their compute.

Qwen3-235B-A22B natively supports a context length of 32,768 tokens. Using the YaRN RoPE-scaling method, it can be extended to 131,072 tokens (128K), which Qwen has validated for long-document and long-conversation use. YaRN is supported by transformers, llama.cpp, vLLM, and SGLang, but the Qwen team recommends enabling it only when you actually need long contexts, since static YaRN can slightly degrade performance on shorter inputs.

Qwen3-235B-A22B can switch between two modes inside a single model. In thinking mode, enabled by default, it produces a reasoning trace wrapped in a <think>...</think> block before its final answer, which improves results on math, code, and logic. In non-thinking mode it replies directly without that block, behaving like a standard instruct model for faster general chat. You can toggle modes with the enable_thinking flag or by adding /think and /no_think to your prompt.