Qwen3-32B

Updated
05.10.2026
Tools
Thinking
Reasoning
Code
Multilingual

Qwen3-32B is a 32.8B-parameter Apache 2.0 model from Alibaba’s Qwen team, with a switchable thinking mode and 100+ languages. Q4_K_M fits a 24 GB card.

At a glance

  • License: Apache 2.0
  • Parameters: 32.8B (31.2B non-embedding)
  • Context length: 32,768 native, 131,072 with YaRN
  • Modalities: Text in, text out
  • Minimum hardware: 24 GB memory (Q4_K_M, 19.76 GB)

What is Qwen3-32B?

Qwen3-32B is a 32.8B-parameter language model from Alibaba's Qwen team, part of the Qwen3 generation of open-weight models. Its defining feature is a switchable thinking mode: the same checkpoint reasons step by step before answering when you want depth on math, code or logic, and skips straight to the answer for fast everyday chat. Alibaba published the weights in April 2025 under Apache 2.0, and the official Q4_K_M build fits on a single 24 GB GPU.

SpecificationQwen3-32B
Total parameters32.8B (31.2B non-embedding)
TypeCausal language model, pretrained and post-trained
Layers64
AttentionGQA, 64 heads for queries, 8 for KV
Context window32,768 tokens native, 131,072 with YaRN
Thinking modeSwitchable per session and per turn
Languages100+ languages and dialects
Release dateApril 2025
LicenseApache 2.0

The thinking switch works at two levels. In the chat template, the enable_thinking flag turns reasoning on or off for the whole session. On top of that, appending /think or /no_think to any message flips the mode for that turn, and the model follows the most recent instruction in a multi-turn conversation. Qwen recommends different sampling per mode, temperature 0.6 with top-p 0.95 while thinking and 0.7 with top-p 0.8 without, and warns against greedy decoding, which can cause endless repetitions. Context is 32,768 tokens natively; Qwen has validated up to 131,072 tokens with YaRN scaling, but advises enabling it only when you actually process long inputs, because static YaRN can degrade quality on short texts.

What Qwen3-32B is good at

Qwen's own card makes four claims for this generation. Reasoning: in thinking mode it surpasses the earlier QwQ, and in non-thinking mode the Qwen2.5-Instruct models, on mathematics, code generation and commonsense logic. Alignment: better human preference in creative writing, role-play, multi-turn dialogue and instruction following. Agents: precise tool calling in both modes, with what Qwen calls leading performance among open-source models on complex agent tasks; the card recommends the Qwen-Agent framework, which ships tool-calling templates and parsers and reads tool definitions straight from an MCP configuration file. Languages: 100+ languages and dialects, with multilingual instruction following and translation named as explicit training targets.

Two practical notes from the card. In multi-turn conversations the history should carry only the final answers, not the thinking content; the official chat template already handles this. And for hard math or coding problems, Qwen suggests giving the model room to think: an output budget of 32,768 tokens for most queries, up to 38,912 for the hardest ones.

Qwen3-32B hardware requirements

The system requirement to check is memory. Qwen publishes official GGUF builds as Qwen/Qwen3-32B-GGUF; these are the real file sizes:

MemoryBuild to pickFile size
24 GBQ4_K_M19.76 GB
32 GBQ5_K_M23.21 GB
48 GBQ6_K26.88 GB
64 GB and upQ8_034.82 GB

When two builds both fit, take the larger one, and leave a few gigabytes of headroom over the file size for the context. The official repo starts at Q4_K_M, so 24 GB is the practical floor here; with less memory, look at the smaller sibling Qwen3-14B. If the GGUF format is new to you, start with what GGUF is.

How to run Qwen3-32B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3-32B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every Qwen model you can run locally, including the lighter sibling Qwen3-30B-A3B.

Qwen3-32B license

Qwen3-32B is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can fine-tune the model, ship it inside a product, and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download Qwen/Qwen3-32B
from transformers import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen3-32B is a 32.8B-parameter dense large language model from Qwen, Alibaba Cloud's model team. It supports a 128K context window and a hybrid thinking mode that runs explicit reasoning for hard problems or answers directly for plain chat. The weights are open under Apache 2.0, so you can download and run it yourself.

A 4-bit quantized build (Q4_K_M) needs roughly 16-19GB of VRAM, so it fits a 24GB GPU like an RTX 3090 or RTX 4090. An 8-bit build needs about 32GB, and full FP16 precision needs around 65GB. Using the full 128K context adds memory on top of the weights, so on a 24GB card you may need to limit context length.

Yes. The model is released as open weights under the Apache 2.0 license, which means free download and use. The license also allows commercial use, modification, and redistribution as long as you keep the license and attribution notices.

Yes. Once the weights finish downloading, Qwen3-32B runs fully on your own hardware with no internet connection required. In Atomic Chat every prompt and response stays on-device, so nothing is sent to an external server.

It is strong at reasoning and math thanks to its thinking mode, which produces a chain of thought before the final answer. It also handles code generation and debugging, supports tool and function calling for agent workflows, and works across 100+ languages. The 128K context lets you feed in long documents or large codebases at once.