What is Qwen3.6-27B?
Qwen3.6-27B is a dense 27.8B vision-language model from the Qwen team and the first open-weight release in the Qwen3.6 series. Qwen pitches it as flagship-level coding in a 27B dense model. The weights went up on Hugging Face on April 21, 2026 under Apache 2.0.
| Specification | Qwen3.6-27B |
|---|---|
| Total parameters | 27.8B |
| Architecture | Dense, hybrid attention (Gated DeltaNet + Gated Attention), with a vision encoder |
| Layers | 64 |
| Layer layout | 3 Gated DeltaNet layers then 1 Gated Attention layer, repeated 16 times |
| Hidden dimension | 5,120 |
| Gated Attention heads | 24 Q, 4 KV, head dimension 256, rotary dimension 64 |
| Gated DeltaNet heads | 48 V, 16 QK, head dimension 128 |
| FFN intermediate dimension | 17,408 |
| Token embedding | 248,320, padded |
| Context window | 262,144 tokens natively, extensible to 1,010,000 with YaRN |
| Modalities | Text, image and video input |
| Reasoning | Thinking on by default, can be switched off; earlier turns can be preserved |
| Multi-Token Prediction | Trained with multi-step MTP |
| Training stages | Pre-training and post-training |
| Serving stacks | Transformers, vLLM, SGLang, KTransformers |
| Release date | April 21, 2026 |
| License | Apache 2.0 |
That layout means 48 of the 64 layers use linear attention and only 16 keep a regular KV cache, the part that normally grows with every token. The model is also trained for multi-step Multi-Token Prediction, so it can draft several tokens per forward pass, the same mechanism behind speculative decoding. New in 3.6 is a preserve_thinking option that keeps reasoning traces from earlier messages, which Qwen added for agent loops to cut redundant re-reasoning and improve KV cache reuse.
Qwen3.6-27B benchmarks
Qwen's launch numbers, from the model card, compare the 27B with Qwen3.5-27B, Qwen3.5-397B-A17B, Qwen3.6-35B-A3B, Gemma4-31B and Claude 4.5 Opus:
| Benchmark | Qwen3.6-27B | Qwen3.5-27B | Qwen3.5-397B-A17B | Qwen3.6-35B-A3B | Gemma4-31B | Claude 4.5 Opus |
|---|---|---|---|---|---|---|
SWE-bench Verified Software engineering | 77.2 | 75.0 | 76.2 | 73.4 | 52.0 | 80.9 |
SWE-bench Pro Harder engineering | 53.5 | 51.2 | 50.9 | 49.5 | 35.7 | 57.1 |
Terminal-Bench 2.0 Terminal agents | 59.3 | 41.6 | 52.5 | 51.5 | 42.9 | 59.3 |
SkillsBench Agent skills | 48.2 | 27.2 | 30.0 | 28.7 | 23.6 | 45.3 |
LiveCodeBench v6 Competitive coding | 83.9 | 80.7 | 83.6 | 80.4 | 80.0 | 84.8 |
GPQA Diamond Expert science | 87.8 | 85.5 | 88.4 | 86.0 | 84.3 | 87.0 |
MMLU-Pro Academic knowledge | 86.2 | 86.1 | 87.8 | 85.2 | 85.2 | 89.5 |
Inside its own family the 27B beats Qwen3.5-27B on every row here and outruns Qwen3.5-397B-A17B on the coding-agent rows. Claude 4.5 Opus still leads most of the table, but the 27B matches it on Terminal-Bench 2.0 and passes it on SkillsBench.
Qwen's own framing is narrower than the table: it claims stability and real-world utility, and says the model handles frontend workflows and repository-level reasoning with greater fluency and precision. The card carries numbers for that: QwenWebBench, Qwen's internal front-end generation benchmark, moves from 1068 for the 3.5 27B to 1487, NL2Repo from 27.3 to 36.2, SWE-bench Multilingual from 69.3 to 71.3. On the vision half it posts 82.9 on MMMU, 87.7 on VideoMME with subtitles and 70.3 on AndroidWorld.
Qwen3.6-27B hardware requirements
The system requirement to check is memory. Quantized builds come from the community repo unsloth/Qwen3.6-27B-GGUF, which runs from a 9.39 GB 2-bit file to a 53.8 GB BF16 export split across two parts.
| Memory | Build to pick | File size |
|---|---|---|
| 12 GB | UD-IQ2_XXS | 9.39 GB |
| 16 GB | Q3_K_M | 13.59 GB |
| 20 GB | IQ4_XS | 15.44 GB |
| 24 GB | Q4_K_M | 16.82 GB |
| 32 GB | Q5_K_M | 19.51 GB |
| 48 GB | Q6_K | 22.52 GB |
| 64 GB and up | Q8_0 | 28.60 GB |
Neighbouring builds differ by a gigabyte or two, so when two of them both fit, take the larger one. That matters most at the bottom, where the 2-bit files are what make the model fit a 12 GB card at all and quality falls fastest. Image and video input also needs the mmproj file from the same repo, 0.93 GB in F16, next to the main GGUF. If the format is new to you, start with what GGUF is, then the best local LLMs for a 16 GB Mac.
For the unquantized weights Qwen names Transformers, vLLM, SGLang and KTransformers, and points at SGLang, KTransformers or vLLM for production throughput. It recommends temperature 1.0 with top_p 0.95 and top_k 20 in thinking mode, temperature 0.6 for precise coding, and temperature 0.7 with top_p 0.80 and presence_penalty 1.5 with thinking off, at 32,768 output tokens for most queries. Under memory pressure it advises shrinking the context window but keeping at least 128K tokens.
How to run Qwen3.6-27B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Qwen3.6-27B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Qwen model you can run locally.
Qwen3.6-27B license
Qwen3.6-27B is released under Apache 2.0, with the license file in the Hugging Face repo. That permits commercial use, modification and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.
