Qwen3.6-27B

Updated
24.08.2026
Tools
Thinking
Embedding
Vision
Reasoning
Code
Multilingual

Qwen3.6-27B is a dense 27.8B vision-language model with a 262K context and thinking on by default. GGUF builds run from 9.39 GB on a 12 GB machine.

At a glance

  • License: Apache 2.0
  • Parameters: 27.8B, dense
  • Context length: 262,144 tokens, extensible to 1,010,000 with YaRN
  • Modalities: Text, image and video input
  • Minimum hardware: 12 GB of RAM or VRAM, smallest GGUF build is 9.39 GB

What is Qwen3.6-27B?

Qwen3.6-27B is a dense 27.8B vision-language model from the Qwen team and the first open-weight release in the Qwen3.6 series. Qwen pitches it as flagship-level coding in a 27B dense model. The weights went up on Hugging Face on April 21, 2026 under Apache 2.0.

SpecificationQwen3.6-27B
Total parameters27.8B
ArchitectureDense, hybrid attention (Gated DeltaNet + Gated Attention), with a vision encoder
Layers64
Layer layout3 Gated DeltaNet layers then 1 Gated Attention layer, repeated 16 times
Hidden dimension5,120
Gated Attention heads24 Q, 4 KV, head dimension 256, rotary dimension 64
Gated DeltaNet heads48 V, 16 QK, head dimension 128
FFN intermediate dimension17,408
Token embedding248,320, padded
Context window262,144 tokens natively, extensible to 1,010,000 with YaRN
ModalitiesText, image and video input
ReasoningThinking on by default, can be switched off; earlier turns can be preserved
Multi-Token PredictionTrained with multi-step MTP
Training stagesPre-training and post-training
Serving stacksTransformers, vLLM, SGLang, KTransformers
Release dateApril 21, 2026
LicenseApache 2.0

That layout means 48 of the 64 layers use linear attention and only 16 keep a regular KV cache, the part that normally grows with every token. The model is also trained for multi-step Multi-Token Prediction, so it can draft several tokens per forward pass, the same mechanism behind speculative decoding. New in 3.6 is a preserve_thinking option that keeps reasoning traces from earlier messages, which Qwen added for agent loops to cut redundant re-reasoning and improve KV cache reuse.

Qwen3.6-27B benchmarks

Qwen's launch numbers, from the model card, compare the 27B with Qwen3.5-27B, Qwen3.5-397B-A17B, Qwen3.6-35B-A3B, Gemma4-31B and Claude 4.5 Opus:

BenchmarkQwen3.6-27BQwen3.5-27BQwen3.5-397B-A17BQwen3.6-35B-A3BGemma4-31BClaude 4.5 Opus
SWE-bench Verified
Software engineering
77.275.076.273.452.080.9
SWE-bench Pro
Harder engineering
53.551.250.949.535.757.1
Terminal-Bench 2.0
Terminal agents
59.341.652.551.542.959.3
SkillsBench
Agent skills
48.227.230.028.723.645.3
LiveCodeBench v6
Competitive coding
83.980.783.680.480.084.8
GPQA Diamond
Expert science
87.885.588.486.084.387.0
MMLU-Pro
Academic knowledge
86.286.187.885.285.289.5

Inside its own family the 27B beats Qwen3.5-27B on every row here and outruns Qwen3.5-397B-A17B on the coding-agent rows. Claude 4.5 Opus still leads most of the table, but the 27B matches it on Terminal-Bench 2.0 and passes it on SkillsBench.

Qwen's own framing is narrower than the table: it claims stability and real-world utility, and says the model handles frontend workflows and repository-level reasoning with greater fluency and precision. The card carries numbers for that: QwenWebBench, Qwen's internal front-end generation benchmark, moves from 1068 for the 3.5 27B to 1487, NL2Repo from 27.3 to 36.2, SWE-bench Multilingual from 69.3 to 71.3. On the vision half it posts 82.9 on MMMU, 87.7 on VideoMME with subtitles and 70.3 on AndroidWorld.

Qwen3.6-27B hardware requirements

The system requirement to check is memory. Quantized builds come from the community repo unsloth/Qwen3.6-27B-GGUF, which runs from a 9.39 GB 2-bit file to a 53.8 GB BF16 export split across two parts.

MemoryBuild to pickFile size
12 GBUD-IQ2_XXS9.39 GB
16 GBQ3_K_M13.59 GB
20 GBIQ4_XS15.44 GB
24 GBQ4_K_M16.82 GB
32 GBQ5_K_M19.51 GB
48 GBQ6_K22.52 GB
64 GB and upQ8_028.60 GB

Neighbouring builds differ by a gigabyte or two, so when two of them both fit, take the larger one. That matters most at the bottom, where the 2-bit files are what make the model fit a 12 GB card at all and quality falls fastest. Image and video input also needs the mmproj file from the same repo, 0.93 GB in F16, next to the main GGUF. If the format is new to you, start with what GGUF is, then the best local LLMs for a 16 GB Mac.

For the unquantized weights Qwen names Transformers, vLLM, SGLang and KTransformers, and points at SGLang, KTransformers or vLLM for production throughput. It recommends temperature 1.0 with top_p 0.95 and top_k 20 in thinking mode, temperature 0.6 for precise coding, and temperature 0.7 with top_p 0.80 and presence_penalty 1.5 with thinking off, at 32,768 output tokens for most queries. Under memory pressure it advises shrinking the context window but keeping at least 128K tokens.

How to run Qwen3.6-27B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Qwen3.6-27B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Qwen model you can run locally.

Qwen3.6-27B license

Qwen3.6-27B is released under Apache 2.0, with the license file in the Hugging Face repo. That permits commercial use, modification and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download Qwen/Qwen3.6-27B
from transformers import AutoModel
model = AutoModel.from_pretrained("Qwen/Qwen3.6-27B")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Qwen3.6-27B is a 27.8B-parameter dense, multimodal language model from Qwen (Alibaba Cloud). It handles text, images, and video, supports reasoning and tool calling, and carries a 262,144-token context window. The weights are open under Apache-2.0, so it can run locally in Atomic Chat without an API.

The Q4_K_M weights are about 16.8 GB, so they cannot be loaded entirely into 16 GB of VRAM. The table uses a 24 GB memory tier as a starting point, with additional memory needed for the KV cache, runtime and vision projector. Q3_K_M is about 13.6 GB on disk, not a 12 GB runtime requirement. CPU offloading and a shorter context can reduce VRAM use.

Yes. Qwen3.6-27B is published under the Apache-2.0 license, which allows free use including commercial deployment, with no royalties or usage fees. You can download the weights from Hugging Face and run them on your own hardware at no cost.

Yes. Once the weights are downloaded, the model runs entirely on your machine with no network connection. In Atomic Chat your prompts, code, and files never leave your device, which suits private or air-gapped work.

Its strongest area is agentic coding, where Qwen reports 77.2 on SWE-bench Verified, ahead of the larger Qwen3.5-397B-A17B model. It is also natively multimodal, reading images, OCR documents, and long video, and it supports tool calling plus a thinking mode for multi-step agent tasks.