DeepSeek-V3.2

Updated
24.08.2026
Thinking
Tools
Reasoning
Code
Multilingual

Run DeepSeek-V3.2, a 685B MoE model, locally in Atomic Chat. Private, offline reasoning and coding, no API keys, no cloud, no limits. Free.

At a glance

  • License: MIT
  • Parameters: 685.4B
  • Context length: Not stated by DeepSeek; DSA is described as optimized for long context
  • Modalities: Text
  • Minimum hardware: ~192 GB memory (smallest GGUF build is 161.28 GB)

What is DeepSeek-V3.2?

DeepSeek-V3.2 is a 685.4B parameter open-weights model from DeepSeek-AI, published on Hugging Face on December 1, 2025 under the MIT license. DeepSeek builds it on three things: DeepSeek Sparse Attention (DSA), an attention mechanism that cuts computational complexity for long-context work, a scaled reinforcement learning framework for post-training, and a synthesis pipeline that generates agentic training data at scale. The weights are free to download and MIT licensed, but this is a server-class model: the smallest quantized build is 161 GB, so plan for a multi-GPU box or a very large unified-memory machine.

SpecificationDeepSeek-V3.2
Total parameters685.4B
ArchitectureSame model structure as DeepSeek-V3.2-Exp, with DeepSeek Sparse Attention
Original precisionFP8 (F8_E4M3) weights
ModalitiesText
ReasoningThinking mode, including thinking with tools
Recommended samplingtemperature 1.0, top_p 0.95
Release dateDecember 1, 2025
LicenseMIT

The chat template changed a lot in this release. Tool calling has a revised format, the model can think while it calls tools, and a new developer role exists solely for search agent scenarios. There is no Jinja template in the repo; DeepSeek ships Python encoding scripts and test cases instead, so check your serving stack handles the new format before you rely on tool calls.

What DeepSeek-V3.2 is good at

DeepSeek publishes no benchmark numbers in text on this model card, only a benchmark image, so what follows are the vendor's own claims. DeepSeek reports that scaling reinforcement learning post-training brings V3.2 to performance comparable to GPT-5, and that the high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 with reasoning on par with Gemini-3.0-Pro. The concrete evidence offered: gold-medal performance at the 2025 International Mathematical Olympiad and the International Olympiad in Informatics, with the model's selected submissions for IMO 2025, IOI 2025, the ICPC World Finals and CMO 2025 published in the repo for the community to verify.

The other focus is agents. The synthesis pipeline exists to push reasoning into tool-use scenarios, and DeepSeek says it improves compliance and generalization in complex interactive environments. Two practical notes from the card: the Speciale variant is built exclusively for deep reasoning and does not support tool calling, and for local deployment DeepSeek recommends temperature 1.0 with top_p 0.95.

DeepSeek-V3.2 hardware requirements

The system requirement to check is memory. DeepSeek ships the original FP8 weights; the GGUF builds below come from unsloth/DeepSeek-V3.2-GGUF. UD-TQ1_0 ships as one single file; the other four are split into shards, so the size below is the total across a build's files:

MemoryBuild to pickFile size
192 GBUD-TQ1_0161.28 GB
256 GBUD-IQ2_M228.11 GB
320 GBUD-IQ3_XXS273.12 GB
384 GBIQ4_XS358.37 GB
512 GB and upUD-Q4_K_XL407.81 GB

When two builds both fit, take the larger one, and leave headroom above the file size for the KV cache and the rest of the system. If the GGUF format is new to you, start with what GGUF is.

How to run DeepSeek-V3.2 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for DeepSeek-V3.2 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

If your machine is smaller than the table above, see every DeepSeek model you can run locally, including DeepSeek-R1.

DeepSeek-V3.2 license

DeepSeek-V3.2 is released under the MIT license, and the card states that it covers both the repository and the model weights. That permits commercial use, modification and redistribution with no royalties, so you can build products on the model and serve it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download deepseek-ai/DeepSeek-V3.2
from transformers import AutoModel
model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-V3.2")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

DeepSeek-V3.2 is a 685.4B-parameter open-weight language model from deepseek-ai. It uses a Mixture-of-Experts design with DeepSeek Sparse Attention for efficient long-context work, and supports a thinking mode, tool calls, code, and multilingual input across a 128K-token window.

It is a large model, so plan for serious memory. Full FP16 weights run into the hundreds of gigabytes, and even 4-bit quantization needs roughly 386 GB, which points to multi-GPU or high-memory workstation setups rather than a single consumer card. The 128K context window adds memory pressure as the conversation grows.

Yes. The weights are published on Hugging Face under the MIT license at no cost. You can download and run it without an API key or per-token fees, and the license allows both personal and commercial use.

Yes. Once the weights are downloaded, DeepSeek-V3.2 runs fully on your own hardware with no internet connection required. In Atomic Chat your prompts stay on the device, which suits private or data-sensitive work.

Download the weights with huggingface-cli download deepseek-ai/DeepSeek-V3.2, then serve them with vLLM or Transformers, or open the model in Atomic Chat with one click. It is a strong fit for step-by-step reasoning, tool-using agents, and coding tasks, and it handles multilingual prompts.