What is DeepSeek-V4-Pro?
DeepSeek-V4-Pro is a 1.6T-parameter Mixture-of-Experts model from DeepSeek, released as a preview of the V4 series alongside the smaller DeepSeek-V4-Flash. It activates 49B parameters per token and reads up to one million tokens of context. What makes it matter for local deployment is efficiency: a new hybrid attention design cuts long-context compute and memory far enough that the 1M window is actually usable, and the MIT license lets you run it anywhere. The weights landed on Hugging Face on April 22, 2026.
| Specification | DeepSeek-V4-Pro |
|---|---|
| Total parameters | 1.6T |
| Activated parameters | 49B per token |
| Architecture | MoE, hybrid attention (Compressed Sparse Attention + Heavily Compressed Attention) |
| Context window | 1M tokens |
| Modalities | Text in, text out |
| Precision | FP4 + FP8 mixed: MoE experts in FP4, most other weights in FP8 |
| Reasoning modes | Non-think, Think High, Think Max |
| Pretraining data | More than 32T tokens |
| Release date | April 22, 2026 |
| License | MIT |
The hybrid attention pairs Compressed Sparse Attention with Heavily Compressed Attention, and the payoff shows up at long context: at the full 1M-token setting, DeepSeek reports the model needs only 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2. Training adds Manifold-Constrained Hyper-Connections on top of standard residual connections, plus the Muon optimizer for faster convergence. Post-training runs in two stages: domain experts are trained separately with SFT and GRPO-based RL, then merged into a single model through on-policy distillation.
DeepSeek-V4-Pro benchmarks
DeepSeek's model card compares the Pro at its highest reasoning effort, DeepSeek-V4-Pro-Max, with Opus 4.6 Max, GPT-5.4 xHigh, Gemini 3.1 Pro High and K2.6 Thinking:
| Benchmark | DeepSeek-V4-Pro Max | Opus 4.6 Max | GPT-5.4 xHigh | Gemini 3.1 Pro High | K2.6 Thinking |
|---|---|---|---|---|---|
LiveCodeBench Competitive coding | 93.5 | 88.8 | - | 91.7 | 89.6 |
Codeforces Contest rating | 3206 | - | 3168 | 3052 | - |
SWE-bench Verified Software engineering | 80.6 | 80.8 | - | 80.6 | 80.2 |
Terminal Bench 2.0 Terminal agents | 67.9 | 65.4 | 75.1 | 68.5 | 66.7 |
GPQA Diamond Expert science | 90.1 | 91.3 | 93.0 | 94.3 | 90.5 |
Humanity's Last Exam Expert questions | 37.7 | 40.0 | 39.8 | 44.4 | 36.4 |
MRCR 1M Long context | 83.5 | 92.9 | - | 76.3 | - |
The Pro takes both competitive-coding rows outright and sits 0.2 points behind Opus 4.6 Max on SWE-bench Verified. Gemini 3.1 Pro still leads the knowledge benchmarks and Opus holds the 1M-token retrieval rows, but this is an MIT-licensed model within a few points of closed frontier models almost everywhere.
DeepSeek-V4-Pro hardware requirements
The system requirement to check is memory, and at this scale that means server memory: the GGUF builds in unsloth/DeepSeek-V4-Pro-0813-GGUF ship as 20 split files each, and even the 4-bit build totals 849.7 GB.
| Memory | Build to pick | File size |
|---|---|---|
| 896 GB | UD-Q4_K_XL | 849.7 GB |
| 1 TB and up | UD-Q8_K_XL | 873.4 GB |
The two builds sit only about 24 GB apart, so when both fit, take UD-Q8_K_XL. Leave headroom above the file size for the KV cache: DeepSeek recommends a context window of at least 384K tokens for the Think Max mode, and for local serving suggests temperature 1.0 and top_p 1.0. If split files and quant names are new to you, start with what GGUF is.
How to run DeepSeek-V4-Pro in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for DeepSeek-V4-Pro in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
The Pro is the top of the family: see every DeepSeek model you can run locally, or the much smaller sibling DeepSeek-V4-Flash, which keeps the 1M context at 284B total and 13B activated parameters.
DeepSeek-V4-Pro license
DeepSeek-V4-Pro is released under the MIT license, covering both the repository and the model weights. MIT permits commercial use, modification and redistribution with no royalties, so you can fine-tune the model, build products on it and serve it from your own hardware without a usage fee.
