DeepSeek-V4-Pro

Updated
05.10.2026
Thinking
Reasoning
Code
Multilingual

DeepSeek-V4-Pro is a 1.6T-parameter MoE model, 49B active per token, with a 1M-token context under MIT. Run it locally, free, with Atomic Chat.

At a glance

  • License: MIT
  • Parameters: 1.6T total, 49B activated per token
  • Context length: 1M tokens
  • Modalities: Text in, text out
  • Minimum hardware: ~850 GB of memory for the smallest GGUF build

What is DeepSeek-V4-Pro?

DeepSeek-V4-Pro is a 1.6T-parameter Mixture-of-Experts model from DeepSeek, released as a preview of the V4 series alongside the smaller DeepSeek-V4-Flash. It activates 49B parameters per token and reads up to one million tokens of context. What makes it matter for local deployment is efficiency: a new hybrid attention design cuts long-context compute and memory far enough that the 1M window is actually usable, and the MIT license lets you run it anywhere. The weights landed on Hugging Face on April 22, 2026.

SpecificationDeepSeek-V4-Pro
Total parameters1.6T
Activated parameters49B per token
ArchitectureMoE, hybrid attention (Compressed Sparse Attention + Heavily Compressed Attention)
Context window1M tokens
ModalitiesText in, text out
PrecisionFP4 + FP8 mixed: MoE experts in FP4, most other weights in FP8
Reasoning modesNon-think, Think High, Think Max
Pretraining dataMore than 32T tokens
Release dateApril 22, 2026
LicenseMIT

The hybrid attention pairs Compressed Sparse Attention with Heavily Compressed Attention, and the payoff shows up at long context: at the full 1M-token setting, DeepSeek reports the model needs only 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2. Training adds Manifold-Constrained Hyper-Connections on top of standard residual connections, plus the Muon optimizer for faster convergence. Post-training runs in two stages: domain experts are trained separately with SFT and GRPO-based RL, then merged into a single model through on-policy distillation.

DeepSeek-V4-Pro benchmarks

DeepSeek's model card compares the Pro at its highest reasoning effort, DeepSeek-V4-Pro-Max, with Opus 4.6 Max, GPT-5.4 xHigh, Gemini 3.1 Pro High and K2.6 Thinking:

BenchmarkDeepSeek-V4-Pro MaxOpus 4.6 MaxGPT-5.4 xHighGemini 3.1 Pro HighK2.6 Thinking
LiveCodeBench
Competitive coding
93.588.8-91.789.6
Codeforces
Contest rating
3206-31683052-
SWE-bench Verified
Software engineering
80.680.8-80.680.2
Terminal Bench 2.0
Terminal agents
67.965.475.168.566.7
GPQA Diamond
Expert science
90.191.393.094.390.5
Humanity's Last Exam
Expert questions
37.740.039.844.436.4
MRCR 1M
Long context
83.592.9-76.3-

The Pro takes both competitive-coding rows outright and sits 0.2 points behind Opus 4.6 Max on SWE-bench Verified. Gemini 3.1 Pro still leads the knowledge benchmarks and Opus holds the 1M-token retrieval rows, but this is an MIT-licensed model within a few points of closed frontier models almost everywhere.

DeepSeek-V4-Pro hardware requirements

The system requirement to check is memory, and at this scale that means server memory: the GGUF builds in unsloth/DeepSeek-V4-Pro-0813-GGUF ship as 20 split files each, and even the 4-bit build totals 849.7 GB.

MemoryBuild to pickFile size
896 GBUD-Q4_K_XL849.7 GB
1 TB and upUD-Q8_K_XL873.4 GB

The two builds sit only about 24 GB apart, so when both fit, take UD-Q8_K_XL. Leave headroom above the file size for the KV cache: DeepSeek recommends a context window of at least 384K tokens for the Think Max mode, and for local serving suggests temperature 1.0 and top_p 1.0. If split files and quant names are new to you, start with what GGUF is.

How to run DeepSeek-V4-Pro in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for DeepSeek-V4-Pro in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

The Pro is the top of the family: see every DeepSeek model you can run locally, or the much smaller sibling DeepSeek-V4-Flash, which keeps the 1M context at 284B total and 13B activated parameters.

DeepSeek-V4-Pro license

DeepSeek-V4-Pro is released under the MIT license, covering both the repository and the model weights. MIT permits commercial use, modification and redistribution with no royalties, so you can fine-tune the model, build products on it and serve it from your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download deepseek-ai/DeepSeek-V4-Pro
from transformers import AutoModel
model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-V4-Pro")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

DeepSeek-V4-Pro is DeepSeek's open-weight MoE model with 1.6T total parameters and 49B activated per token. It supports a context window of up to one million tokens. The released weights use mixed FP4 and FP8 precision: FP4 for MoE experts and FP8 for most other parameters.

This model requires a server-class memory budget. Capacity depends on the checkpoint and quantization, with additional memory needed for the KV cache and runtime. Check long-context requirements separately before attempting the full one-million-token window. The official mixed FP4 and FP8 checkpoint should not be described or sized as an entirely FP8 model.

Yes. The weights are released under the MIT license, which is one of the most permissive open-source licenses. You can download, run, fine-tune, and redistribute the model, including in commercial products, with no royalty as long as the license notice is kept.

Yes. Once you download the weights from Hugging Face, the model runs entirely on your own machine and needs no internet connection to generate responses. In Atomic Chat that means your prompts and outputs stay on-device.

Download the weights with huggingface-cli download deepseek-ai/DeepSeek-V4-Pro, then serve them with vLLM (which handles MoE expert parallelism) or load them in Transformers; in Atomic Chat it's a one-click model. It's strongest at step-by-step reasoning, code generation and review across large repositories thanks to its 1M-token context, and multilingual text work.