What is DeepSeek-R1?
DeepSeek-R1 is a 671B parameter Mixture-of-Experts reasoning model from DeepSeek, published on January 20, 2025 under the MIT license. It activates 37B parameters per token, carries a 128K context window, and is trained on top of DeepSeek-V3-Base with large-scale reinforcement learning. DeepSeek reports performance comparable to OpenAI o1 across math, code and reasoning tasks, and the weights are open under MIT, so you can run the model, modify it, and distill it into your own models without a usage fee.
| Specification | DeepSeek-R1 |
|---|---|
| Total parameters | 671B |
| Activated parameters | 37B per token |
| Architecture | Mixture-of-Experts, trained on DeepSeek-V3-Base |
| Context window | 128K tokens |
| Training | Large-scale reinforcement learning: two RL stages plus two SFT stages |
| Distilled variants | 1.5B, 7B, 8B, 14B, 32B and 70B, based on Qwen2.5 and Llama3 |
| Release date | January 20, 2025 |
| License | MIT |
The training run is the notable part. DeepSeek first built R1-Zero with reinforcement learning alone, no supervised fine-tuning step, and the model developed self-verification, reflection and long chains of thought on its own. It also picked up endless repetition and language mixing, so the released R1 adds cold-start data before RL. Alongside the big model, DeepSeek open-sourced six dense distills from 1.5B to 70B; the 32B Qwen-based distill beats OpenAI o1-mini across various benchmarks. Practical notes from the model card: keep temperature between 0.5 and 0.7, put all instructions in the user prompt instead of a system prompt, and force the response to open with a think tag, since the model sometimes skips its reasoning block otherwise.
DeepSeek-R1 benchmarks
DeepSeek's launch numbers, from the model card, put R1 against OpenAI o1-1217, o1-mini, DeepSeek V3 and Claude 3.5 Sonnet:
| Benchmark | DeepSeek-R1 | OpenAI o1-1217 | OpenAI o1-mini | DeepSeek V3 | Claude 3.5 Sonnet |
|---|---|---|---|---|---|
MMLU Academic knowledge | 90.8 | 91.8 | 85.2 | 88.5 | 88.3 |
GPQA Diamond Expert science | 71.5 | 75.7 | 60.0 | 59.1 | 65.0 |
AIME 2024 Competition math | 79.8 | 79.2 | 63.6 | 39.2 | 16.0 |
MATH-500 Math problems | 97.3 | 96.4 | 90.0 | 90.2 | 78.3 |
LiveCodeBench Competitive coding | 65.9 | 63.4 | 53.8 | - | 33.8 |
Codeforces Contest percentile | 96.3 | 96.6 | 93.4 | 58.7 | 20.3 |
SWE-bench Verified Software engineering | 49.2 | 48.9 | 41.6 | 42.0 | 50.8 |
R1 takes the math rows and LiveCodeBench outright, and sits within a fraction of a point of o1-1217 on the Codeforces percentile. o1-1217 keeps the lead on GPQA Diamond and MMLU, and Claude 3.5 Sonnet stays ahead on SWE-bench Verified.
DeepSeek-R1 hardware requirements
The system requirement to check is memory, and at 671B parameters the bar is high: even the 1-bit builds are over 140 GB. The GGUF builds below come from unsloth/DeepSeek-R1-GGUF, with file sizes summed across the split parts.
| Memory | Build to pick | File size |
|---|---|---|
| 160 GB | UD-IQ1_S | 140.2 GB |
| 192 GB | UD-IQ1_M | 168.9 GB |
| 256 GB | UD-Q2_K_XL | 226.6 GB |
| 384 GB | Q3_K_M | 319.2 GB |
| 512 GB | Q4_K_M | 404.4 GB |
| 768 GB and up | Q8_0 | 713.3 GB |
When two builds both fit, take the larger one, and leave headroom above the file size for the context cache. If none of these tiers matches hardware you have, DeepSeek's distilled 1.5B to 70B models are the realistic local route. If the GGUF format is new to you, start with what GGUF is.
How to run DeepSeek-R1 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for DeepSeek-R1 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every DeepSeek model you can run locally, or jump to the refreshed checkpoint DeepSeek-R1-0528.
DeepSeek-R1 license
DeepSeek-R1 is released under the MIT license, which covers both the code repository and the model weights. That permits commercial use, modification and derivative works, explicitly including distillation for training other LLMs. The distilled checkpoints inherit their base model terms: the Qwen-based distills are Apache 2.0, and the Llama-based ones carry the Llama 3.1 and Llama 3.3 licenses.
