DeepSeek-R1

Updated
24.08.2026
Thinking
Reasoning
Code
Multilingual

Run DeepSeek-R1, a 671B reasoning model, locally with Atomic Chat. Private, offline chain-of-thought with no cloud and no limits. Free.

At a glance

  • License: MIT
  • Parameters: 671B total, 37B activated per token
  • Context length: 128K tokens
  • Modalities: Text
  • Minimum hardware: 160 GB of memory for the 140.2 GB 1-bit GGUF build

What is DeepSeek-R1?

DeepSeek-R1 is a 671B parameter Mixture-of-Experts reasoning model from DeepSeek, published on January 20, 2025 under the MIT license. It activates 37B parameters per token, carries a 128K context window, and is trained on top of DeepSeek-V3-Base with large-scale reinforcement learning. DeepSeek reports performance comparable to OpenAI o1 across math, code and reasoning tasks, and the weights are open under MIT, so you can run the model, modify it, and distill it into your own models without a usage fee.

SpecificationDeepSeek-R1
Total parameters671B
Activated parameters37B per token
ArchitectureMixture-of-Experts, trained on DeepSeek-V3-Base
Context window128K tokens
TrainingLarge-scale reinforcement learning: two RL stages plus two SFT stages
Distilled variants1.5B, 7B, 8B, 14B, 32B and 70B, based on Qwen2.5 and Llama3
Release dateJanuary 20, 2025
LicenseMIT

The training run is the notable part. DeepSeek first built R1-Zero with reinforcement learning alone, no supervised fine-tuning step, and the model developed self-verification, reflection and long chains of thought on its own. It also picked up endless repetition and language mixing, so the released R1 adds cold-start data before RL. Alongside the big model, DeepSeek open-sourced six dense distills from 1.5B to 70B; the 32B Qwen-based distill beats OpenAI o1-mini across various benchmarks. Practical notes from the model card: keep temperature between 0.5 and 0.7, put all instructions in the user prompt instead of a system prompt, and force the response to open with a think tag, since the model sometimes skips its reasoning block otherwise.

DeepSeek-R1 benchmarks

DeepSeek's launch numbers, from the model card, put R1 against OpenAI o1-1217, o1-mini, DeepSeek V3 and Claude 3.5 Sonnet:

BenchmarkDeepSeek-R1OpenAI o1-1217OpenAI o1-miniDeepSeek V3Claude 3.5 Sonnet
MMLU
Academic knowledge
90.891.885.288.588.3
GPQA Diamond
Expert science
71.575.760.059.165.0
AIME 2024
Competition math
79.879.263.639.216.0
MATH-500
Math problems
97.396.490.090.278.3
LiveCodeBench
Competitive coding
65.963.453.8-33.8
Codeforces
Contest percentile
96.396.693.458.720.3
SWE-bench Verified
Software engineering
49.248.941.642.050.8

R1 takes the math rows and LiveCodeBench outright, and sits within a fraction of a point of o1-1217 on the Codeforces percentile. o1-1217 keeps the lead on GPQA Diamond and MMLU, and Claude 3.5 Sonnet stays ahead on SWE-bench Verified.

DeepSeek-R1 hardware requirements

The system requirement to check is memory, and at 671B parameters the bar is high: even the 1-bit builds are over 140 GB. The GGUF builds below come from unsloth/DeepSeek-R1-GGUF, with file sizes summed across the split parts.

MemoryBuild to pickFile size
160 GBUD-IQ1_S140.2 GB
192 GBUD-IQ1_M168.9 GB
256 GBUD-Q2_K_XL226.6 GB
384 GBQ3_K_M319.2 GB
512 GBQ4_K_M404.4 GB
768 GB and upQ8_0713.3 GB

When two builds both fit, take the larger one, and leave headroom above the file size for the context cache. If none of these tiers matches hardware you have, DeepSeek's distilled 1.5B to 70B models are the realistic local route. If the GGUF format is new to you, start with what GGUF is.

How to run DeepSeek-R1 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for DeepSeek-R1 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every DeepSeek model you can run locally, or jump to the refreshed checkpoint DeepSeek-R1-0528.

DeepSeek-R1 license

DeepSeek-R1 is released under the MIT license, which covers both the code repository and the model weights. That permits commercial use, modification and derivative works, explicitly including distillation for training other LLMs. The distilled checkpoints inherit their base model terms: the Qwen-based distills are Apache 2.0, and the Llama-based ones carry the Llama 3.1 and Llama 3.3 licenses.

Get the weights from Hugging Face

huggingface-cli download deepseek-ai/DeepSeek-R1
from transformers import AutoModel
model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-R1")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

DeepSeek-R1 is an open-weight reasoning model from deepseek-ai, built on a Mixture-of-Experts architecture totaling 684.5B parameters with a 128K-token context window. It was trained with reinforcement learning to produce a step-by-step thinking trace before answering. On math, code, and reasoning benchmarks it performs in the same range as OpenAI's o1.

The full 684.5B MoE model is data-center scale and needs a multi-GPU or server setup, so most people run the distilled versions instead. A 7B distill fits in about 8 GB of VRAM, a 14B needs roughly 8.5 GB, a 32B around 17.5 GB, and the 70B distill about 36 GB. Leave extra headroom, since peak VRAM during long reasoning prompts can run well above the base weight size.

Yes. DeepSeek-R1 is released under the MIT license, so the weights are free to download and run. When you run it locally through Atomic Chat there are no subscription fees and no per-token charges, and the MIT terms also allow commercial use and modification.

Yes. Once you have downloaded the weights, DeepSeek-R1 runs entirely on your own machine with no internet connection required. This keeps your prompts and data on-device, which suits private documents, secure environments, or just working without a network.

Atomic Chat lets you load DeepSeek-R1 with one click, which handles the download and setup for you. If you prefer to do it manually, pull the weights with huggingface-cli download deepseek-ai/DeepSeek-R1 and serve them through vLLM or Transformers. For most desktops, picking a distilled variant that fits your VRAM is the simpler route.