DeepSeek-R1-0528

Updated
24.08.2026
Thinking
Reasoning
Code
Tools

A 671B-parameter MoE reasoning model from DeepSeek with 37B active params, MIT-licensed, strong at math, code, and long chain-of-thought.

At a glance

  • License: MIT
  • Parameters: 684.5B total, native FP8
  • Context length: Not stated by DeepSeek; evaluations use a 64K token maximum generation length
  • Modalities: Text in, text out (the vendor card documents no other modality)
  • Minimum hardware: 192 GB memory for the 161.6 GB UD-TQ1_0 GGUF

What is DeepSeek-R1-0528?

DeepSeek-R1-0528 is the May 28, 2025 update of DeepSeek's R1 reasoning model. DeepSeek calls it a minor version upgrade, but the post-training round behind it, more compute plus new algorithmic optimization mechanisms, improved its depth of reasoning across mathematics, programming and general logic. DeepSeek states that its overall performance is now approaching that of leading models such as O3 and Gemini 2.5 Pro. The weights ship in native FP8, 684.5B parameters by safetensors count, under the MIT license. The smallest GGUF build is 161.62 GB, so memory is the first thing to check.

SpecificationDeepSeek-R1-0528
Total parameters684.5B (safetensors count)
Weight breakdown680.6B in FP8, 3.9B in BF16, 41.6M in FP32
Native precisionFP8 (F8_E4M3)
TypeReasoning model, long chain-of-thought
ModalitiesText in, text out
Maximum generation length64K tokens (vendor eval setting)
Temperature in DeepSeek's app0.6
Distilled variantDeepSeek-R1-0528-Qwen3-8B
PaperarXiv 2501.12948
Release dateMay 28, 2025
LicenseMIT

The extra accuracy comes from longer thinking. On the AIME test set the previous R1 averaged 12K tokens per question; 0528 averages 23K. DeepSeek also reports a reduced hallucination rate, enhanced support for function calling and a better experience for vibe coding. Two usage changes matter locally: a system prompt is now supported, and you no longer need to start the output with a think tag to force the thinking pattern. DeepSeek also distilled the 0528 chain-of-thought into Qwen3 8B Base. The result, DeepSeek-R1-0528-Qwen3-8B, keeps the Qwen3-8B architecture with the 0528 tokenizer configuration, scores 86.0 on AIME 2024 against 76.0 for Qwen3-8B by DeepSeek's numbers, and is run the same way as Qwen3-8B.

DeepSeek-R1-0528 benchmarks

DeepSeek's published numbers, from the model card, compare 0528 with the original DeepSeek R1. Maximum generation length was 64K tokens, and for the benchmarks that require sampling DeepSeek used temperature 0.6, top-p 0.95 and 16 responses per query to estimate pass@1. The metric is not the same in every row: SWE Verified is scored as problems resolved, Aider-Polyglot as accuracy.

BenchmarkDeepSeek-R1-0528DeepSeek R1
AIME 2025
Competition math
87.570.0
HMMT 2025
Olympiad math
79.441.7
GPQA Diamond
Expert science
81.071.5
LiveCodeBench
Competitive coding
73.363.5
SWE Verified
Software engineering
57.649.2
Aider-Polyglot
Code editing
71.653.3
Humanity's Last Exam
Expert questions
17.78.5

The update wins every row here. HMMT 2025 goes from 41.7 to 79.4, Humanity's Last Exam from 8.5 to 17.7, and Aider-Polyglot from 53.3 to 71.6. The one regression in the vendor's full table is SimpleQA, down from 30.1 to 27.8.

The rest of that table reads: AIME 2024 from 79.8 to 91.4, CNMO 2024 from 78.8 to 86.9, the Codeforces Div1 rating from 1530 to 1930. General knowledge moves less: MMLU-Redux 92.9 to 93.4, MMLU-Pro 84.0 to 85.0, FRAMES 82.5 to 83.0. Tool scores are published for 0528 only, with no R1 column: 37.0 on BFCL v3 MultiTurn, and 53.5 on the airline split of Tau-Bench against 63.9 on retail, with GPT-4.1 playing the user. SWE Verified was run with the Agentless framework, and only text-only prompts were scored on Humanity's Last Exam.

DeepSeek-R1-0528 hardware requirements

The system requirement to check is memory. The GGUF builds come from the unsloth/DeepSeek-R1-0528-GGUF repository on Hugging Face. UD-TQ1_0 is one 161.62 GB file. Every other build is split into parts, so the sizes below are the total across a build's parts.

MemoryBuild to pickFile size
192 GBUD-TQ1_0161.6 GB
256 GBUD-IQ2_M228.5 GB
320 GBUD-Q3_K_XL295.9 GB
384 GBIQ4_XS357.7 GB
512 GBQ4_K_M404.9 GB
640 GBQ6_K551.1 GB
768 GB and upQ8_0713.3 GB

When two builds both fit with room left for context, take the larger one. The table skips rungs the repo has. Between UD-TQ1_0 and UD-IQ2_M it carries UD-IQ1_S at 185.5 GB, UD-IQ1_M at 200.5 GB and UD-IQ2_XXS at 216.8 GB. Above UD-IQ2_M it carries UD-Q2_K_XL at 251.1 GB, UD-IQ3_XXS at 272.9 GB and UD-Q4_K_XL at 384.0 GB. Check the listing for a build closer to the memory you have. If none of these tiers fit your hardware, the distilled DeepSeek-R1-0528-Qwen3-8B is the practical way to get the 0528 chain-of-thought on a desktop. If the GGUF format is new to you, start with what GGUF is.

How to run DeepSeek-R1-0528 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for DeepSeek-R1-0528 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

DeepSeek runs the model at temperature 0.6 in its own web app, behind a two-line system prompt: one line names the assistant as DeepSeek-R1 from DeepSeek, the next gives the current date. DeepSeek publishes prompt templates for file uploads and web search, and its DeepSeek Platform endpoint is OpenAI-compatible. For self-serving it points to the DeepSeek-R1 GitHub repository.

For the rest of the family, see every DeepSeek model you can run locally, or browse the full local model catalogue.

DeepSeek-R1-0528 license

DeepSeek-R1-0528 is released under the MIT License, and DeepSeek states that the R1 series, Base and Chat alike, supports commercial use and distillation. That covers building products on the model, redistributing quantized builds, and training smaller models on its outputs, with no usage fee.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download deepseek-ai/DeepSeek-R1-0528
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-R1-0528",
    "messages": [{"role": "user", "content": "Prove 1+2+...+n = n(n+1)/2."}],
    "temperature": 0.6
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-R1-0528", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-R1-0528")
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "deepseek-ai/DeepSeek-R1-0528",
  messages: [{ role: "user", content: "Solve: integral of x^2 dx." }],
  temperature: 0.6,
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

DeepSeek-R1-0528 is a May 28, 2025 update to DeepSeek AI's R1 reasoning model. It uses the DeepSeek-V3 mixture-of-experts architecture with 671B total parameters and 37B active per token. The update deepened the model's chain-of-thought reasoning, raising AIME 2025 accuracy from 70% to 87.5%, and added support for system prompts and function calling.

The full 671B model in its native FP8 weights needs roughly 700 GB of memory, so it targets multi-GPU servers or high-memory hosts. Quantized GGUF builds shrink this considerably: community 1.5-2 bit dynamic quants from Unsloth bring it into the 150-250 GB range, runnable on a CPU with enough system RAM or a workstation with several GPUs. For a single consumer card, the distilled DeepSeek-R1-0528-Qwen3-8B is the practical option.

Yes. The model weights are published on Hugging Face under the MIT License, which permits commercial use, modification, and distillation. You can download and self-host it at no cost, or call it through DeepSeek's OpenAI-compatible API and third-party providers. The MIT terms explicitly allow training other models on its outputs.

DeepSeek-R1-0528 supports a 128K-token context window, inherited from the DeepSeek-V3 architecture it is built on. Because it is a reasoning model that produces long chains of thought, it can spend a large share of that budget on internal thinking. DeepSeek caps maximum generation length at 64K tokens in its own evaluations.

DeepSeek-R1-0528 is a post-training refresh of the same base model, so it keeps the 671B parameter count but thinks harder. On AIME 2025 it jumps from 70% to 87.5%, on LiveCodeBench from 63.5 to 73.3, and its Codeforces rating climbs from 1530 to 1930. The trade-off is longer outputs: it averages about 23K tokens per AIME question versus 12K before. It also lowers hallucination and improves function calling.