What is DeepSeek-R1-0528?
DeepSeek-R1-0528 is the May 28, 2025 update of DeepSeek's R1 reasoning model. DeepSeek calls it a minor version upgrade, but the post-training round behind it, more compute plus new algorithmic optimization mechanisms, improved its depth of reasoning across mathematics, programming and general logic. DeepSeek states that its overall performance is now approaching that of leading models such as O3 and Gemini 2.5 Pro. The weights ship in native FP8, 684.5B parameters by safetensors count, under the MIT license. The smallest GGUF build is 161.62 GB, so memory is the first thing to check.
| Specification | DeepSeek-R1-0528 |
|---|---|
| Total parameters | 684.5B (safetensors count) |
| Weight breakdown | 680.6B in FP8, 3.9B in BF16, 41.6M in FP32 |
| Native precision | FP8 (F8_E4M3) |
| Type | Reasoning model, long chain-of-thought |
| Modalities | Text in, text out |
| Maximum generation length | 64K tokens (vendor eval setting) |
| Temperature in DeepSeek's app | 0.6 |
| Distilled variant | DeepSeek-R1-0528-Qwen3-8B |
| Paper | arXiv 2501.12948 |
| Release date | May 28, 2025 |
| License | MIT |
The extra accuracy comes from longer thinking. On the AIME test set the previous R1 averaged 12K tokens per question; 0528 averages 23K. DeepSeek also reports a reduced hallucination rate, enhanced support for function calling and a better experience for vibe coding. Two usage changes matter locally: a system prompt is now supported, and you no longer need to start the output with a think tag to force the thinking pattern. DeepSeek also distilled the 0528 chain-of-thought into Qwen3 8B Base. The result, DeepSeek-R1-0528-Qwen3-8B, keeps the Qwen3-8B architecture with the 0528 tokenizer configuration, scores 86.0 on AIME 2024 against 76.0 for Qwen3-8B by DeepSeek's numbers, and is run the same way as Qwen3-8B.
DeepSeek-R1-0528 benchmarks
DeepSeek's published numbers, from the model card, compare 0528 with the original DeepSeek R1. Maximum generation length was 64K tokens, and for the benchmarks that require sampling DeepSeek used temperature 0.6, top-p 0.95 and 16 responses per query to estimate pass@1. The metric is not the same in every row: SWE Verified is scored as problems resolved, Aider-Polyglot as accuracy.
| Benchmark | DeepSeek-R1-0528 | DeepSeek R1 |
|---|---|---|
AIME 2025 Competition math | 87.5 | 70.0 |
HMMT 2025 Olympiad math | 79.4 | 41.7 |
GPQA Diamond Expert science | 81.0 | 71.5 |
LiveCodeBench Competitive coding | 73.3 | 63.5 |
SWE Verified Software engineering | 57.6 | 49.2 |
Aider-Polyglot Code editing | 71.6 | 53.3 |
Humanity's Last Exam Expert questions | 17.7 | 8.5 |
The update wins every row here. HMMT 2025 goes from 41.7 to 79.4, Humanity's Last Exam from 8.5 to 17.7, and Aider-Polyglot from 53.3 to 71.6. The one regression in the vendor's full table is SimpleQA, down from 30.1 to 27.8.
The rest of that table reads: AIME 2024 from 79.8 to 91.4, CNMO 2024 from 78.8 to 86.9, the Codeforces Div1 rating from 1530 to 1930. General knowledge moves less: MMLU-Redux 92.9 to 93.4, MMLU-Pro 84.0 to 85.0, FRAMES 82.5 to 83.0. Tool scores are published for 0528 only, with no R1 column: 37.0 on BFCL v3 MultiTurn, and 53.5 on the airline split of Tau-Bench against 63.9 on retail, with GPT-4.1 playing the user. SWE Verified was run with the Agentless framework, and only text-only prompts were scored on Humanity's Last Exam.
DeepSeek-R1-0528 hardware requirements
The system requirement to check is memory. The GGUF builds come from the unsloth/DeepSeek-R1-0528-GGUF repository on Hugging Face. UD-TQ1_0 is one 161.62 GB file. Every other build is split into parts, so the sizes below are the total across a build's parts.
| Memory | Build to pick | File size |
|---|---|---|
| 192 GB | UD-TQ1_0 | 161.6 GB |
| 256 GB | UD-IQ2_M | 228.5 GB |
| 320 GB | UD-Q3_K_XL | 295.9 GB |
| 384 GB | IQ4_XS | 357.7 GB |
| 512 GB | Q4_K_M | 404.9 GB |
| 640 GB | Q6_K | 551.1 GB |
| 768 GB and up | Q8_0 | 713.3 GB |
When two builds both fit with room left for context, take the larger one. The table skips rungs the repo has. Between UD-TQ1_0 and UD-IQ2_M it carries UD-IQ1_S at 185.5 GB, UD-IQ1_M at 200.5 GB and UD-IQ2_XXS at 216.8 GB. Above UD-IQ2_M it carries UD-Q2_K_XL at 251.1 GB, UD-IQ3_XXS at 272.9 GB and UD-Q4_K_XL at 384.0 GB. Check the listing for a build closer to the memory you have. If none of these tiers fit your hardware, the distilled DeepSeek-R1-0528-Qwen3-8B is the practical way to get the 0528 chain-of-thought on a desktop. If the GGUF format is new to you, start with what GGUF is.
How to run DeepSeek-R1-0528 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for DeepSeek-R1-0528 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
DeepSeek runs the model at temperature 0.6 in its own web app, behind a two-line system prompt: one line names the assistant as DeepSeek-R1 from DeepSeek, the next gives the current date. DeepSeek publishes prompt templates for file uploads and web search, and its DeepSeek Platform endpoint is OpenAI-compatible. For self-serving it points to the DeepSeek-R1 GitHub repository.
For the rest of the family, see every DeepSeek model you can run locally, or browse the full local model catalogue.
DeepSeek-R1-0528 license
DeepSeek-R1-0528 is released under the MIT License, and DeepSeek states that the R1 series, Base and Chat alike, supports commercial use and distillation. That covers building products on the model, redistributing quantized builds, and training smaller models on its outputs, with no usage fee.
