DeepSeek-V3-0324

Updated
24.08.2026
Reasoning
Code
Multilingual
Tools

A 671B-parameter Mixture-of-Experts LLM from DeepSeek-AI (37B active) with a 128K context, strong coding and improved function calling.

At a glance

  • License: MIT
  • Parameters: 684.5B (safetensors count)
  • Context length: Not stated on the vendor card
  • Modalities: Not stated on the vendor card
  • Minimum hardware: At least 186 GB of memory (smallest GGUF build, UD-IQ1_S)

What is DeepSeek-V3-0324?

DeepSeek-V3-0324 is an open-weight chat model from DeepSeek, published on Hugging Face on March 24, 2025 under the MIT license. It keeps exactly the same model structure as the original DeepSeek-V3 and swaps in a stronger checkpoint: DeepSeek reports higher scores on reasoning and coding benchmarks, more executable front-end code, and more accurate function calling. The safetensors listing counts 684.5B parameters, so running it locally is a memory question before it is anything else, and the MIT terms put nothing in the way of self-hosting.

SpecificationDeepSeek-V3-0324
Total parameters684.5B (safetensors count)
ArchitectureSame model structure as DeepSeek-V3
Native precisionFP8: 680.6B of the 684.5B parameters stored as F8_E4M3
Other weight dtypes3.9B in BF16, 41.6M in FP32
Supported featuresFunction calling, JSON output, FIM completion
Prompt formatsDocumented in the DeepSeek-V2.5 repository
Recommended temperature0.3 (the API maps a request of 1.0 down to 0.3)
System promptOfficial app prepends a dated prompt naming the assistant DeepSeek Chat
ServingDeepSeek-V3 GitHub repository; Transformers not supported directly
Technical reportDeepSeek-V3 Technical Report, arXiv 2412.19437
Release dateMarch 24, 2025
LicenseMIT

The released weights are native FP8: 680.6B of the 684.5B parameters are stored as F8_E4M3, with 3.9B left in BF16 and 41.6M in FP32, so the checkpoint is already compact for its parameter count. DeepSeek runs the model at temperature 0.3 in its own web app, and the API remaps a requested 1.0 to that value, so 0.3 is the setting to copy locally. The model supports function calling, JSON output and fill-in-the-middle completion, with the prompt formats documented in the DeepSeek-V2.5 repository.

For serving, the vendor points to the DeepSeek-V3 GitHub repository and notes that Hugging Face Transformers does not support the model directly, so the GGUF builds below are the practical local route. DeepSeek also publishes the prompt scaffolding it runs in production: the official app prepends a system prompt that names the assistant DeepSeek Chat and gives the current date, and the card carries file-upload and web-search templates, the search ones in Chinese and English.

DeepSeek-V3-0324 benchmarks

DeepSeek's numbers, from the model card, compare the 0324 checkpoint with the original DeepSeek-V3 it replaces:

BenchmarkDeepSeek-V3-0324DeepSeek-V3
MMLU-Pro
Academic knowledge
81.275.9
GPQA
Expert science
68.459.1
AIME
Competition math
59.439.6
LiveCodeBench
Competitive coding
49.239.2

Every row improves, with AIME up 19.8 points and LiveCodeBench up 10. These are DeepSeek's own numbers against its own predecessor; the card publishes no comparison with other vendors' models.

The rest of the release notes carry no numbers. DeepSeek says front-end code is more executable and that web pages and game front-ends come out better looking; that medium-to-long-form Chinese writing is aligned with the R1 writing style; that multi-turn interactive rewriting, translation quality and letter writing all improved; that Chinese search returns more detailed outputs on report analysis requests; and that function calling accuracy is up, fixing issues from previous V3 versions. All of that is vendor claim, not measurement.

DeepSeek-V3-0324 hardware requirements

The system requirement to check is memory, and at this scale it is measured in hundreds of gigabytes. The GGUF builds below come from the unsloth/DeepSeek-V3-0324-GGUF repository; each build ships as a set of split files, and the size listed is the total for the set.

MemoryBuild to pickFile size
192 GBUD-IQ1_S186.3 GB
224 GBUD-IQ2_XXS218.7 GB
256 GBUD-Q2_K_XL247.6 GB
384 GBUD-Q3_K_XL320.8 GB
512 GBUD-Q4_K_XL404.9 GB
640 GBQ6_K550.8 GB
768 GB and upQ8_0713.3 GB

The dynamic UD builds stop at UD-Q4_K_XL. The repo also carries plain K-quants: Q2_K at 244.0 GB, Q3_K_M at 319.2 GB, Q4_K_M at 404.4 GB, Q5_K_M at 475.4 GB, plus Q6_K, Q8_0 and a full BF16 conversion spread over 30 files at 1,342.4 GB. When two builds both fit, take the larger one. The low-bit ladder steps tens of gigabytes at a time: UD-IQ1_S at 186.3 GB, UD-IQ1_M at 196.5 GB, UD-IQ2_XXS at 218.7 GB, UD-Q2_K_XL at 247.6 GB, a 61 GB spread across four builds. Even that smallest set needs around 186 GB, so plan for a multi-GPU server or a machine with very large unified memory; if split GGUF files are new to you, start with what GGUF is.

How to run DeepSeek-V3-0324 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for DeepSeek-V3-0324 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

If your hardware is smaller than this model, see every DeepSeek model you can run locally, or the newer DeepSeek-V3.2. For a laptop that tops out at 16 GB, start from our shortlist of the best local LLMs for a 16 GB Mac instead.

DeepSeek-V3-0324 license

The repository and the model weights are licensed under MIT. That permits commercial use, modification, and redistribution with no royalties, so you can self-host the model, fine-tune it, and build products on it without separate licensing terms. The card also carries a citation block for the DeepSeek-V3 Technical Report, arXiv 2412.19437.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download deepseek-ai/DeepSeek-V3-0324
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek-ai/DeepSeek-V3-0324", "messages": [{"role": "user", "content": "Hello"}]}'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-V3-0324", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V3-0324", trust_remote_code=True)
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "deepseek-ai/DeepSeek-V3-0324",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

DeepSeek-V3-0324 is an open-weight large language model released in March 2025 by DeepSeek-AI. It is a Mixture-of-Experts (MoE) model with 671B total parameters, of which about 37B are activated per token. The 0324 checkpoint is a post-training update of DeepSeek-V3 that improves reasoning, coding, and function calling over the original release.

The full FP8 weights are around 700 GB, so running the model unquantized requires a multi-GPU server with roughly 700 GB or more of combined VRAM (for example 8 to 16 high-memory data-center GPUs). Community 4-bit and 1.58-bit GGUF quantizations from projects like Unsloth bring the footprint down to roughly 130 to 400 GB, which can run across high-RAM workstations or smaller GPU clusters at reduced speed. This is not a model that fits on a single consumer card.

Yes. The repository and the model weights are released under the MIT License, which permits commercial use, modification, and redistribution. You can download the weights from Hugging Face at no cost, or access the model through DeepSeek's API and third-party inference providers, which charge per token.

DeepSeek-V3-0324 supports a 128K token context window. That capacity lets it work over long documents, large codebases, and extended multi-turn conversations within a single request.

DeepSeek-V3-0324 is a refreshed post-training checkpoint built on the same architecture as DeepSeek-V3. DeepSeek reports clear benchmark gains, including MMLU-Pro rising from 75.9 to 81.2 and GPQA from 59.1 to 68.4, along with stronger front-end code generation, more reliable multi-turn rewriting, and improved function calling. It is not a separate reasoning model like DeepSeek-R1; it remains a general-purpose chat and instruction model.