What is DeepSeek-Coder-V2-Lite-Instruct?
DeepSeek-Coder-V2-Lite-Instruct is a 16B Mixture-of-Experts code model from DeepSeek AI, the smaller of the two sizes in the DeepSeek-Coder-V2 release, with 2.4B parameters active per token. DeepSeek built it by continuing pre-training from an intermediate checkpoint of DeepSeek-V2 on an additional 6 trillion tokens, a run aimed at coding and mathematical reasoning that, by DeepSeek's account, kept general language performance comparable to DeepSeek-V2. Against the earlier DeepSeek-Coder-33B, the V2 family widens programming language coverage from 86 to 338 and stretches the context window from 16K to 128K tokens. Running DeepSeek-Coder-V2 in BF16 takes eight 80 GB GPUs, and the Lite size is the one DeepSeek writes local examples for: the weights went up on Hugging Face on June 14, 2024 under the DeepSeek Model License, which supports commercial use.
| Specification | DeepSeek-Coder-V2-Lite-Instruct |
|---|---|
| Total parameters | 16B |
| Active parameters | 2.4B per token |
| Architecture | Mixture-of-Experts (DeepSeekMoE framework) |
| Context window | 128K tokens |
| Programming languages | 338 |
| Modalities | Text |
| Chat format | User and Assistant template, optional system message |
| Release date | June 14, 2024 |
| License | DeepSeek Model License, code under MIT |
What DeepSeek-Coder-V2-Lite-Instruct is good at
DeepSeek presents Coder-V2 as an open model that closes the gap with closed ones on code. The card claims performance comparable to GPT-4 Turbo on code-specific tasks, and says that in standard benchmark evaluations DeepSeek-Coder-V2 beats closed-source models such as GPT-4 Turbo, Claude 3 Opus and Gemini 1.5 Pro on coding and math benchmarks. Those claims are written about DeepSeek-Coder-V2 as a release; the card publishes no scores per size. It also credits the continued pre-training with significant gains over DeepSeek-Coder-33B across code-related tasks, reasoning and general capabilities.
For the Lite Instruct checkpoint specifically, the card's own example is chat completion. Messages go through the chat template shipped in tokenizer_config.json, which in plain form is a User and Assistant transcript wrapped in begin and end of sentence tokens, with an optional system message at the top. The code completion and code insertion examples on the same card load DeepSeek-Coder-V2-Lite-Base, not Instruct. DeepSeek shows two ways to serve it: Hugging Face Transformers, which it says you can employ directly, and vLLM, which it marks as recommended and which needs pull request 4650 merged into your vLLM codebase first.
DeepSeek-Coder-V2-Lite-Instruct hardware requirements
The system requirement to check is memory. DeepSeek ships no official GGUF repo, so the sizes below are the real file sizes from the community build bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 8 GB | IQ2_M | 6.33 GB |
| 12 GB | IQ4_XS | 8.57 GB |
| 16 GB | Q5_K_M | 11.85 GB |
| 24 GB and up | Q8_0 | 16.70 GB |
Neighbouring builds sit a gigabyte or two apart, so when two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.
How to run DeepSeek-Coder-V2-Lite-Instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for DeepSeek-Coder-V2-Lite-Instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every DeepSeek model you can run locally, or compare it with a newer MoE coder of a similar footprint, Qwen3-Coder-30B-A3B-Instruct.
DeepSeek-Coder-V2-Lite-Instruct license
The code repository is MIT licensed, and use of the DeepSeek-Coder-V2 Base and Instruct models is subject to the DeepSeek Model License. DeepSeek states that the whole Coder-V2 series, Base and Instruct included, supports commercial use, so you can ship products on top of the model; read the model license itself for the exact terms.
