DeepSeek-Coder-V2-Lite-Instruct

Updated
24.08.2026
Code
Reasoning
Tools

A 16B Mixture-of-Experts code model from DeepSeek AI with 2.4B active params, 128K context, and support for 338 programming languages.

At a glance

  • License: DeepSeek Model License, commercial use supported
  • Parameters: 16B total, 2.4B active (MoE)
  • Context length: 128K tokens
  • Modalities: Text in, text out
  • Minimum hardware: 8 GB memory with the IQ2_M GGUF (6.33 GB)

What is DeepSeek-Coder-V2-Lite-Instruct?

DeepSeek-Coder-V2-Lite-Instruct is a 16B Mixture-of-Experts code model from DeepSeek AI, the smaller of the two sizes in the DeepSeek-Coder-V2 release, with 2.4B parameters active per token. DeepSeek built it by continuing pre-training from an intermediate checkpoint of DeepSeek-V2 on an additional 6 trillion tokens, a run aimed at coding and mathematical reasoning that, by DeepSeek's account, kept general language performance comparable to DeepSeek-V2. Against the earlier DeepSeek-Coder-33B, the V2 family widens programming language coverage from 86 to 338 and stretches the context window from 16K to 128K tokens. Running DeepSeek-Coder-V2 in BF16 takes eight 80 GB GPUs, and the Lite size is the one DeepSeek writes local examples for: the weights went up on Hugging Face on June 14, 2024 under the DeepSeek Model License, which supports commercial use.

SpecificationDeepSeek-Coder-V2-Lite-Instruct
Total parameters16B
Active parameters2.4B per token
ArchitectureMixture-of-Experts (DeepSeekMoE framework)
Context window128K tokens
Programming languages338
ModalitiesText
Chat formatUser and Assistant template, optional system message
Release dateJune 14, 2024
LicenseDeepSeek Model License, code under MIT

What DeepSeek-Coder-V2-Lite-Instruct is good at

DeepSeek presents Coder-V2 as an open model that closes the gap with closed ones on code. The card claims performance comparable to GPT-4 Turbo on code-specific tasks, and says that in standard benchmark evaluations DeepSeek-Coder-V2 beats closed-source models such as GPT-4 Turbo, Claude 3 Opus and Gemini 1.5 Pro on coding and math benchmarks. Those claims are written about DeepSeek-Coder-V2 as a release; the card publishes no scores per size. It also credits the continued pre-training with significant gains over DeepSeek-Coder-33B across code-related tasks, reasoning and general capabilities.

For the Lite Instruct checkpoint specifically, the card's own example is chat completion. Messages go through the chat template shipped in tokenizer_config.json, which in plain form is a User and Assistant transcript wrapped in begin and end of sentence tokens, with an optional system message at the top. The code completion and code insertion examples on the same card load DeepSeek-Coder-V2-Lite-Base, not Instruct. DeepSeek shows two ways to serve it: Hugging Face Transformers, which it says you can employ directly, and vLLM, which it marks as recommended and which needs pull request 4650 merged into your vLLM codebase first.

DeepSeek-Coder-V2-Lite-Instruct hardware requirements

The system requirement to check is memory. DeepSeek ships no official GGUF repo, so the sizes below are the real file sizes from the community build bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF.

MemoryBuild to pickFile size
8 GBIQ2_M6.33 GB
12 GBIQ4_XS8.57 GB
16 GBQ5_K_M11.85 GB
24 GB and upQ8_016.70 GB

Neighbouring builds sit a gigabyte or two apart, so when two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.

How to run DeepSeek-Coder-V2-Lite-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for DeepSeek-Coder-V2-Lite-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every DeepSeek model you can run locally, or compare it with a newer MoE coder of a similar footprint, Qwen3-Coder-30B-A3B-Instruct.

DeepSeek-Coder-V2-Lite-Instruct license

The code repository is MIT licensed, and use of the DeepSeek-Coder-V2 Base and Instruct models is subject to the DeepSeek Model License. DeepSeek states that the whole Coder-V2 series, Base and Instruct included, supports commercial use, so you can ship products on top of the model; read the model license itself for the exact terms.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct",
    "messages": [{"role": "user", "content": "Write a quicksort in Python."}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct", trust_remote_code=True, torch_dtype=torch.bfloat16).cuda()
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct", trust_remote_code=True)
messages = [{"role": "user", "content": "Write a quicksort in Python."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(out[0][len(inputs[0]):], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "not-needed" });
const res = await client.chat.completions.create({
  model: "deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct",
  messages: [{ role: "user", content: "Write a quicksort in Python." }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

DeepSeek-Coder-V2-Lite-Instruct is an open-source code language model from DeepSeek AI. It uses a Mixture-of-Experts (MoE) design with 16B total parameters but only 2.4B active per token, and it was continued-pretrained from DeepSeek-V2 on an extra 6 trillion tokens. The instruct variant is tuned for chat-style coding assistance and supports a 128K context window.

Because only 2.4B of the 16B parameters are active per token, the Lite model runs far lighter than its size suggests. In BF16 it needs roughly 32 GB of memory, but 4-bit GGUF quants bring it down to around 10-12 GB, so it fits on a single 16 GB GPU or a Mac with enough unified memory. The larger 236B DeepSeek-Coder-V2 model, by contrast, requires 8x80 GB GPUs for BF16 inference.

Yes. The weights are published on Hugging Face and can be downloaded at no cost. The code repository is MIT-licensed, and the model weights are released under the DeepSeek Model License, which explicitly permits commercial use. You should review the model license terms before deploying it in a product.

DeepSeek-Coder-V2 supports 338 programming languages, up from 86 in the original DeepSeek-Coder. The same training applies to the Lite variant. It handles code completion, code insertion (fill-in-the-middle), generation, and debugging across that language set, and the 128K context lets it work over large files and repositories.

For its small active footprint it is strong. The Lite-Instruct model scores around 81% pass@1 on HumanEval and performs well on MBPP and math benchmarks like GSM8K. The full 236B DeepSeek-Coder-V2 reaches performance comparable to GPT-4-Turbo on code-specific tasks, while the Lite model trades some of that accuracy for the ability to run on a single consumer GPU.