What is Kimi-K2-Instruct?
Kimi-K2-Instruct is the post-trained chat model of Kimi K2, a mixture-of-experts model from Moonshot AI with 1 trillion total parameters and 32 billion activated per token. The weights went up on Hugging Face on 11 July 2025, stored in block-fp8 format and released under a Modified MIT License. Moonshot built the model for agentic work: tool use, reasoning and autonomous problem-solving. It is also a "reflex-grade" model in their wording, meaning it answers directly, without a long thinking phase.
| Specification | Kimi-K2-Instruct |
|---|---|
| Total parameters | 1T |
| Activated parameters | 32B |
| Architecture | Mixture-of-Experts, 384 experts, 8 selected per token plus 1 shared |
| Layers | 61 (1 dense) |
| Attention | MLA |
| Context window | 128K tokens |
| Vocabulary | 160K |
| Modalities | Text |
| Recommended temperature | 0.6 |
| Release date | 11 July 2025 |
| License | Modified MIT |
Of the 1 trillion parameters, only 32 billion run on any given token: the router picks 8 of the 384 experts, plus one shared expert that is always on. Moonshot pre-trained the model on 15.5T tokens with zero training instability, which it credits to MuonClip, the Muon optimizer applied at what Moonshot calls an unprecedented scale with new techniques to resolve instabilities while scaling up. The result is a model that ships tool calling natively: you pass the tool list in each request and the model decides when and how to invoke them.
Kimi-K2-Instruct benchmarks
Moonshot's launch numbers, from the model card, put Kimi-K2-Instruct against DeepSeek-V3-0324, Qwen3-235B-A22B in non-thinking mode, Claude Sonnet 4 and Claude Opus 4 without extended thinking, and GPT-4.1. The vendor table also carries a Gemini 2.5 Flash Preview column, which tops none of the rows below:
| Benchmark | Kimi-K2-Instruct | DeepSeek-V3-0324 | Qwen3-235B-A22B | Claude Sonnet 4 | Claude Opus 4 | GPT-4.1 |
|---|---|---|---|---|---|---|
LiveCodeBench v6 Competitive coding | 53.7 | 46.9 | 37.0 | 48.5 | 47.4 | 44.7 |
SWE-bench Verified Software engineering | 65.8 | 38.8 | 34.4 | 72.7 | 72.5 | 54.6 |
Tau2 telecom Agentic tools | 65.8 | 32.5 | 22.1 | 45.2 | 57.0 | 38.6 |
AceBench Function calling | 76.5 | 72.7 | 70.5 | 76.2 | 75.6 | 80.1 |
AIME 2025 Competition math | 49.5 | 46.7 | 24.7 | 33.1 | 33.9 | 37.0 |
GPQA-Diamond Expert science | 75.1 | 68.4 | 62.9 | 70.0 | 74.9 | 66.3 |
MMLU General knowledge | 89.5 | 89.4 | 87.0 | 91.5 | 92.9 | 90.4 |
Kimi-K2-Instruct takes four of the seven rows and stays close on the rest. Claude Opus 4 leads MMLU, Claude Sonnet 4 leads the agentic SWE-bench Verified run, and GPT-4.1 leads AceBench. The SWE-bench score here is single-attempt with bash and editor tools; with multiple attempts and an internal scoring model, Moonshot reports 71.6.
Kimi-K2-Instruct hardware requirements
The system requirement to check is memory, and at this scale that means a server or a very large workstation: the smallest GGUF build is still about a quarter of a terabyte. The community quantizations with real file listings live at unsloth/Kimi-K2-Instruct-GGUF; the sizes below are the summed multi-part files from that repo.
| Memory | Build to pick | File size |
|---|---|---|
| 256 GB | UD-TQ1_0 | 243.6 GB |
| 320 GB | UD-IQ1_M | 304.2 GB |
| 384 GB | UD-IQ2_M | 347.1 GB |
| 512 GB | UD-Q3_K_XL | 452.1 GB |
| 640 GB and up | UD-Q4_K_XL | 587.1 GB |
Leave headroom above the file size for the KV cache, and when two builds both fit, take the larger one. For the original block-fp8 checkpoint, Moonshot recommends vLLM, SGLang, KTransformers or TensorRT-LLM; if the quantized route is new to you, start with what GGUF is.
How to run Kimi-K2-Instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Kimi-K2-Instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
Moonshot has since published an updated checkpoint, Kimi-K2-Instruct-0905, and the rest of the family is on our Kimi models page.
Kimi-K2-Instruct license
Both the code repository and the model weights are released under the Modified MIT License, listed in the Hugging Face metadata as license: other, license_name: modified-mit. Moonshot does not spell out the terms on the model card, so read the LICENSE file in the repo before you ship a product on the weights.
