What is Kimi-K2-Instruct-0905?
Kimi-K2-Instruct-0905 is Moonshot AI's September 2025 update to Kimi K2, a mixture-of-experts language model with 1 trillion total parameters, 32 billion of them activated per token. The 0905 build has one clear focus: agentic coding. Moonshot reports gains on public benchmarks and on real-world coding-agent tasks, better frontend output, and a context window raised from 128K to 256K tokens for long-horizon work. The weights are open under a Modified MIT license, so the whole thing can run on your own hardware, provided you have a lot of it.
| Specification | Kimi-K2-Instruct-0905 |
|---|---|
| Architecture | Mixture-of-Experts (MoE) |
| Total parameters | 1T |
| Activated parameters | 32B |
| Layers | 61 (1 dense) |
| Experts | 384, 8 selected per token, 1 shared |
| Attention | MLA, 64 heads |
| Attention hidden dimension | 7168 |
| MoE hidden dimension per expert | 2048 |
| Activation function | SwiGLU |
| Context window | 256K tokens |
| Vocabulary | 160K |
| Checkpoint format | Block-FP8 |
| Recommended temperature | 0.6 |
| Release date | September 3, 2025 |
| License | Modified MIT |
Only 32B of the trillion parameters do work on any given token: the router picks 8 of the 384 experts, plus 1 shared expert that every token passes through. Each expert is narrow at 2048 hidden units, while attention runs at 7168 with 64 heads and MLA across 61 layers, one of them dense. Per-token compute stays at 32B-model levels while memory still has to hold the full 1T weights, which is why the hardware table below is measured in hundreds of gigabytes rather than tens.
Kimi-K2-Instruct-0905 benchmarks
Moonshot AI's launch numbers, reported as means over five full-test-set runs, compare the 0905 build with its predecessor K2-Instruct-0711, three rival models and Claude Sonnet 4:
| Benchmark | Kimi-K2-Instruct-0905 | K2-Instruct-0711 | Qwen3-Coder-480B-A35B | GLM-4.5 | DeepSeek-V3.1 | Claude Sonnet 4 |
|---|---|---|---|---|---|---|
SWE-bench Verified Software engineering | 69.2 | 65.8 | 69.6 | 64.2 | 66.0 | 72.7 |
SWE-bench Multilingual Multilingual engineering | 55.9 | 47.3 | 54.7 | 52.7 | 54.5 | 53.3 |
Multi-SWE-Bench Cross-language engineering | 33.5 | 31.3 | 32.7 | 31.7 | 29.0 | 35.7 |
Terminal-Bench Terminal agents | 44.5 | 37.5 | 37.5 | 39.9 | 31.3 | 36.4 |
SWE-Dev Feature development | 66.6 | 61.9 | 64.7 | 63.2 | 53.3 | 67.1 |
The 0905 build takes Terminal-Bench and SWE-bench Multilingual outright and improves on 0711 in every row, while Claude Sonnet 4 keeps SWE-bench Verified, Multi-SWE-Bench and SWE-Dev. Among the open models in the table it leads everywhere except SWE-bench Verified, where Qwen3-Coder finishes 0.4 points ahead.
The methodology matters here. Before each run Moonshot prunes the repository so that every Git object unreachable from the target commit disappears, so the agent sees only code that existed then; on SWE-Dev it also deletes the test files that exercise the functions the agent must write, removing hints about the intended implementation. Every result except Terminal-Bench, which runs on Terminus-2, comes from an in-house harness derived from SWE-agent with the Bash and Edit tool context windows clamped and the system prompt rewritten for the task, and some baseline figures are quoted from the other vendors' own reports rather than rerun. Moonshot also claims strong tool calling: you pass the available tools with each request and the model decides on its own when and how to invoke them, which needs an engine that supports Kimi K2's native tool parsing.
Kimi-K2-Instruct-0905 hardware requirements
The system requirement to check is memory, and here that means combined RAM plus VRAM in the hundreds of gigabytes. The community GGUF builds come from unsloth/Kimi-K2-Instruct-0905-GGUF; each size below is the sum of that build's file shards.
| Memory | Build to pick | File size |
|---|---|---|
| 256 GB | UD-TQ1_0 | 248.6 GB |
| 320 GB | UD-IQ1_M | 304.8 GB |
| 384 GB | UD-IQ2_M | 349.1 GB |
| 448 GB | UD-IQ3_XXS | 417.3 GB |
| 512 GB | UD-Q3_K_XL | 452.4 GB |
| 640 GB | UD-Q4_K_XL | 588.1 GB |
| 768 GB | UD-Q5_K_XL | 734.1 GB |
| 1 TB and up | UD-Q6_K_XL | 879.1 GB |
Sizes cover the weights only, so keep the gap between the file and your memory ceiling free for the KV cache; when two builds both fit with room to spare, take the larger one. The rungs are not evenly spaced: UD-IQ2_M down to UD-TQ1_0 sheds about 100 GB, and that bottom end of the ladder is where quality falls fastest, so go below UD-IQ2_M only when the memory is genuinely not there. The same repo goes on up to BF16 at 2053.2 GB. Moonshot ships the original block-FP8 checkpoint for vLLM, SGLang, KTransformers and TensorRT-LLM, with deployment examples for the first two. If the format is new to you, start with what GGUF is.
How to run Kimi-K2-Instruct-0905 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Kimi-K2-Instruct-0905 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
Moonshot recommends temperature 0.6, with the system prompt "You are Kimi, an AI assistant created by Moonshot AI." as the default.
For the rest of the family, see every Kimi model you can run locally, or the earlier Kimi-K2-Instruct that this release supersedes. If the ladder above rules it out, our catalogue of local models has smaller options.
Kimi-K2-Instruct-0905 license
Both the code repository and the model weights ship under a Modified MIT License. The base is MIT, one of the most permissive licenses around; the modified part means Moonshot AI attached its own terms on top, so read the LICENSE file in the Hugging Face repo before you build a commercial product on the weights.
