Kimi-K2-Instruct-0905

Updated
24.08.2026
Tools
Reasoning
Code

Run Kimi-K2-Instruct-0905, a 1T-parameter MoE model, locally in Atomic Chat. Private, offline agentic tool use, no API keys, no cloud. Free.

At a glance

  • License: Modified MIT
  • Parameters: 1T total, 32B activated (MoE)
  • Context length: 256K tokens
  • Modalities: Text
  • Minimum hardware: ~249 GB combined RAM and VRAM for the smallest GGUF (UD-TQ1_0)

What is Kimi-K2-Instruct-0905?

Kimi-K2-Instruct-0905 is Moonshot AI's September 2025 update to Kimi K2, a mixture-of-experts language model with 1 trillion total parameters, 32 billion of them activated per token. The 0905 build has one clear focus: agentic coding. Moonshot reports gains on public benchmarks and on real-world coding-agent tasks, better frontend output, and a context window raised from 128K to 256K tokens for long-horizon work. The weights are open under a Modified MIT license, so the whole thing can run on your own hardware, provided you have a lot of it.

SpecificationKimi-K2-Instruct-0905
ArchitectureMixture-of-Experts (MoE)
Total parameters1T
Activated parameters32B
Layers61 (1 dense)
Experts384, 8 selected per token, 1 shared
AttentionMLA, 64 heads
Attention hidden dimension7168
MoE hidden dimension per expert2048
Activation functionSwiGLU
Context window256K tokens
Vocabulary160K
Checkpoint formatBlock-FP8
Recommended temperature0.6
Release dateSeptember 3, 2025
LicenseModified MIT

Only 32B of the trillion parameters do work on any given token: the router picks 8 of the 384 experts, plus 1 shared expert that every token passes through. Each expert is narrow at 2048 hidden units, while attention runs at 7168 with 64 heads and MLA across 61 layers, one of them dense. Per-token compute stays at 32B-model levels while memory still has to hold the full 1T weights, which is why the hardware table below is measured in hundreds of gigabytes rather than tens.

Kimi-K2-Instruct-0905 benchmarks

Moonshot AI's launch numbers, reported as means over five full-test-set runs, compare the 0905 build with its predecessor K2-Instruct-0711, three rival models and Claude Sonnet 4:

BenchmarkKimi-K2-Instruct-0905K2-Instruct-0711Qwen3-Coder-480B-A35BGLM-4.5DeepSeek-V3.1Claude Sonnet 4
SWE-bench Verified
Software engineering
69.265.869.664.266.072.7
SWE-bench Multilingual
Multilingual engineering
55.947.354.752.754.553.3
Multi-SWE-Bench
Cross-language engineering
33.531.332.731.729.035.7
Terminal-Bench
Terminal agents
44.537.537.539.931.336.4
SWE-Dev
Feature development
66.661.964.763.253.367.1

The 0905 build takes Terminal-Bench and SWE-bench Multilingual outright and improves on 0711 in every row, while Claude Sonnet 4 keeps SWE-bench Verified, Multi-SWE-Bench and SWE-Dev. Among the open models in the table it leads everywhere except SWE-bench Verified, where Qwen3-Coder finishes 0.4 points ahead.

The methodology matters here. Before each run Moonshot prunes the repository so that every Git object unreachable from the target commit disappears, so the agent sees only code that existed then; on SWE-Dev it also deletes the test files that exercise the functions the agent must write, removing hints about the intended implementation. Every result except Terminal-Bench, which runs on Terminus-2, comes from an in-house harness derived from SWE-agent with the Bash and Edit tool context windows clamped and the system prompt rewritten for the task, and some baseline figures are quoted from the other vendors' own reports rather than rerun. Moonshot also claims strong tool calling: you pass the available tools with each request and the model decides on its own when and how to invoke them, which needs an engine that supports Kimi K2's native tool parsing.

Kimi-K2-Instruct-0905 hardware requirements

The system requirement to check is memory, and here that means combined RAM plus VRAM in the hundreds of gigabytes. The community GGUF builds come from unsloth/Kimi-K2-Instruct-0905-GGUF; each size below is the sum of that build's file shards.

MemoryBuild to pickFile size
256 GBUD-TQ1_0248.6 GB
320 GBUD-IQ1_M304.8 GB
384 GBUD-IQ2_M349.1 GB
448 GBUD-IQ3_XXS417.3 GB
512 GBUD-Q3_K_XL452.4 GB
640 GBUD-Q4_K_XL588.1 GB
768 GBUD-Q5_K_XL734.1 GB
1 TB and upUD-Q6_K_XL879.1 GB

Sizes cover the weights only, so keep the gap between the file and your memory ceiling free for the KV cache; when two builds both fit with room to spare, take the larger one. The rungs are not evenly spaced: UD-IQ2_M down to UD-TQ1_0 sheds about 100 GB, and that bottom end of the ladder is where quality falls fastest, so go below UD-IQ2_M only when the memory is genuinely not there. The same repo goes on up to BF16 at 2053.2 GB. Moonshot ships the original block-FP8 checkpoint for vLLM, SGLang, KTransformers and TensorRT-LLM, with deployment examples for the first two. If the format is new to you, start with what GGUF is.

How to run Kimi-K2-Instruct-0905 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Kimi-K2-Instruct-0905 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

Moonshot recommends temperature 0.6, with the system prompt "You are Kimi, an AI assistant created by Moonshot AI." as the default.

For the rest of the family, see every Kimi model you can run locally, or the earlier Kimi-K2-Instruct that this release supersedes. If the ladder above rules it out, our catalogue of local models has smaller options.

Kimi-K2-Instruct-0905 license

Both the code repository and the model weights ship under a Modified MIT License. The base is MIT, one of the most permissive licenses around; the modified part means Moonshot AI attached its own terms on top, so read the LICENSE file in the Hugging Face repo before you build a commercial product on the weights.

Get the weights from Hugging Face

huggingface-cli download moonshotai/Kimi-K2-Instruct-0905
from transformers import AutoModel
model = AutoModel.from_pretrained("moonshotai/Kimi-K2-Instruct-0905")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

It is the September 2025 release of Moonshot AI's Kimi K2 instruct model, a mixture-of-experts LLM with about 1 trillion total parameters and roughly 32 billion active per token. It is tuned for chat, tool calling, and agentic coding, and supports a 256K-token context window. The 0905 update focused on stronger coding-agent performance and better frontend code generation.

It is demanding. Quantized GGUF builds need roughly 250GB or more of combined system RAM and VRAM, and higher-precision versions need much more, so it targets workstations and servers rather than a typical laptop. A common rule is that your RAM plus VRAM should about match the quant size; if it falls short the model still runs but slows down as layers offload to disk.

Yes. The weights are openly published on Hugging Face and free to download, under a Modified MIT License that Moonshot AI lists as a custom ("other") license. Running it locally costs only your own hardware and electricity, with no per-token API fees.

Yes, once you have downloaded the weights it runs fully offline with no internet connection. In Atomic Chat the model executes on-device, so your prompts and data never leave your machine. The only step that needs a connection is the initial download of the model files.

It is strongest at agentic coding, where it plans and carries out multi-step programming tasks while calling tools, including frontend work that improved in the 0905 release. The 256K context also makes it well suited to long sessions over large codebases or document sets. Its tool-calling support lets it drive local agents and automations.