Kimi-K2-Instruct

Updated
24.08.2026
Tools
Code
Reasoning
Multilingual

A 1T-parameter MoE chat model from Moonshot AI with 32B active parameters, built for agentic tool use and strong coding.

At a glance

  • License: Modified MIT
  • Parameters: 1T total, 32B activated (MoE)
  • Context length: 128K tokens
  • Modalities: Text
  • Minimum hardware: about 244 GB of combined memory for the smallest GGUF build (UD-TQ1_0)

What is Kimi-K2-Instruct?

Kimi-K2-Instruct is the post-trained chat model of Kimi K2, a mixture-of-experts model from Moonshot AI with 1 trillion total parameters and 32 billion activated per token. The weights went up on Hugging Face on 11 July 2025, stored in block-fp8 format and released under a Modified MIT License. Moonshot built the model for agentic work: tool use, reasoning and autonomous problem-solving. It is also a "reflex-grade" model in their wording, meaning it answers directly, without a long thinking phase.

SpecificationKimi-K2-Instruct
Total parameters1T
Activated parameters32B
ArchitectureMixture-of-Experts, 384 experts, 8 selected per token plus 1 shared
Layers61 (1 dense)
AttentionMLA
Context window128K tokens
Vocabulary160K
ModalitiesText
Recommended temperature0.6
Release date11 July 2025
LicenseModified MIT

Of the 1 trillion parameters, only 32 billion run on any given token: the router picks 8 of the 384 experts, plus one shared expert that is always on. Moonshot pre-trained the model on 15.5T tokens with zero training instability, which it credits to MuonClip, the Muon optimizer applied at what Moonshot calls an unprecedented scale with new techniques to resolve instabilities while scaling up. The result is a model that ships tool calling natively: you pass the tool list in each request and the model decides when and how to invoke them.

Kimi-K2-Instruct benchmarks

Moonshot's launch numbers, from the model card, put Kimi-K2-Instruct against DeepSeek-V3-0324, Qwen3-235B-A22B in non-thinking mode, Claude Sonnet 4 and Claude Opus 4 without extended thinking, and GPT-4.1. The vendor table also carries a Gemini 2.5 Flash Preview column, which tops none of the rows below:

BenchmarkKimi-K2-InstructDeepSeek-V3-0324Qwen3-235B-A22BClaude Sonnet 4Claude Opus 4GPT-4.1
LiveCodeBench v6
Competitive coding
53.746.937.048.547.444.7
SWE-bench Verified
Software engineering
65.838.834.472.772.554.6
Tau2 telecom
Agentic tools
65.832.522.145.257.038.6
AceBench
Function calling
76.572.770.576.275.680.1
AIME 2025
Competition math
49.546.724.733.133.937.0
GPQA-Diamond
Expert science
75.168.462.970.074.966.3
MMLU
General knowledge
89.589.487.091.592.990.4

Kimi-K2-Instruct takes four of the seven rows and stays close on the rest. Claude Opus 4 leads MMLU, Claude Sonnet 4 leads the agentic SWE-bench Verified run, and GPT-4.1 leads AceBench. The SWE-bench score here is single-attempt with bash and editor tools; with multiple attempts and an internal scoring model, Moonshot reports 71.6.

Kimi-K2-Instruct hardware requirements

The system requirement to check is memory, and at this scale that means a server or a very large workstation: the smallest GGUF build is still about a quarter of a terabyte. The community quantizations with real file listings live at unsloth/Kimi-K2-Instruct-GGUF; the sizes below are the summed multi-part files from that repo.

MemoryBuild to pickFile size
256 GBUD-TQ1_0243.6 GB
320 GBUD-IQ1_M304.2 GB
384 GBUD-IQ2_M347.1 GB
512 GBUD-Q3_K_XL452.1 GB
640 GB and upUD-Q4_K_XL587.1 GB

Leave headroom above the file size for the KV cache, and when two builds both fit, take the larger one. For the original block-fp8 checkpoint, Moonshot recommends vLLM, SGLang, KTransformers or TensorRT-LLM; if the quantized route is new to you, start with what GGUF is.

How to run Kimi-K2-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Kimi-K2-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

Moonshot has since published an updated checkpoint, Kimi-K2-Instruct-0905, and the rest of the family is on our Kimi models page.

Kimi-K2-Instruct license

Both the code repository and the model weights are released under the Modified MIT License, listed in the Hugging Face metadata as license: other, license_name: modified-mit. Moonshot does not spell out the terms on the model card, so read the LICENSE file in the repo before you ship a product on the weights.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download moonshotai/Kimi-K2-Instruct
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "moonshotai/Kimi-K2-Instruct", "messages": [{"role": "user", "content": "Hello"}]}'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("moonshotai/Kimi-K2-Instruct", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("moonshotai/Kimi-K2-Instruct", trust_remote_code=True)
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "moonshotai/Kimi-K2-Instruct",
  messages: [{ role: "user", content: "Hello" }]
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Kimi-K2-Instruct is a post-trained, instruction-tuned large language model from Moonshot AI. It is a mixture-of-experts model with 1 trillion total parameters and 32 billion activated per token, tuned for general-purpose chat and agentic tool use. Moonshot describes it as a reflex-grade model without long chain-of-thought thinking.

Running Kimi-K2-Instruct locally is demanding because the full model has 1 trillion parameters. The weights ship in block-fp8 format and still occupy roughly 1 TB, so a single multi-GPU server with hundreds of gigabytes of combined VRAM is needed even at fp8. Most people run it through hosted inference providers or Moonshot's own API rather than on personal hardware.

Yes. Moonshot AI released both the code and the model weights under a Modified MIT License, and the checkpoints are available on Hugging Face. The license permits commercial use; the main added condition is an attribution clause that applies to very large-scale commercial deployments. You can also use it through Moonshot's paid API.

Kimi-K2-Instruct supports a 128K-token context window. It uses Multi-head Latent Attention (MLA) across 61 layers, which helps keep long-context inference efficient relative to the model's scale.

Kimi-K2-Instruct is tuned specifically for coding and tool use. It scores 65.8% pass@1 on SWE-bench Verified and 47.3% on SWE-bench Multilingual with bash and editor tools, single attempt and no test-time compute. It also has native tool-calling: you pass the available tools in each request and the model decides when to invoke them.