Kimi-K2.7-Code

Updated
24.08.2026
Tools
Thinking
Vision
Reasoning
Code

Run Kimi-K2.7-Code, a 1T-parameter coding model, locally with Atomic Chat. Offline agentic coding, fully private. Download free.

At a glance

  • License: Modified MIT
  • Parameters: 1T total, 32B activated per token
  • Context length: 256K tokens
  • Modalities: Text and image input, experimental video
  • Minimum hardware: ~320 GB combined RAM and VRAM (UD-IQ1_M GGUF, 303.9 GB)

What is Kimi-K2.7-Code?

Kimi-K2.7-Code is a coding-focused agentic model from Moonshot AI, built on top of Kimi K2.6. It is a 1T-parameter Mixture-of-Experts model that activates 32B parameters per token, tuned for long-horizon software engineering: finishing multi-step coding tasks end to end instead of answering one prompt at a time. Moonshot also cut thinking-token usage by roughly 30% compared with K2.6, so the same job burns fewer tokens. The weights landed on Hugging Face on June 11, 2026 under a Modified MIT license.

SpecificationKimi-K2.7-Code
Total parameters1T
Activated parameters32B per token
ArchitectureMixture-of-Experts: 384 experts, 8 selected per token plus 1 shared, MLA attention
Context window256K tokens
ModalitiesText and image input, video input experimental
Vision encoderMoonViT, 400M parameters
ReasoningThinking forced on, reasoning kept across turns
Native quantizationINT4
Release dateJune 11, 2026
LicenseModified MIT

Per token the router picks 8 of the 384 experts plus 1 shared expert, so a forward pass does the work of a 32B model while the full 1T parameters sit in memory. Two design choices matter for agent work. Thinking is forced on and the model keeps its reasoning content across turns, a mode Moonshot calls preserve_thinking and says raises performance in coding-agent scenarios. And the weights ship in native INT4, the same quantization method as Kimi-K2-Thinking. The model carries a 400M-parameter MoonViT vision encoder, so it takes image input alongside text; video input exists but only as an experimental feature of Moonshot's official API.

Kimi-K2.7-Code benchmarks

Moonshot's model-card numbers compare K2.7 Code with its predecessor Kimi K2.6, GPT-5.5 running in Codex, and Claude Opus 4.8 running in Claude Code:

BenchmarkKimi-K2.7-CodeKimi K2.6GPT-5.5Claude Opus 4.8
Kimi Code Bench v2
Realistic coding
62.050.969.067.4
Program Bench
Program recreation
53.648.369.163.8
MLS Bench Lite
ML research
35.126.735.542.8
Kimi Claw 24/7 Bench
Long-horizon agents
46.942.952.850.4
MCP Atlas
MCP tools
76.069.479.481.3
MCP Mark Verified
MCP servers
81.172.892.976.4

K2.7 Code improves on K2.6 in all six rows while spending about 30% fewer thinking tokens, and it beats Claude Opus 4.8 on MCP Mark Verified. It takes no top scores though: GPT-5.5 or Opus 4.8 leads every row. The honest read is an open-weights coding agent one step below the closed frontier, and this one you can run on your own hardware.

Kimi-K2.7-Code hardware requirements

The system requirement to check is memory, and a 1T-parameter model needs a lot of it: count RAM and VRAM together. The GGUF builds below come from unsloth/Kimi-K2.7-Code-GGUF, each sharded into files of roughly 48 GB; the sizes listed are the total per build.

MemoryBuild to pickFile size
320 GBUD-IQ1_M303.9 GB
384 GBUD-Q2_K_XL339.5 GB
512 GBUD-Q3_K_XL463.9 GB
640 GBUD-Q4_K_XL583.7 GB
768 GB and upUD-Q8_K_XL594.5 GB

When two builds both fit, take the larger one, and leave headroom for the KV cache: the 256K context is not free. If the format is new to you, start with what GGUF is.

How to run Kimi-K2.7-Code in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Kimi-K2.7-Code in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

The rest of the family is at every Kimi model you can run locally, including the general-purpose sibling Kimi-K2-Instruct-0905.

Kimi-K2.7-Code license

Both the code repository and the model weights are released under Moonshot's Modified MIT License, listed on Hugging Face as license "other". The MIT base permits use, modification and redistribution, including commercial use; the "modified" part is Moonshot's own added terms, so read the LICENSE file in the repository before you build a product on it.

Get the weights from Hugging Face

huggingface-cli download moonshotai/Kimi-K2.7-Code
from transformers import AutoModel
model = AutoModel.from_pretrained("moonshotai/Kimi-K2.7-Code")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Kimi-K2.7-Code is an open-weight coding model from Moonshot AI in the Kimi K2 family. It is a Mixture-of-Experts model with about 1058.6B total parameters and a 262144-token context window, tuned for agentic software engineering tasks like multi-step code generation and tool use. Moonshot reports it beats Claude Opus 4.8 on the MCP Mark Verified agent benchmark.

This is a large model, so plan for a lot of memory. A 2-bit quantized build is around 350GB and needs roughly that much combined system RAM and VRAM. People run it on a single 24GB GPU paired with 256GB of system RAM using CPU offloading at a few tokens per second, while multi-GPU rigs run it faster.

Yes. The weights are published openly on Hugging Face under a Modified MIT License, so you can download and run them at no cost. The license permits free use, modification, and redistribution, including commercial use, under its terms. Running it locally in Atomic Chat is free; the only cost is the hardware to host it.

Yes. Once you download the weights, the model runs fully on-device with no internet connection. Loading it in Atomic Chat keeps every prompt and response on your own machine, so your code stays private. You only need a connection for the initial download.

It is built for coding and agentic work: multi-step code generation, tool calling, and reasoning across large codebases in 10+ languages. It reports +21.8% on Kimi Code Bench v2 over K2.6 while using about 30% fewer thinking tokens. For general writing and conversation, Moonshot recommends the more well-rounded K2.6 instead.