MiMo-V2.5-Pro

Updated
05.10.2026
Tools
Thinking
Reasoning
Code
Multilingual

MiMo-V2.5-Pro is Xiaomi’s 1.02T-parameter MoE with 42B active and a 1M-token context, MIT licensed. The smallest GGUF build is 304 GB.

At a glance

  • License: MIT
  • Parameters: 1.02T total, 42B active
  • Context length: 1M tokens
  • Modalities: Text (English and Chinese)
  • Minimum hardware: 304 GB of memory for the smallest GGUF build (UD-IQ1_M)

What is MiMo-V2.5-Pro?

MiMo-V2.5-Pro is Xiaomi's most capable open-source model to date: a Mixture-of-Experts network with 1.02T total parameters, 42B of them active per token, and a context window of up to 1M tokens. Xiaomi built it for demanding agentic work, complex software engineering and long-horizon tasks, and says it stays coherent across thousands of tool calls. The weights went up on Hugging Face in April 2026 under the MIT license.

SpecificationMiMo-V2.5-Pro
Total parameters1.02T
Active parameters42B per token
ArchitectureMoE, 384 routed experts, 8 per token
Layers70 (1 dense + 69 MoE)
Hidden size6144
AttentionHybrid: 60 sliding-window, 10 full-attention layers
Attention heads128 query, 8 KV (GQA)
Head dimension192 QK, 128 V
Feed-forward width2048 per expert, 16384 dense (layer 0)
Context window1M tokens (256K on the Base checkpoint)
Multi-Token Prediction3 MTP modules with dense FFNs
Pre-training27T tokens, FP8 mixed precision, 32k native length
Post-trainingSFT, agentic RL, Multi-Teacher On-Policy Distillation
Weight precisionFP8 (E4M3) mixed
LanguagesEnglish and Chinese
Release dateApril 27, 2026
LicenseMIT

The attention design is what makes the 1M window practical. 60 of the 70 layers use sliding window attention with a 128-token window and only 10 keep global attention, a 6:1 interleave that Xiaomi says cuts KV-cache storage by nearly 7x, with a learnable attention sink bias preserving long-context quality. The three MTP modules are lightweight dense FFNs that Xiaomi integrates natively for training and inference, unlike traditional speculative decoding, and it says they triple output speed during inference.

MiMo-V2.5-Pro benchmarks

Xiaomi's model card publishes scores for the pre-trained base model against MiMo-V2.5, DeepSeek-V4-Pro, DeepSeek-V4-Flash and Kimi-K2, all compared in their base versions:

BenchmarkMiMo-V2.5-ProMiMo-V2.5DeepSeek-V4-ProDeepSeek-V4-FlashKimi-K2
MMLU
General knowledge
89.486.390.188.787.8
MMLU-Pro
Academic knowledge
68.565.873.568.369.2
GPQA-Diamond
Expert science
66.758.1--48.1
GSM8K
Grade-school math
99.683.392.690.892.1
MATH
Advanced math
86.267.764.557.470.2
LiveCodeBench v6
Competitive coding
39.635.5--26.3
SWE-Bench (AgentLess)
Software engineering
35.730.8--28.2

The base model takes every math and coding row above, GSM8K at 99.6 and a 16-point gap on MATH over the closest listed rival. DeepSeek-V4-Pro, a bigger 1.6T model, keeps the lead on broad knowledge in MMLU and MMLU-Pro, and the DeepSeek columns are blank on the science and coding rows, so those wins are counted against MiMo-V2.5 and Kimi-K2 only. It is not a clean sweep: Xiaomi's fuller table gives Kimi-K2 HumanEval+ at 84.8 against 75.6.

Xiaomi also publishes a long-context read on GraphWalks, an OpenAI benchmark that fills the prompt with a directed graph of hex-hash nodes and asks for a breadth-first search to a given depth or a node's parents, scored across the full 32k to 1M span with the fixes Anthropic described. MiMo-V2.5-Pro gets 0.56 BFS and 0.92 Parents at 512k, 0.37 and 0.62 at 1M, where the earlier V2 Pro degrades rapidly past 128k and reaches 0.00 on both subtasks. The agentic behaviour comes from the three-stage recipe of MiMo-V2-Flash: SFT on curated pairs, a domain-specific RL stage with separate teacher models for math, safety and tool use, then Multi-Teacher On-Policy Distillation, where one student learns from its own outputs under token-level guidance from those teachers.

MiMo-V2.5-Pro hardware requirements

The system requirement to check is memory: even heavily quantized, a 1.02T-parameter model starts around 304 GB. Xiaomi ships FP8 weights for server-side serving; the GGUF builds below are community quants from unsloth/MiMo-V2.5-Pro-GGUF, summed across their split files.

MemoryBuild to pickFile size
320 GBUD-IQ1_M304.1 GB
384 GBUD-Q2_K_XL338.4 GB
512 GBUD-IQ3_XXS412.7 GB
640 GBUD-Q4_K_S588.6 GB
768 GBUD-Q4_K_M629.6 GB
896 GBUD-Q5_K_M758.0 GB
1 TBUD-Q6_K846.6 GB
1.5 TB and upQ8_01087.5 GB

When two builds both fit with room left for the KV cache, take the larger one. That matters most in the first two rows: UD-IQ1_M is barely more than a quarter the size of Q8_0, and quality falls fastest there. Xiaomi's deployment notes name SGLang and vLLM, both with published cookbooks, and recommend temperature 1.0 with top_p 0.95 for local runs; their SGLang command serves the FP8 weights at the full 1,048,576-token context with EAGLE speculative decoding across 16-way tensor parallelism. If the format is new to you, start with what GGUF is.

How to run MiMo-V2.5-Pro in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for MiMo-V2.5-Pro in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every MiMo model you can run locally, or browse the full model catalogue for something smaller.

MiMo-V2.5-Pro license

MiMo-V2.5-Pro is released under the MIT license, one of the most permissive there is. It allows commercial use, modification and redistribution with no royalties; the only obligation is keeping the copyright and license notice with the code. Xiaomi also serves the model on its own MiMo API platform, but the MIT terms leave you free to run the weights yourself, quantized or not.

Get the weights from Hugging Face

huggingface-cli download XiaomiMiMo/MiMo-V2.5-Pro
from transformers import AutoModel
model = AutoModel.from_pretrained("XiaomiMiMo/MiMo-V2.5-Pro")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

MiMo-V2.5-Pro is an open-weight Mixture-of-Experts language model from Xiaomi (XiaomiMiMo), with about 1023.2B total parameters and roughly 42B active per token. It was built for agentic tool use, software engineering, and long-horizon reasoning, with a 1,048,576-token context window. The weights are public on Hugging Face under the MIT license.

The full model is large, so a single consumer GPU is not enough. Running the full weights realistically calls for a multi-GPU setup with 80GB-class cards (such as 2x H100) or several RTX 4090s. Quantized builds cut the memory needed, but this remains a heavy model compared to small local LLMs.

Yes. The weights are released under the MIT license, so downloading and running the model yourself costs nothing in license fees. Your only cost is the hardware or electricity to run it. The MIT license also permits commercial use and modification.

Yes. After you download the weights from Hugging Face once, all inference happens on your own machine with no network connection required. In Atomic Chat the model loads and runs entirely on-device, so prompts and outputs never leave your computer.

It is aimed at long, multi-step agent workflows, coding tasks, and reasoning over large inputs. It supports native tool calling and can sustain task chains across many tool calls, and the 1M-token context lets it work over big codebases or document sets in one pass. It handles both English and Chinese.