What is MiMo-V2.5-Pro?
MiMo-V2.5-Pro is Xiaomi's most capable open-source model to date: a Mixture-of-Experts network with 1.02T total parameters, 42B of them active per token, and a context window of up to 1M tokens. Xiaomi built it for demanding agentic work, complex software engineering and long-horizon tasks, and says it stays coherent across thousands of tool calls. The weights went up on Hugging Face in April 2026 under the MIT license.
| Specification | MiMo-V2.5-Pro |
|---|---|
| Total parameters | 1.02T |
| Active parameters | 42B per token |
| Architecture | MoE, 384 routed experts, 8 per token |
| Layers | 70 (1 dense + 69 MoE) |
| Hidden size | 6144 |
| Attention | Hybrid: 60 sliding-window, 10 full-attention layers |
| Attention heads | 128 query, 8 KV (GQA) |
| Head dimension | 192 QK, 128 V |
| Feed-forward width | 2048 per expert, 16384 dense (layer 0) |
| Context window | 1M tokens (256K on the Base checkpoint) |
| Multi-Token Prediction | 3 MTP modules with dense FFNs |
| Pre-training | 27T tokens, FP8 mixed precision, 32k native length |
| Post-training | SFT, agentic RL, Multi-Teacher On-Policy Distillation |
| Weight precision | FP8 (E4M3) mixed |
| Languages | English and Chinese |
| Release date | April 27, 2026 |
| License | MIT |
The attention design is what makes the 1M window practical. 60 of the 70 layers use sliding window attention with a 128-token window and only 10 keep global attention, a 6:1 interleave that Xiaomi says cuts KV-cache storage by nearly 7x, with a learnable attention sink bias preserving long-context quality. The three MTP modules are lightweight dense FFNs that Xiaomi integrates natively for training and inference, unlike traditional speculative decoding, and it says they triple output speed during inference.
MiMo-V2.5-Pro benchmarks
Xiaomi's model card publishes scores for the pre-trained base model against MiMo-V2.5, DeepSeek-V4-Pro, DeepSeek-V4-Flash and Kimi-K2, all compared in their base versions:
| Benchmark | MiMo-V2.5-Pro | MiMo-V2.5 | DeepSeek-V4-Pro | DeepSeek-V4-Flash | Kimi-K2 |
|---|---|---|---|---|---|
MMLU General knowledge | 89.4 | 86.3 | 90.1 | 88.7 | 87.8 |
MMLU-Pro Academic knowledge | 68.5 | 65.8 | 73.5 | 68.3 | 69.2 |
GPQA-Diamond Expert science | 66.7 | 58.1 | - | - | 48.1 |
GSM8K Grade-school math | 99.6 | 83.3 | 92.6 | 90.8 | 92.1 |
MATH Advanced math | 86.2 | 67.7 | 64.5 | 57.4 | 70.2 |
LiveCodeBench v6 Competitive coding | 39.6 | 35.5 | - | - | 26.3 |
SWE-Bench (AgentLess) Software engineering | 35.7 | 30.8 | - | - | 28.2 |
The base model takes every math and coding row above, GSM8K at 99.6 and a 16-point gap on MATH over the closest listed rival. DeepSeek-V4-Pro, a bigger 1.6T model, keeps the lead on broad knowledge in MMLU and MMLU-Pro, and the DeepSeek columns are blank on the science and coding rows, so those wins are counted against MiMo-V2.5 and Kimi-K2 only. It is not a clean sweep: Xiaomi's fuller table gives Kimi-K2 HumanEval+ at 84.8 against 75.6.
Xiaomi also publishes a long-context read on GraphWalks, an OpenAI benchmark that fills the prompt with a directed graph of hex-hash nodes and asks for a breadth-first search to a given depth or a node's parents, scored across the full 32k to 1M span with the fixes Anthropic described. MiMo-V2.5-Pro gets 0.56 BFS and 0.92 Parents at 512k, 0.37 and 0.62 at 1M, where the earlier V2 Pro degrades rapidly past 128k and reaches 0.00 on both subtasks. The agentic behaviour comes from the three-stage recipe of MiMo-V2-Flash: SFT on curated pairs, a domain-specific RL stage with separate teacher models for math, safety and tool use, then Multi-Teacher On-Policy Distillation, where one student learns from its own outputs under token-level guidance from those teachers.
MiMo-V2.5-Pro hardware requirements
The system requirement to check is memory: even heavily quantized, a 1.02T-parameter model starts around 304 GB. Xiaomi ships FP8 weights for server-side serving; the GGUF builds below are community quants from unsloth/MiMo-V2.5-Pro-GGUF, summed across their split files.
| Memory | Build to pick | File size |
|---|---|---|
| 320 GB | UD-IQ1_M | 304.1 GB |
| 384 GB | UD-Q2_K_XL | 338.4 GB |
| 512 GB | UD-IQ3_XXS | 412.7 GB |
| 640 GB | UD-Q4_K_S | 588.6 GB |
| 768 GB | UD-Q4_K_M | 629.6 GB |
| 896 GB | UD-Q5_K_M | 758.0 GB |
| 1 TB | UD-Q6_K | 846.6 GB |
| 1.5 TB and up | Q8_0 | 1087.5 GB |
When two builds both fit with room left for the KV cache, take the larger one. That matters most in the first two rows: UD-IQ1_M is barely more than a quarter the size of Q8_0, and quality falls fastest there. Xiaomi's deployment notes name SGLang and vLLM, both with published cookbooks, and recommend temperature 1.0 with top_p 0.95 for local runs; their SGLang command serves the FP8 weights at the full 1,048,576-token context with EAGLE speculative decoding across 16-way tensor parallelism. If the format is new to you, start with what GGUF is.
How to run MiMo-V2.5-Pro in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for MiMo-V2.5-Pro in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the lineup, see every MiMo model you can run locally, or browse the full model catalogue for something smaller.
MiMo-V2.5-Pro license
MiMo-V2.5-Pro is released under the MIT license, one of the most permissive there is. It allows commercial use, modification and redistribution with no royalties; the only obligation is keeping the copyright and license notice with the code. Xiaomi also serves the model on its own MiMo API platform, but the MIT terms leave you free to run the weights yourself, quantized or not.
