What is MiniMax-M2.5?
MiniMax-M2.5 is an open-weight model from MiniMax with 228.7B total parameters, trained with reinforcement learning in hundreds of thousands of complex real-world environments. MiniMax pitches it as SOTA in coding, agentic tool use and search, and office work, serves it natively at 100 tokens per second, and prices it so that an hour of continuous generation at that rate costs $1. The weights landed on Hugging Face on February 12, 2026 under a Modified MIT license, and quantized GGUF builds start at 55.8 GB.
| Specification | MiniMax-M2.5 |
|---|---|
| Total parameters | 228.7B |
| Modalities | Text in, text out |
| Versions | M2.5 (50 tokens/s) and M2.5-Lightning (100 tokens/s), identical capability |
| Training | Reinforcement learning across 200,000+ real-world environments |
| API pricing | $0.30 per million input tokens, $2.40 per million output (Lightning); M2.5 costs half |
| Recommended sampling | temperature 1.0, top_p 0.95, top_k 40 |
| Release date | February 12, 2026 |
| License | Modified MIT |
The coding training is the distinctive part. M2.5 learned in more than 200,000 environments spanning over ten languages, including Go, C++, TypeScript, Rust, Kotlin, Python, Java and PHP, and MiniMax trained it across the whole lifecycle: system design and environment setup, feature iteration, code review and system testing, on Web, Android, iOS and Windows projects with server-side APIs and databases rather than frontend demos alone. One habit emerged during training: before writing code the model decomposes the project and plans features, structure and UI the way an experienced software architect would. It is quick in practice too. An average SWE-bench Verified task takes it 22.8 minutes, on par with Claude Opus 4.6 at 22.9 minutes, at about a tenth of the cost per task.
MiniMax-M2.5 benchmarks
MiniMax's headline numbers are agentic: 80.2 on SWE-bench Verified, 51.3 on Multi-SWE-Bench and 76.3 on BrowseComp with context management. The appendix of the model card adds a general-capability table against Claude, Gemini and GPT models:
| Benchmark | MiniMax-M2.5 | MiniMax-M2.1 | Claude Sonnet 4.5 | Claude Opus 4.5 | Claude Opus 4.6 | Gemini 3 Pro | GPT-5.2 (thinking) |
|---|---|---|---|---|---|---|---|
AIME25 Competition math | 86.3 | 83.0 | 88.0 | 91.0 | 95.6 | 96.0 | 98.0 |
GPQA-D Expert science | 85.2 | 83.0 | 83.0 | 87.0 | 90.0 | 91.0 | 90.0 |
HLE w/o tools Expert questions | 19.4 | 22.2 | 17.3 | 28.4 | 30.7 | 37.2 | 31.4 |
SciCode Scientific coding | 44.4 | 41.0 | 45.0 | 50.0 | 52.0 | 56.0 | 52.0 |
IFBench Instruction following | 70.0 | 70.0 | 57.0 | 58.0 | 53.0 | 70.0 | 75.0 |
AA-LCR Long-context reasoning | 69.5 | 62.0 | 66.0 | 74.0 | 71.0 | 71.0 | 73.0 |
M2.5 tops none of these rows, but it is not last either: it beats all three Claude models on IFBench and ties Gemini 3 Pro there at 70.0, and it is ahead of Claude Sonnet 4.5 on GPQA-D, HLE without tools and AA-LCR. The exam scores are not the sales pitch. The coding section of the same model card reports two results that MiniMax does lead on, both on SWE-bench Verified run under different coding agent harnesses: 79.7 for M2.5 against 78.9 for Claude Opus 4.6 on Droid, and 76.1 against 75.9 on OpenCode.
MiniMax-M2.5 hardware requirements
The system requirement to check is memory. The builds below come from the unsloth/MiniMax-M2.5-GGUF repo; most ship as multi-part files, so the sizes listed are the totals of all parts. File size covers the weights only, so the memory column leaves headroom for the KV cache and the rest of the system.
| Memory | Build to pick | File size |
|---|---|---|
| 96 GB | UD-TQ1_0 | 55.8 GB |
| 128 GB | UD-Q2_K_XL | 85.9 GB |
| 192 GB | UD-Q4_K_XL | 131.3 GB |
| 256 GB | Q5_K_M | 162.3 GB |
| 384 GB and up | Q8_0 | 243.1 GB |
When two builds both fit, take the larger one; the 1-bit and 2-bit files give up the most quality, so move up as soon as your memory allows. If you serve the original weights instead, MiniMax recommends SGLang, vLLM, Transformers or KTransformers. New to quantized files? Start with what GGUF is.
How to run MiniMax-M2.5 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for MiniMax-M2.5 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the lineup, see every MiniMax model you can run locally, including the sibling page for MiniMax-M3.
MiniMax-M2.5 license
MiniMax-M2.5 ships under a Modified MIT license; the exact text lives in the LICENSE file of the MiniMax-M2.5 GitHub repository. The MIT base permits commercial use, modification and redistribution without royalties; read the modified terms before you build a product on it.
