MiniMax-M2.5

Updated
24.08.2026
Thinking
Tools
Reasoning
Code
Web

A 229B-parameter (10B active) MoE model from MiniMax built for agentic coding, tool use, and search, with a 200K context window.

At a glance

  • License: Modified MIT
  • Parameters: 228.7B total
  • Context length: Not stated in the MiniMax model card
  • Modalities: Text in, text out
  • Minimum hardware: 96 GB memory, running the smallest GGUF build, UD-TQ1_0 at 55.8 GB

What is MiniMax-M2.5?

MiniMax-M2.5 is an open-weight model from MiniMax with 228.7B total parameters, trained with reinforcement learning in hundreds of thousands of complex real-world environments. MiniMax pitches it as SOTA in coding, agentic tool use and search, and office work, serves it natively at 100 tokens per second, and prices it so that an hour of continuous generation at that rate costs $1. The weights landed on Hugging Face on February 12, 2026 under a Modified MIT license, and quantized GGUF builds start at 55.8 GB.

SpecificationMiniMax-M2.5
Total parameters228.7B
ModalitiesText in, text out
VersionsM2.5 (50 tokens/s) and M2.5-Lightning (100 tokens/s), identical capability
TrainingReinforcement learning across 200,000+ real-world environments
API pricing$0.30 per million input tokens, $2.40 per million output (Lightning); M2.5 costs half
Recommended samplingtemperature 1.0, top_p 0.95, top_k 40
Release dateFebruary 12, 2026
LicenseModified MIT

The coding training is the distinctive part. M2.5 learned in more than 200,000 environments spanning over ten languages, including Go, C++, TypeScript, Rust, Kotlin, Python, Java and PHP, and MiniMax trained it across the whole lifecycle: system design and environment setup, feature iteration, code review and system testing, on Web, Android, iOS and Windows projects with server-side APIs and databases rather than frontend demos alone. One habit emerged during training: before writing code the model decomposes the project and plans features, structure and UI the way an experienced software architect would. It is quick in practice too. An average SWE-bench Verified task takes it 22.8 minutes, on par with Claude Opus 4.6 at 22.9 minutes, at about a tenth of the cost per task.

MiniMax-M2.5 benchmarks

MiniMax's headline numbers are agentic: 80.2 on SWE-bench Verified, 51.3 on Multi-SWE-Bench and 76.3 on BrowseComp with context management. The appendix of the model card adds a general-capability table against Claude, Gemini and GPT models:

BenchmarkMiniMax-M2.5MiniMax-M2.1Claude Sonnet 4.5Claude Opus 4.5Claude Opus 4.6Gemini 3 ProGPT-5.2 (thinking)
AIME25
Competition math
86.383.088.091.095.696.098.0
GPQA-D
Expert science
85.283.083.087.090.091.090.0
HLE w/o tools
Expert questions
19.422.217.328.430.737.231.4
SciCode
Scientific coding
44.441.045.050.052.056.052.0
IFBench
Instruction following
70.070.057.058.053.070.075.0
AA-LCR
Long-context reasoning
69.562.066.074.071.071.073.0

M2.5 tops none of these rows, but it is not last either: it beats all three Claude models on IFBench and ties Gemini 3 Pro there at 70.0, and it is ahead of Claude Sonnet 4.5 on GPQA-D, HLE without tools and AA-LCR. The exam scores are not the sales pitch. The coding section of the same model card reports two results that MiniMax does lead on, both on SWE-bench Verified run under different coding agent harnesses: 79.7 for M2.5 against 78.9 for Claude Opus 4.6 on Droid, and 76.1 against 75.9 on OpenCode.

MiniMax-M2.5 hardware requirements

The system requirement to check is memory. The builds below come from the unsloth/MiniMax-M2.5-GGUF repo; most ship as multi-part files, so the sizes listed are the totals of all parts. File size covers the weights only, so the memory column leaves headroom for the KV cache and the rest of the system.

MemoryBuild to pickFile size
96 GBUD-TQ1_055.8 GB
128 GBUD-Q2_K_XL85.9 GB
192 GBUD-Q4_K_XL131.3 GB
256 GBQ5_K_M162.3 GB
384 GB and upQ8_0243.1 GB

When two builds both fit, take the larger one; the 1-bit and 2-bit files give up the most quality, so move up as soon as your memory allows. If you serve the original weights instead, MiniMax recommends SGLang, vLLM, Transformers or KTransformers. New to quantized files? Start with what GGUF is.

How to run MiniMax-M2.5 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for MiniMax-M2.5 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every MiniMax model you can run locally, including the sibling page for MiniMax-M3.

MiniMax-M2.5 license

MiniMax-M2.5 ships under a Modified MIT license; the exact text lives in the LICENSE file of the MiniMax-M2.5 GitHub repository. The MIT base permits commercial use, modification and redistribution without royalties; read the modified terms before you build a product on it.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download MiniMaxAI/MiniMax-M2.5
# serve with vLLM:
vllm serve MiniMaxAI/MiniMax-M2.5 --trust-remote-code --tensor-parallel-size 2
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMaxAI/MiniMax-M2.5",
    "messages": [{"role": "user", "content": "Refactor this function for readability."}],
    "temperature": 1.0,
    "top_p": 0.95
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("MiniMaxAI/MiniMax-M2.5", trust_remote_code=True, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("MiniMaxAI/MiniMax-M2.5", trust_remote_code=True)
messages = [{"role": "user", "content": "Write a Python function to reverse a linked list."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512, temperature=1.0, top_p=0.95)
print(tokenizer.decode(out[0]))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "EMPTY" });
const res = await client.chat.completions.create({
  model: "MiniMaxAI/MiniMax-M2.5",
  messages: [{ role: "user", content: "Plan the file structure for a CLI todo app." }],
  temperature: 1.0,
  top_p: 0.95,
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

MiniMax-M2.5 is a Mixture-of-Experts (MoE) model with 229B total parameters, of which roughly 10B are active per token. It activates 8 of its experts on each forward pass, which keeps inference cost low while retaining the capacity of a much larger dense model.

Despite activating only 10B parameters, MiniMax-M2.5 still has to hold all 229B weights in memory. The fp8 weights are about 230 GB, so a single 24 GB consumer GPU is not enough. Practical local setups need 96 GB or more of combined VRAM plus system RAM, for example 2x H100 80 GB or several consumer GPUs with CPU offload. Quantized GGUF builds shrink it: a 3-bit quant lands around 101 GB and runs at roughly 25 tokens per second on an 80 GB H100.

The weights are released on Hugging Face under a Modified MIT license, so you can download, run, fine-tune, and use the model commercially. It is open-weight rather than fully open-source, since the training data and full pipeline are not published. MiniMax also offers a hosted API where M2.5 costs about $0.30 per million input tokens and $2.40 per million output tokens.

MiniMax-M2.5 supports a context window of about 204,800 tokens (roughly 200K). That headroom suits long agentic runs, large codebases, and multi-document search tasks. For very long browsing sessions the model is designed to discard history when token usage gets high, which is how MiniMax reports its BrowseComp results.

MiniMax-M2.5 is built for agentic coding and tool use. It reports 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, and was trained across more than 200,000 real-world environments in over ten programming languages. It also handles office tasks such as Word, PowerPoint, and Excel financial modeling. The model thinks step by step before acting and runs natively at up to 100 tokens per second.