MiniMax-M2.7

Updated
24.08.2026
Thinking
Tools
Reasoning
Code

Run MiniMax-M2.7, a 229B MoE model, locally in Atomic Chat. Private, offline reasoning and coding, no API keys, no cloud, no limits. Free.

At a glance

  • License: Custom ("other"), full text in the repo LICENSE file
  • Parameters: 228.7B
  • Release date: April 9, 2026
  • Modalities: Text in, text out
  • Minimum hardware: 64 GB RAM (smallest GGUF build is 60.7 GB)

What is MiniMax-M2.7?

MiniMax-M2.7 is a 228.7B-parameter open-weight model from MiniMax, published on Hugging Face on April 9, 2026. MiniMax describes it as the first of its models to take a hand in its own development: during training it updated its own memory, built dozens of complex skills for RL experiments, and changed its own learning process based on the results. One internal version rewrote a programming scaffold over 100+ rounds, analyzing failure trajectories, editing code, running evaluations, and deciding to keep or revert each change, and came out 30% better for it. The released model is built for the same kind of work: complex agent harnesses, long productivity tasks, native Agent Teams for multi-agent collaboration, and dynamic tool search.

SpecificationMiniMax-M2.7
Total parameters228.7B
ModalitiesText in, text out
Agent featuresAgent Teams, complex Skills, dynamic tool search
Recommended servingSGLang, vLLM or Transformers; also on NVIDIA NIM
Recommended samplingtemperature 1.0, top_p 0.95, top_k 40
Release dateApril 9, 2026
LicenseCustom, listed as "other" on Hugging Face

The card positions M2.7 well beyond code generation. MiniMax claims system-level engineering skills: correlating monitoring metrics, running trace analysis, verifying root causes in databases, and making SRE-level decisions, and says the model has cut live production incident recovery to under three minutes on multiple occasions. On the office side it edits Word, Excel and PPT files over multiple rounds while keeping the deliverables editable, and MiniMax reports 97% skill compliance across 40+ complex skills on MM Claw.

MiniMax-M2.7 benchmarks

All numbers below are MiniMax's own, from the model card. The vendor published no side-by-side competitor table, so the scores stand alone: MLE Bench Lite is a medal rate across 22 ML competitions, GDPval-AA is an ELO rating, the rest are percentages.

BenchmarkMiniMax-M2.7
SWE-Pro
Harder engineering
56.22
SWE Multilingual
Multilingual engineering
76.5
Terminal Bench 2
Terminal agents
57.0
NL2Repo
Repo-level coding
39.8
VIBE-Pro
Vibe coding
55.6
Toolathon
Long-horizon tools
46.3
MLE Bench Lite
ML competitions
66.6
GDPval-AA
Professional tasks
1495

MiniMax's framing for these: SWE-Pro matches GPT-5.3-Codex, VIBE-Pro is nearly on par with Opus 4.6, and the GDPval-AA ELO is the highest among open-weight models, ahead of GPT5.3. The card also concedes ground: the MLE Bench Lite medal rate is second to Opus-4.6 and GPT-5.4, and on the MM Claw end-to-end benchmark its 62.7% lands close to, not above, Sonnet 4.6.

MiniMax-M2.7 hardware requirements

The system requirement to check is memory. The community GGUF builds come from unsloth/MiniMax-M2.7-GGUF, shipped as multi-part files, so the sizes below are per-build totals. The unquantized BF16 conversion is 457.5 GB; the quantized ladder is what makes local use realistic.

MemoryBuild to pickFile size
64 GBUD-IQ1_M60.7 GB
96 GBUD-IQ3_S83.6 GB
128 GBUD-IQ4_XS108.4 GB
192 GBUD-Q5_K_XL169.5 GB
256 GB and upQ8_0243.1 GB

When two builds both fit, take the larger one. That matters most in the 1-bit and 2-bit range, where quality falls fastest per gigabyte saved: on a 96 GB machine UD-IQ3_S at 83.6 GB is worth the squeeze over UD-IQ2_M at 70.2 GB. If quant names and split files are new territory, start with what GGUF is.

How to run MiniMax-M2.7 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for MiniMax-M2.7 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the lineup, see every MiniMax model you can run locally, or the sibling MiniMax M3.

MiniMax-M2.7 license

MiniMax-M2.7 ships under MiniMax's own license: Hugging Face lists it as "other" rather than a standard permissive license, and the full text lives in the LICENSE file of the model repository. Read that file before you build a commercial product on the model; the terms are the vendor's, not Apache or MIT.

Get the weights from Hugging Face

huggingface-cli download MiniMaxAI/MiniMax-M2.7
from transformers import AutoModel
model = AutoModel.from_pretrained("MiniMaxAI/MiniMax-M2.7")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

MiniMax-M2.7 is an open-weight 228.7B-parameter Mixture-of-Experts model from MiniMaxAI, built for coding and agentic workflows. It activates only about 10B parameters per token, supports tool calls and a 200K-token context, and uses <think> tags to separate its reasoning from its final output. You can run it locally in Atomic Chat instead of through a hosted API.

Plan for at least 128GB of combined memory (VRAM plus system RAM for offloading). The full bf16 weights are around 457GB, while a 4-bit GGUF quant drops to roughly 108GB and fits on a 128GB-RAM machine. On Apple Silicon, the MLX build needs a Mac Studio with 128GB or more of unified memory.

The weights are free to download from Hugging Face and free for personal, non-commercial use. Commercial use is different: the license requires prior written authorization from MiniMaxAI and a visible "Built with MiniMax M2.7" attribution. Running it in Atomic Chat on your own machine costs nothing beyond your hardware.

Yes. Once you download the weights, the model runs entirely on your hardware with no internet connection required. In Atomic Chat the model executes on-device, so your prompts, code, and files stay local and are never sent to an external server.

It is built for agentic coding: multi-file edits, code-run-fix loops, and test-validated repairs, plus planning and tool use across shell, browser, retrieval, and code runners. The 200K context window suits large codebases and long documents. Its published capabilities are thinking, reasoning, tool use, and code.