What is MiniMax-M3?
MiniMax-M3 is a native multimodal Mixture of Experts model from MiniMax, the lab behind the MiniMax Agent platform and the earlier M2 series. It carries about 428B total parameters but activates only about 23B per token, reads text, images and video, and holds a 1M-token context window. For local use that combination is the point: a frontier-scale model that decodes at the cost of a ~23B one, with an attention design that keeps the huge context affordable. The weights landed on Hugging Face on June 2, 2026 under the MiniMax Community License.
| Specification | MiniMax-M3 |
|---|---|
| Total parameters | ~428B |
| Active parameters | ~23B per token |
| Architecture | Mixture of Experts with MiniMax Sparse Attention (MSA) |
| Context window | 1M tokens |
| Modalities | Text, image and video input, text output |
| Reasoning | Three modes: enabled, adaptive, disabled |
| Release date | June 2, 2026 |
| License | MiniMax Community License |
The attention layer is the distinctive part. MiniMax Sparse Attention (MSA) is a sparse attention operator built for million-token contexts, and MiniMax puts numbers on it: at 1M context, M3 prefills 9x and decodes 15x faster than M2, with per-token compute cut to 1/20. Reasoning is controlled by a thinking parameter with three modes: enabled always reasons, adaptive lets the model decide when extra thinking pays off, and disabled skips it for the lowest latency. MiniMax recommends temperature 1.0 and top_p 0.95.
What MiniMax-M3 is good at
MiniMax pitches M3 at coding and cowork: long-horizon agentic work where a model plans, calls tools and iterates over many steps. The model card claims frontier-level performance across long-horizon agentic benchmarks, but publishes the comparison as a chart image rather than a numbers table, so there are no exact scores to reprint here. The second pillar is native multimodality: M3 went through mixed-modality training from the very first step instead of getting a vision encoder bolted on later, which MiniMax credits for deeper semantic fusion across text, image and video.
The third is long context. The MSA efficiency figures are measured at the full 1M window, so the context is built to be used, not just advertised. Outside Atomic Chat, MiniMax recommends SGLang, vLLM, Transformers and KTransformers for serving, plus Unsloth for GGUF and ATOM for MXFP4/MXFP8.
MiniMax-M3 hardware requirements
The system requirement to check is memory, and a 428B model needs a lot of it. The GGUF builds MiniMax points to live at unsloth/MiniMax-M3-GGUF; each build ships as a set of ~50 GB shards, and the sizes below are the totals per build.
| Memory | Build to pick | File size |
|---|---|---|
| 160 GB | UD-IQ2_M | 134.2 GB |
| 192 GB | UD-IQ3_XXS | 159.4 GB |
| 256 GB | UD-IQ4_NL | 211.8 GB |
| 384 GB | UD-Q5_K_XL | 318.5 GB |
| 512 GB and up | Q8_0 | 452.7 GB |
The smallest build in the repo, UD-IQ1_M, totals 128.4 GB, so a 128 GB machine misses the floor once the system and context cache take their share. When two builds both fit, take the larger one, but leave real headroom: a context window this large costs memory on top of the weights. If the format is new to you, start with what GGUF is.
How to run MiniMax-M3 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for MiniMax-M3 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
If your machine is not in the table above, see every MiniMax model you can run locally, or the previous generation MiniMax-M2.7.
MiniMax-M3 license
MiniMax-M3 ships under the MiniMax Community License, the vendor's own terms rather than a standard open-source license; Hugging Face files it under "other". The weights are open to download and the model card documents local deployment directly, but the exact conditions live in the LICENSE file of the repo, so read it before building a commercial product on top of the model.
