MiniMax-M3

Updated
24.08.2026
Thinking
Vision
Reasoning
Code

Run MiniMax-M3, a 427B MoE model, locally in Atomic Chat. Private, offline long-context reasoning, no API keys, no limits. Free.

At a glance

  • License: MiniMax Community License
  • Parameters: ~428B total, ~23B active (MoE)
  • Context length: 1M tokens
  • Modalities: Text, image and video input
  • Minimum hardware: 160 GB of memory, smallest GGUF totals 128.4 GB

What is MiniMax-M3?

MiniMax-M3 is a native multimodal Mixture of Experts model from MiniMax, the lab behind the MiniMax Agent platform and the earlier M2 series. It carries about 428B total parameters but activates only about 23B per token, reads text, images and video, and holds a 1M-token context window. For local use that combination is the point: a frontier-scale model that decodes at the cost of a ~23B one, with an attention design that keeps the huge context affordable. The weights landed on Hugging Face on June 2, 2026 under the MiniMax Community License.

SpecificationMiniMax-M3
Total parameters~428B
Active parameters~23B per token
ArchitectureMixture of Experts with MiniMax Sparse Attention (MSA)
Context window1M tokens
ModalitiesText, image and video input, text output
ReasoningThree modes: enabled, adaptive, disabled
Release dateJune 2, 2026
LicenseMiniMax Community License

The attention layer is the distinctive part. MiniMax Sparse Attention (MSA) is a sparse attention operator built for million-token contexts, and MiniMax puts numbers on it: at 1M context, M3 prefills 9x and decodes 15x faster than M2, with per-token compute cut to 1/20. Reasoning is controlled by a thinking parameter with three modes: enabled always reasons, adaptive lets the model decide when extra thinking pays off, and disabled skips it for the lowest latency. MiniMax recommends temperature 1.0 and top_p 0.95.

What MiniMax-M3 is good at

MiniMax pitches M3 at coding and cowork: long-horizon agentic work where a model plans, calls tools and iterates over many steps. The model card claims frontier-level performance across long-horizon agentic benchmarks, but publishes the comparison as a chart image rather than a numbers table, so there are no exact scores to reprint here. The second pillar is native multimodality: M3 went through mixed-modality training from the very first step instead of getting a vision encoder bolted on later, which MiniMax credits for deeper semantic fusion across text, image and video.

The third is long context. The MSA efficiency figures are measured at the full 1M window, so the context is built to be used, not just advertised. Outside Atomic Chat, MiniMax recommends SGLang, vLLM, Transformers and KTransformers for serving, plus Unsloth for GGUF and ATOM for MXFP4/MXFP8.

MiniMax-M3 hardware requirements

The system requirement to check is memory, and a 428B model needs a lot of it. The GGUF builds MiniMax points to live at unsloth/MiniMax-M3-GGUF; each build ships as a set of ~50 GB shards, and the sizes below are the totals per build.

MemoryBuild to pickFile size
160 GBUD-IQ2_M134.2 GB
192 GBUD-IQ3_XXS159.4 GB
256 GBUD-IQ4_NL211.8 GB
384 GBUD-Q5_K_XL318.5 GB
512 GB and upQ8_0452.7 GB

The smallest build in the repo, UD-IQ1_M, totals 128.4 GB, so a 128 GB machine misses the floor once the system and context cache take their share. When two builds both fit, take the larger one, but leave real headroom: a context window this large costs memory on top of the weights. If the format is new to you, start with what GGUF is.

How to run MiniMax-M3 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for MiniMax-M3 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

If your machine is not in the table above, see every MiniMax model you can run locally, or the previous generation MiniMax-M2.7.

MiniMax-M3 license

MiniMax-M3 ships under the MiniMax Community License, the vendor's own terms rather than a standard open-source license; Hugging Face files it under "other". The weights are open to download and the model card documents local deployment directly, but the exact conditions live in the LICENSE file of the repo, so read it before building a commercial product on top of the model.

Get the weights from Hugging Face

huggingface-cli download MiniMaxAI/MiniMax-M3
from transformers import AutoModel
model = AutoModel.from_pretrained("MiniMaxAI/MiniMax-M3")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

MiniMax-M3 is an open-weight Mixture-of-Experts language model from MiniMaxAI, with about 427B total parameters and roughly 23B active per token. It is natively multimodal, handles a 1,048,576-token context through MiniMax Sparse Attention, and targets coding, agentic, and reasoning tasks. You can download the weights and run it locally through Atomic Chat.

A lot. The smallest GGUF quant is around 128GB on disk, so plan for at least 130GB or more of RAM or unified memory once you add the KV cache. A 512GB Mac Studio M3 Ultra can run long generations, and GPU servers typically use 8-way tensor parallelism. This is not a model for an average laptop or a single consumer GPU.

The weights are openly available to download and run yourself, so local use through Atomic Chat has no per-token cost. It is released under a custom "other" license, so read MiniMaxAI's terms on Hugging Face before commercial use. MiniMaxAI's own hosted API is separate and priced per token.

Yes. Once you download the weights from Hugging Face, the model runs entirely on your own hardware with no internet connection required. In Atomic Chat your prompts and files stay on-device, which is the main reason to self-host rather than call a cloud API.

Its strongest areas are coding and agentic workflows, where MiniMaxAI reports about 59% on SWE-Bench Pro. The 1M-token context makes it useful for reasoning over whole codebases or large document sets, and its native vision support lets it read screenshots, diagrams, and charts alongside text.