What is gpt-oss-120b?
gpt-oss-120b is the larger of OpenAI's two open-weight gpt-oss models, a Mixture of Experts with 117B total parameters and 5.1B active per token. OpenAI positions it for production, general purpose and high reasoning use, sized to fit a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. The weights were published on Hugging Face on August 4, 2025 under Apache 2.0, alongside the smaller gpt-oss-20b, which OpenAI aims at lower latency and local use cases.
| Specification | gpt-oss-120b |
|---|---|
| Total parameters | 117B |
| Active parameters | 5.1B per token |
| Architecture | Mixture of Experts |
| Native quantization | MXFP4 on the MoE weights, applied in post-training |
| Reasoning | Configurable effort: low, medium or high, set in the system prompt |
| Chain of thought | Fully exposed, not intended for end users |
| Chat format | harmony response format, required |
| Modalities | Text in, text out |
| Release date | August 4, 2025 |
| License | Apache 2.0 |
Two details separate this release from most open weights. First, the MoE weights were post-trained with MXFP4 quantization, a 4-bit format, which is what makes a 117B model fit in 80 GB. OpenAI ran all of its evals on that same MXFP4 build, so the quantized model is the model, not a lossy copy of it. Second, both gpt-oss models were trained on the harmony response format and, in OpenAI's words, will not work correctly without it. A runtime that applies the chat template handles harmony for you; call the model directly and you have to apply the format yourself, through the chat template or OpenAI's openai-harmony package.
What gpt-oss-120b is good at
OpenAI publishes no benchmark table on the model card, so the honest summary is the list of stated capabilities. The model is built for agentic work: function calling with defined schemas, web browsing through built-in browsing tools, Python code execution, and Structured Outputs. OpenAI singles out browser tasks as the kind of agentic operation these models are meant for.
Reasoning effort is a setting rather than a separate model. You put a line like "Reasoning: high" in the system prompt and pick one of three levels: low for fast responses in general dialogue, medium for balanced speed and detail, high for deep and detailed analysis. The chain of thought is fully exposed rather than hidden, which OpenAI frames as a way to debug the model and trust its output; it is not intended to be shown to end users. The model is also fine-tunable, and OpenAI states that gpt-oss-120b can be fine-tuned on a single H100 node, while the 20B fits on consumer hardware.
gpt-oss-120b hardware requirements
The system requirement to check is memory, and OpenAI names one figure: a single 80 GB GPU, an NVIDIA H100 or AMD MI300X. Community GGUF builds are published as unsloth/gpt-oss-120b-GGUF, where F16 is a single file and every quant is a two-part split.
| Memory | Build to pick | File size |
|---|---|---|
| 80 GB | F16 | 65.37 GB |
One row covers it, because the ladder in that repo is unusually flat. The smallest build, Q3_K_S, is 62.56 GB across its two parts, and the largest quant, UD-Q8_K_XL, is 64.47 GB, so single-file F16 at 65.37 GB sits less than 3 GB above the bottom of the ladder. The expert weights already ship in 4-bit MXFP4, which leaves the lower quants almost nothing to shrink. When two builds both fit, take the larger one, and on the 80 GB OpenAI names, every build here fits. If the format is new to you, start with what GGUF is.
How to run gpt-oss-120b in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for gpt-oss-120b in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
If 65 GB of weights is more than your machine holds, the smaller gpt-oss-20b runs within 16 GB of memory, and the rest of the lineup is on our gpt-oss page.
gpt-oss-120b license
gpt-oss-120b is released under Apache 2.0. OpenAI describes it on the model card as a permissive license you can build freely on, with no copyleft restrictions and no patent risk, and names experimentation, customization and commercial deployment as what it is meant for.
