gpt-oss-120b

Updated
05.10.2026
Tools
Thinking
Reasoning
Code

gpt-oss-120b is OpenAI’s open-weight MoE: 117B parameters, 5.1B active, sized for a single 80 GB GPU. Run it locally, free, with Atomic Chat.

At a glance

  • License: Apache 2.0
  • Parameters: 117B total, 5.1B active
  • Context length: Not stated on the vendor model card
  • Modalities: Text in, text out
  • Minimum hardware: Single 80 GB GPU (H100 or MI300X class)

What is gpt-oss-120b?

gpt-oss-120b is the larger of OpenAI's two open-weight gpt-oss models, a Mixture of Experts with 117B total parameters and 5.1B active per token. OpenAI positions it for production, general purpose and high reasoning use, sized to fit a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. The weights were published on Hugging Face on August 4, 2025 under Apache 2.0, alongside the smaller gpt-oss-20b, which OpenAI aims at lower latency and local use cases.

Specificationgpt-oss-120b
Total parameters117B
Active parameters5.1B per token
ArchitectureMixture of Experts
Native quantizationMXFP4 on the MoE weights, applied in post-training
ReasoningConfigurable effort: low, medium or high, set in the system prompt
Chain of thoughtFully exposed, not intended for end users
Chat formatharmony response format, required
ModalitiesText in, text out
Release dateAugust 4, 2025
LicenseApache 2.0

Two details separate this release from most open weights. First, the MoE weights were post-trained with MXFP4 quantization, a 4-bit format, which is what makes a 117B model fit in 80 GB. OpenAI ran all of its evals on that same MXFP4 build, so the quantized model is the model, not a lossy copy of it. Second, both gpt-oss models were trained on the harmony response format and, in OpenAI's words, will not work correctly without it. A runtime that applies the chat template handles harmony for you; call the model directly and you have to apply the format yourself, through the chat template or OpenAI's openai-harmony package.

What gpt-oss-120b is good at

OpenAI publishes no benchmark table on the model card, so the honest summary is the list of stated capabilities. The model is built for agentic work: function calling with defined schemas, web browsing through built-in browsing tools, Python code execution, and Structured Outputs. OpenAI singles out browser tasks as the kind of agentic operation these models are meant for.

Reasoning effort is a setting rather than a separate model. You put a line like "Reasoning: high" in the system prompt and pick one of three levels: low for fast responses in general dialogue, medium for balanced speed and detail, high for deep and detailed analysis. The chain of thought is fully exposed rather than hidden, which OpenAI frames as a way to debug the model and trust its output; it is not intended to be shown to end users. The model is also fine-tunable, and OpenAI states that gpt-oss-120b can be fine-tuned on a single H100 node, while the 20B fits on consumer hardware.

gpt-oss-120b hardware requirements

The system requirement to check is memory, and OpenAI names one figure: a single 80 GB GPU, an NVIDIA H100 or AMD MI300X. Community GGUF builds are published as unsloth/gpt-oss-120b-GGUF, where F16 is a single file and every quant is a two-part split.

MemoryBuild to pickFile size
80 GBF1665.37 GB

One row covers it, because the ladder in that repo is unusually flat. The smallest build, Q3_K_S, is 62.56 GB across its two parts, and the largest quant, UD-Q8_K_XL, is 64.47 GB, so single-file F16 at 65.37 GB sits less than 3 GB above the bottom of the ladder. The expert weights already ship in 4-bit MXFP4, which leaves the lower quants almost nothing to shrink. When two builds both fit, take the larger one, and on the 80 GB OpenAI names, every build here fits. If the format is new to you, start with what GGUF is.

How to run gpt-oss-120b in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for gpt-oss-120b in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

If 65 GB of weights is more than your machine holds, the smaller gpt-oss-20b runs within 16 GB of memory, and the rest of the lineup is on our gpt-oss page.

gpt-oss-120b license

gpt-oss-120b is released under Apache 2.0. OpenAI describes it on the model card as a permissive license you can build freely on, with no copyleft restrictions and no patent risk, and names experimentation, customization and commercial deployment as what it is meant for.

Get the weights from Hugging Face

huggingface-cli download openai/gpt-oss-120b
from transformers import AutoModel
model = AutoModel.from_pretrained("openai/gpt-oss-120b")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

gpt-oss-120b is an open-weight language model from OpenAI built on a Mixture-of-Experts architecture, with about 120B total parameters and roughly 5B active per token. It is tuned for reasoning, tool use, and code, supports a 128K context window, and is released under the Apache-2.0 license so anyone can download and run it.

The practical target is a single GPU with 80 GB of VRAM, such as an NVIDIA H100, since OpenAI ships the model in MXFP4 form at around 61 GiB. You can run it on less VRAM by offloading layers to system RAM, but expect slower generation. Plan for ample system memory as well to handle long contexts.

Yes. The weights are published under the Apache-2.0 license, so the model is free to download, run, and even use commercially. Running it locally in Atomic Chat means there is no API key and no per-token cost.

Yes. Once the weights are downloaded to your machine, gpt-oss-120b runs entirely on-device with no internet connection required. Your prompts and data stay local, which is the point of running it in Atomic Chat instead of a hosted API.

It is strongest at reasoning, agentic tool use, and coding. OpenAI reports it reaches near-parity with o4-mini on core reasoning benchmarks, and its native function calling lets it drive external tools. The 128K context window also helps with long documents and extended code files.