What is gpt-oss-20b?
gpt-oss-20b is an open-weight Mixture of Experts model from OpenAI and the smaller of the two models in the gpt-oss release. It carries 21B total parameters with 3.6B active. OpenAI positions it for lower latency and for local or specialized use cases, while the larger gpt-oss-120b, at 117B parameters with 5.1B active, targets production work on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. The MoE weights were post-trained with MXFP4 quantization, which is why the whole model runs within 16 GB of memory as shipped. OpenAI published the weights on August 4, 2025 under Apache 2.0.
| Specification | gpt-oss-20b |
|---|---|
| Total parameters | 21B |
| Active parameters | 3.6B |
| Checkpoint tensors | 20,914,757,184 parameters in safetensors |
| Architecture | Mixture of Experts, MoE weights post-trained in MXFP4 |
| Modalities | Text |
| Reasoning | Configurable effort: low, medium or high, set in the system prompt |
| Response format | harmony, required for correct output |
| Built-in tool use | Function calling, web browsing, Python execution, Structured Outputs |
| Fine-tuning | Supported, and for the 20B on consumer hardware |
| Release date | August 4, 2025 |
| License | Apache 2.0 |
Two details matter in practice. Both gpt-oss models were trained on OpenAI's harmony response format and will not work correctly without it; the chat template in Transformers applies the format automatically, so a normal chat setup already does the right thing. And all of OpenAI's evals were performed with the same MXFP4 quantization, so the MXFP4 checkpoint OpenAI ships is the exact model it measured, not a compressed copy of it.
What gpt-oss-20b is good at
OpenAI built the gpt-oss series for reasoning and agentic tasks. Reasoning effort is configurable straight from the system prompt across three levels: low for fast responses in general dialogue, medium for balanced speed and detail, and high for deep, detailed analysis. A line like "Reasoning: high" is all it takes. The model also exposes its full chain-of-thought, which OpenAI recommends for debugging and for trust in outputs rather than for showing to end users.
The agentic side is native: function calling with defined schemas, web browsing through built-in browsing tools, Python code execution, and Structured Outputs. OpenAI lists web browsing, function calling against defined schemas and agentic browser operations as the tasks these models are excellent for. It also singles out the 20B as the gpt-oss model you can fine-tune on consumer hardware, while the 120B can be fine-tuned on a single H100 node. The published model card for both models is arXiv 2508.10925.
gpt-oss-20b hardware requirements
The system requirement to check is memory. OpenAI states that its own MXFP4 checkpoint runs within 16 GB of memory; it says nothing about third-party GGUF builds, so the tiers below leave room for the KV cache and the rest of the system. The files are real sizes from the unsloth/gpt-oss-20b-GGUF repository on Hugging Face.
| Memory | Build to pick | File size |
|---|---|---|
| 16 GB | Q4_K_M | 11.62 GB |
| 24 GB | UD-Q8_K_XL | 13.20 GB |
| 24 GB | F16 | 13.79 GB |
The ladder is unusually flat because the MoE weights already ship in MXFP4: the smallest file in the repository, Q3_K_S, is 11.46 GB, Q2_K is 11.47 GB, Q8_0 is 12.11 GB and full precision F16 is 13.79 GB, a spread of about 2.3 GB across the whole set. Dropping below Q4 buys almost nothing here, and when two builds both fit, take the larger one.
OpenAI names the stacks it supports for running the weights directly: Transformers, including transformers serve for an OpenAI-compatible webserver; vLLM, which downloads the model and starts a server from vllm serve openai/gpt-oss-20b; Ollama, with ollama pull gpt-oss:20b for consumer hardware; LM Studio, with lms get openai/gpt-oss-20b; and reference PyTorch and Triton implementations in the gpt-oss repository. If the GGUF format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for how a 16 GB machine handles models this size.
How to run gpt-oss-20b in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for gpt-oss-20b in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the bigger sibling, see gpt-oss-120b, or browse every gpt-oss model you can run locally.
gpt-oss-20b license
gpt-oss-20b is released under Apache 2.0. OpenAI calls it a permissive license and spells out the intent: build freely, without copyleft restrictions or patent risk, which it describes as ideal for experimentation, customization and commercial deployment. Fine-tuning falls under the same terms, and OpenAI says the 20B can be fine-tuned on consumer hardware.
