gpt-oss-20b

Updated
05.10.2026
Tools
Thinking
Reasoning
Code

gpt-oss-20b is OpenAI’s open-weight 21B MoE model with 3.6B active parameters, with its MoE weights post-trained in MXFP4 to run within 16 GB of memory.

At a glance

  • License: Apache 2.0
  • Parameters: 21B total, 3.6B active
  • Context length: Not stated in OpenAI's model card
  • Modalities: Text
  • Minimum hardware: 16 GB of memory (MXFP4 build)

What is gpt-oss-20b?

gpt-oss-20b is an open-weight Mixture of Experts model from OpenAI and the smaller of the two models in the gpt-oss release. It carries 21B total parameters with 3.6B active. OpenAI positions it for lower latency and for local or specialized use cases, while the larger gpt-oss-120b, at 117B parameters with 5.1B active, targets production work on a single 80 GB GPU such as an NVIDIA H100 or AMD MI300X. The MoE weights were post-trained with MXFP4 quantization, which is why the whole model runs within 16 GB of memory as shipped. OpenAI published the weights on August 4, 2025 under Apache 2.0.

Specificationgpt-oss-20b
Total parameters21B
Active parameters3.6B
Checkpoint tensors20,914,757,184 parameters in safetensors
ArchitectureMixture of Experts, MoE weights post-trained in MXFP4
ModalitiesText
ReasoningConfigurable effort: low, medium or high, set in the system prompt
Response formatharmony, required for correct output
Built-in tool useFunction calling, web browsing, Python execution, Structured Outputs
Fine-tuningSupported, and for the 20B on consumer hardware
Release dateAugust 4, 2025
LicenseApache 2.0

Two details matter in practice. Both gpt-oss models were trained on OpenAI's harmony response format and will not work correctly without it; the chat template in Transformers applies the format automatically, so a normal chat setup already does the right thing. And all of OpenAI's evals were performed with the same MXFP4 quantization, so the MXFP4 checkpoint OpenAI ships is the exact model it measured, not a compressed copy of it.

What gpt-oss-20b is good at

OpenAI built the gpt-oss series for reasoning and agentic tasks. Reasoning effort is configurable straight from the system prompt across three levels: low for fast responses in general dialogue, medium for balanced speed and detail, and high for deep, detailed analysis. A line like "Reasoning: high" is all it takes. The model also exposes its full chain-of-thought, which OpenAI recommends for debugging and for trust in outputs rather than for showing to end users.

The agentic side is native: function calling with defined schemas, web browsing through built-in browsing tools, Python code execution, and Structured Outputs. OpenAI lists web browsing, function calling against defined schemas and agentic browser operations as the tasks these models are excellent for. It also singles out the 20B as the gpt-oss model you can fine-tune on consumer hardware, while the 120B can be fine-tuned on a single H100 node. The published model card for both models is arXiv 2508.10925.

gpt-oss-20b hardware requirements

The system requirement to check is memory. OpenAI states that its own MXFP4 checkpoint runs within 16 GB of memory; it says nothing about third-party GGUF builds, so the tiers below leave room for the KV cache and the rest of the system. The files are real sizes from the unsloth/gpt-oss-20b-GGUF repository on Hugging Face.

MemoryBuild to pickFile size
16 GBQ4_K_M11.62 GB
24 GBUD-Q8_K_XL13.20 GB
24 GBF1613.79 GB

The ladder is unusually flat because the MoE weights already ship in MXFP4: the smallest file in the repository, Q3_K_S, is 11.46 GB, Q2_K is 11.47 GB, Q8_0 is 12.11 GB and full precision F16 is 13.79 GB, a spread of about 2.3 GB across the whole set. Dropping below Q4 buys almost nothing here, and when two builds both fit, take the larger one.

OpenAI names the stacks it supports for running the weights directly: Transformers, including transformers serve for an OpenAI-compatible webserver; vLLM, which downloads the model and starts a server from vllm serve openai/gpt-oss-20b; Ollama, with ollama pull gpt-oss:20b for consumer hardware; LM Studio, with lms get openai/gpt-oss-20b; and reference PyTorch and Triton implementations in the gpt-oss repository. If the GGUF format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for how a 16 GB machine handles models this size.

How to run gpt-oss-20b in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for gpt-oss-20b in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the bigger sibling, see gpt-oss-120b, or browse every gpt-oss model you can run locally.

gpt-oss-20b license

gpt-oss-20b is released under Apache 2.0. OpenAI calls it a permissive license and spells out the intent: build freely, without copyleft restrictions or patent risk, which it describes as ideal for experimentation, customization and commercial deployment. Fine-tuning falls under the same terms, and OpenAI says the 20B can be fine-tuned on consumer hardware.

Get the weights from Hugging Face

huggingface-cli download openai/gpt-oss-20b
from transformers import AutoModel
model = AutoModel.from_pretrained("openai/gpt-oss-20b")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

gpt-oss-20b is an open-weight language model released by OpenAI in 2025. It uses a Mixture-of-Experts architecture with about 21.5B total parameters (roughly 3.6B active per token) and a 128K-token context window. It is built for local, on-device use with reasoning, tool calling, and code generation.

With its built-in 4-bit MXFP4 quantization, gpt-oss-20b fits in about 16GB of memory. A 16GB card like the RTX 5080 is the practical minimum, and a 24GB GPU such as the RTX 4090 gives more room for longer context. If you don't have a suitable GPU, around 24GB of system RAM lets it run on CPU, though slower.

Yes. gpt-oss-20b is released under the Apache 2.0 license, so the weights are free to download, run, modify, and use commercially. Running it locally through Atomic Chat means there are no API fees or per-token charges.

Yes. Once the weights are downloaded, gpt-oss-20b runs entirely on your own machine and needs no internet connection. Prompts and outputs stay on-device, which keeps your data private. Atomic Chat loads it locally so it keeps working with the network off.

It is strong at chain-of-thought reasoning, tool and function calling, and writing or debugging code. OpenAI reports it performs near o3-mini on common benchmarks while staying small enough for consumer hardware. The 128K context window also lets it handle long documents and larger codebases in a single session.