GPT-OSS-20B

OpenAI

Parameters

20.9B total / 3.6B active

MoE: total / active

Architecture

Mixture-of-Experts (MoE) Transformer

Released

04.08.2025

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16, U8 (MXFP4) GPT-OSS en

Input Modalities

text

Output Modalities

text

Context (native)

131,072 tokens

Context (extended)

131,072 tokens

Openness Index Score 100.0/100

About

gpt-oss-20b (openai/gpt-oss-20b) is OpenAI's smaller open-weight model, released August 4, 2025 under Apache 2.0 alongside gpt-oss-120b - a 20.91B-parameter Mixture-of-Experts transformer with 3.61B activated parameters per token, a 131,072-token context, and text-only input/output. Its MoE weights are post-trained with MXFP4 quantization, letting the model run within 16GB of memory; all reported evals use that same quantization.

It shares the 120b feature set: configurable reasoning effort (low, medium, high) tuned to latency needs, fully exposed chain-of-thought (for debugging and trust, not for end users), native agentic capabilities (function calling, web browsing, Python code execution, Structured Outputs), and full fine-tunability. It was trained on a text-only dataset of trillions of tokens focused on STEM, coding and general knowledge (knowledge cutoff June 2024), with harmful content filtered via CBRN pre-training filters from GPT-4o; training required ~10x fewer H100-hours than gpt-oss-120b. Model card paper: arXiv 2508.10925.

Training Data Trained on a text-only dataset with trillions of tokens, with a focus on STEM, coding, and general knowledge. Knowledge cutoff: June 2024. Pre-training data filtered for harmful content using CBRN pre-training filters from GPT-4o. Training required ~10x fewer H100-hours than gpt-oss-120b.

Benchmark Scores

Benchmark Score Date
AIME 2024 (with tools)
stem_reasoning
100.00%
05.08.2025
AIME 2025
stem_reasoning
98.73%
05.08.2025
AIME 2025 (with tools)
stem_reasoning
100.00%
05.08.2025
GPQA Diamond
stem_reasoning
56.89%
05.08.2025
GPQA Diamond (with tools)
stem_reasoning
100.00%
05.08.2025
Humanity's Last Exam
stem_reasoning
16.67%
05.08.2025
HLE (with tools)
stem_reasoning
17.30
05.08.2025
MMLU
knowledge
84.26%
05.08.2025
SWE-bench Verified
coding_agent
69.15%
05.08.2025
Tau-Bench Retail
general_agent
54.80
05.08.2025
Tau-Bench Airline
general_agent
38.00
05.08.2025
Aider Polyglot
coding_agent
34.20
05.08.2025
AIME 2024
stem_reasoning
92.10
05.08.2025
MMMLU
multilingual
36.24%
05.08.2025
HealthBench
general_capabilities
42.50
05.08.2025
HealthBench Hard
general_capabilities
10.80
05.08.2025
HealthBench Consensus
general_capabilities
82.60
05.08.2025
CodeForces
stem_reasoning
63.08%
05.08.2025
CodeForces (with tools)
stem_reasoning
100.00%
05.08.2025

Model Tree, Spaces and Paper

Model tree for openai/gpt-oss-20b

Adapters

458 models

Finetunes

552 models

Merges

7 models

Quantizations

235 models

Spaces using openai/gpt-oss-20b 100

Collection including openai/gpt-oss-20b

[

gpt-oss

Collection

Open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. • 2 items • Updated Aug 7, 2025 • 472

](https://huggingface.co/collections/openai/gpt-oss)

Paper for openai/gpt-oss-20b

[

gpt-oss-120b & gpt-oss-20b Model Card

Paper • 2508.10925 • Published Aug 8, 2025 • 29

](https://huggingface.co/papers/2508.10925)

Citation

Citation

@misc{openai2025gptoss120bgptoss20bmodel,
      title={gpt-oss-120b & gpt-oss-20b Model Card}, 
      author={OpenAI},
      year={2025},
      eprint={2508.10925},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2508.10925}, 
}

Safetensors

Model size

21B params

Tensor type

BF16

·

U8

·

Fine-Tuning

Fine-tuning

Both gpt-oss models can be fine-tuned for a variety of specialized use cases.

This smaller model gpt-oss-20b can be fine-tuned on consumer hardware, whereas the larger gpt-oss-120b can be fine-tuned on a single H100 node.

Tool Use

Tool use

The gpt-oss models are excellent for:

  • Web browsing (using built-in browsing tools)
  • Function calling with defined schemas
  • Agentic operations like browser tasks

Reasoning Levels

Reasoning levels

You can adjust the reasoning level that suits your task across three levels:

  • Low: Fast responses for general dialogue.
  • Medium: Balanced speed and detail.
  • High: Deep and detailed analysis.

The reasoning level can be set in the system prompts, e.g., "Reasoning: high".

Inference Examples: Ollama

Ollama

If you are trying to run gpt-oss on consumer hardware, you can use Ollama by running the following commands after installing Ollama.

## gpt-oss-20b
ollama pull gpt-oss:20b
ollama run gpt-oss:20b

Learn more about how to use gpt-oss with Ollama.

LM Studio

If you are using LM Studio you can use the following commands to download.

## gpt-oss-20b
lms get openai/gpt-oss-20b

Check out our awesome list for a broader collection of gpt-oss resources and inference partners.


Download the model

You can download the model weights from the Hugging Face Hub directly from Hugging Face CLI:

## gpt-oss-20b
huggingface-cli download openai/gpt-oss-20b --include "original/*" --local-dir gpt-oss-20b/
pip install gpt-oss
python -m gpt_oss.chat model/

Inference Examples: PyTorch / Triton

PyTorch / Triton

To learn about how to use this model with PyTorch and Triton, check out our reference implementations in the gpt-oss repository.

Inference Examples: vLLM

vLLM

vLLM recommends using uv for Python dependency management. You can use vLLM to spin up an OpenAI-compatible webserver. The following command will automatically download the model and start the server.

uv pip install --pre vllm==0.10.1+gptoss \
    --extra-index-url https://wheels.vllm.ai/gpt-oss/ \
    --extra-index-url https://download.pytorch.org/whl/nightly/cu128 \
    --index-strategy unsafe-best-match

vllm serve openai/gpt-oss-20b

Learn more about how to use gpt-oss with vLLM.

Inference Examples: Transformers

Transformers

You can use gpt-oss-120b and gpt-oss-20b with Transformers. If you use the Transformers chat template, it will automatically apply the harmony response format. If you use model.generate directly, you need to apply the harmony format manually using the chat template or use our openai-harmony package.

To get started, install the necessary dependencies to setup your environment:

pip install -U transformers kernels torch 

Once, setup you can proceed to run the model by running the snippet below:

from transformers import pipeline
import torch

model_id = "openai/gpt-oss-20b"

pipe = pipeline(
    "text-generation",
    model=model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Explain quantum mechanics clearly and concisely."},
]

outputs = pipe(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])

Alternatively, you can run the model via Transformers Serve to spin up a OpenAI-compatible webserver:

transformers serve
transformers chat localhost:8000 --model-name-or-path openai/gpt-oss-20b

Learn more about how to use gpt-oss with Transformers.

Highlights (Apache 2.0, reasoning effort, MXFP4)

Highlights

  • Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk—ideal for experimentation, customization, and commercial deployment.
  • Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs.
  • Full chain-of-thought: Gain complete access to the model’s reasoning process, facilitating easier debugging and increased trust in outputs. It’s not intended to be shown to end users.
  • Fine-tunable: Fully customize models to your specific use case through parameter fine-tuning.
  • Agentic capabilities: Use the models’ native capabilities for function calling, web browsing, Python code execution, and Structured Outputs.
  • MXFP4 quantization: The models were post-trained with MXFP4 quantization of the MoE weights, making gpt-oss-120b run on a single 80GB GPU (like NVIDIA H100 or AMD MI300X) and the gpt-oss-20b model run within 16GB of memory. All evals were performed with the same MXFP4 quantization.

Architecture

Decoder Block ×24 input Embedding vocab 201K · d 2880 Sliding Window Attn Hybrid 64:8 · dₕ 64 · win 128 ×12 Full Attention Hybrid 64:8 · dₕ 64 · win 128 ×12 MoE FFN 32 experts · top-4 · dᴻ 2880 Final Norm LM Head vocab 201K output
Attention
Hybrid Attention (64:8)
MoE
32 experts · top-4 per token
Layers
24
Hidden size
2880
Context
131K tokens
RoPE θ
150K
Parameters
20910M
Active params
3610M

Source: Hugging Face config.json · GptOssForCausalLM · exact layer pattern · model repo

Type: Mixture-of-Experts (MoE) Transformer
Attention: Grouped Query Attention (GQA) with RoPE, alternating banded window (128 tokens) and dense attention, learned bias in softmax (attention sinks)
Decoder: Autoregressive decoder, Pre-LN with RMSNorm
MoE: yes (32 experts)
Routing: Linear router with top-4 expert selection, softmax weighting over selected experts
Layers 24
Context length 131K
Extended context 131K
Experts 32
Experts per token 4
Experts per token 4
Attention heads 64
KV heads 8
Hidden size 2880
Vocabulary 201K
Checkpoint size 12.8
RoPE dim 64

Training Pipeline

  1. 1
    pretraining

    Pretraining

  2. 2
    sft

    Post-Training: Distillation and SFT

  3. 3
    rl

    Reinforcement Learning for Reasoning and Tool Use

  4. 4
    other

    Safety Training: Deliberative Alignment

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Text-only pre-training corpus (STEM, coding, general knowledge) pretraining — —

Trend Analysis

24h Change

+0.1%

7d Change

+1.0%

Current

40,435

huggingface

downloads

+1.1%

huggingface

likes

+0.0%

ollama

downloads

+0.8%

huggingface

downloads_all_time

+0.3%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 12,500,000 pulls daily 01.09.2026
huggingface downloads_all_time 94,001,122 daily 01.09.2026
huggingface followers 40,435 daily 01.09.2026
huggingface likes 4,974 daily 01.09.2026
huggingface downloads 6,519,052 daily 01.09.2026
ollama downloads 12,400,000 pulls daily 31.08.2026
huggingface downloads_all_time 93,740,312 daily 31.08.2026
huggingface followers 40,378 daily 31.08.2026
huggingface likes 4,973 daily 31.08.2026
huggingface downloads 6,450,692 daily 31.08.2026
ollama downloads 12,400,000 pulls daily 30.08.2026
huggingface downloads_all_time 93,563,232 daily 30.08.2026
huggingface followers 40,323 daily 30.08.2026
huggingface likes 4,969 daily 30.08.2026
huggingface downloads 6,520,972 daily 30.08.2026
ollama downloads 12,400,000 pulls daily 29.08.2026
huggingface followers 40,265 daily 29.08.2026
huggingface likes 4,967 daily 29.08.2026
huggingface downloads 6,434,202 daily 29.08.2026
ollama downloads 12,300,000 pulls daily 28.08.2026
huggingface followers 40,216 daily 28.08.2026
huggingface likes 4,963 daily 28.08.2026
huggingface downloads 6,346,634 daily 28.08.2026
ollama downloads 12,300,000 pulls daily 27.08.2026
huggingface followers 40,157 daily 27.08.2026
huggingface likes 4,957 daily 27.08.2026
huggingface downloads 6,702,545 daily 27.08.2026
ollama downloads 12,200,000 pulls daily 26.08.2026
huggingface followers 40,076 daily 26.08.2026
huggingface likes 4,955 daily 26.08.2026
huggingface downloads 6,950,594 daily 26.08.2026
ollama downloads 12,200,000 pulls daily 25.08.2026
huggingface followers 40,023 daily 25.08.2026
huggingface likes 4,954 daily 25.08.2026
huggingface downloads 7,056,476 daily 25.08.2026
ollama downloads 12,200,000 pulls daily 24.08.2026
huggingface followers 39,962 daily 24.08.2026
huggingface likes 4,949 daily 24.08.2026
huggingface downloads 7,036,240 daily 24.08.2026

View full metric history →

Related Models