GPT-OSS-120B

OpenAI

Parameters

116.8B total / 5.1B active

MoE: total / active

Architecture

Mixture-of-Experts (MoE) Transformer

Released

04.08.2025

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16, U8 (MXFP4) GPT-OSS en

Input Modalities

text

Output Modalities

text

Context (native)

131,072 tokens

Context (extended)

131,072 tokens

Openness Index Score 100.0/100

About

gpt-oss-120b (openai/gpt-oss-120b) is OpenAI's larger open-weight model, released August 4, 2025 under Apache 2.0 - a 116.83B-parameter Mixture-of-Experts transformer with 5.13B activated parameters per token, a 131,072-token context, and text-only input/output. The MoE weights are post-trained with MXFP4 quantization, so the model runs on a single 80GB GPU (NVIDIA H100 or AMD MI300X); all reported evals use that same MXFP4 quantization.

It offers configurable reasoning effort (low, medium, high) tuned to latency needs, exposes its full chain-of-thought (intended for debugging and trust, not for end users), and carries native agentic capabilities: function calling, web browsing, Python code execution and Structured Outputs. It was trained on a text-only dataset of trillions of tokens focused on STEM, coding and general knowledge (knowledge cutoff June 2024), with pre-training data filtered for harmful content using CBRN pre-training filters from GPT-4o. The models are fully fine-tunable. Model card paper: arXiv 2508.10925.

Training Data Trained on a text-only dataset with trillions of tokens, with a focus on STEM, coding, and general knowledge. Knowledge cutoff: June 2024. Pre-training data filtered for harmful content using CBRN pre-training filters from GPT-4o.

Benchmark Scores

Benchmark Score Date
AIME 2024
stem_reasoning
100.00%
08.08.2025
AIME 2025
stem_reasoning
100.00%
08.08.2025
GPQA Diamond
stem_reasoning
73.15%
08.08.2025
Humanity's Last Exam
stem_reasoning
24.24%
08.08.2025
HLE (with tools)
stem_reasoning
3.60%
08.08.2025
MMLU
knowledge
99.67%
08.08.2025
SWE-bench Verified
coding_agent
71.10%
08.08.2025
Tau-Bench Retail
general_agent
87.84%
08.08.2025
Tau-Bench Airline
general_agent
100.00%
08.08.2025
Aider Polyglot
coding_agent
100.00%
08.08.2025
MMMLU
multilingual
60.70%
08.08.2025
HealthBench Hard
general_capabilities
100.00%
08.08.2025
HealthBench Consensus
general_capabilities
100.00%
08.08.2025
CodeForces
stem_reasoning
70.01%
08.08.2025
HealthBench
general_capabilities
85.31%
08.08.2025

Model Tree, Spaces and Paper

Model tree for openai/gpt-oss-120b

Adapters

220 models

Finetunes

111 models

Merges

1 model

Quantizations

131 models

Spaces using openai/gpt-oss-120b 100

Collection including openai/gpt-oss-120b

[

gpt-oss

Collection

Open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. • 2 items • Updated Aug 7, 2025 • 472

](https://huggingface.co/collections/openai/gpt-oss)

Paper for openai/gpt-oss-120b

[

gpt-oss-120b & gpt-oss-20b Model Card

Paper • 2508.10925 • Published Aug 8, 2025 • 29

](https://huggingface.co/papers/2508.10925)

Citation

Citation

@misc{openai2025gptoss120bgptoss20bmodel,
      title={gpt-oss-120b & gpt-oss-20b Model Card}, 
      author={OpenAI},
      year={2025},
      eprint={2508.10925},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2508.10925}, 
}

Safetensors

Model size

117B params

Tensor type

BF16

·

U8

·

Fine-Tuning

Fine-tuning

Both gpt-oss models can be fine-tuned for a variety of specialized use cases.

This larger model gpt-oss-120b can be fine-tuned on a single H100 node, whereas the smaller gpt-oss-20b can even be fine-tuned on consumer hardware.

Tool Use

Tool use

The gpt-oss models are excellent for:

  • Web browsing (using built-in browsing tools)
  • Function calling with defined schemas
  • Agentic operations like browser tasks

Reasoning Levels

Reasoning levels

You can adjust the reasoning level that suits your task across three levels:

  • Low: Fast responses for general dialogue.
  • Medium: Balanced speed and detail.
  • High: Deep and detailed analysis.

The reasoning level can be set in the system prompts, e.g., "Reasoning: high".

Inference Examples: Ollama

Ollama

If you are trying to run gpt-oss on consumer hardware, you can use Ollama by running the following commands after installing Ollama.

## gpt-oss-120b
ollama pull gpt-oss:120b
ollama run gpt-oss:120b

Learn more about how to use gpt-oss with Ollama.

LM Studio

If you are using LM Studio you can use the following commands to download.

## gpt-oss-120b
lms get openai/gpt-oss-120b

Check out our awesome list for a broader collection of gpt-oss resources and inference partners.


Download the model

You can download the model weights from the Hugging Face Hub directly from Hugging Face CLI:

## gpt-oss-120b
huggingface-cli download openai/gpt-oss-120b --include "original/*" --local-dir gpt-oss-120b/
pip install gpt-oss
python -m gpt_oss.chat model/

Inference Examples: PyTorch / Triton

PyTorch / Triton

To learn about how to use this model with PyTorch and Triton, check out our reference implementations in the gpt-oss repository.

Inference Examples: vLLM

vLLM

vLLM recommends using uv for Python dependency management. You can use vLLM to spin up an OpenAI-compatible webserver. The following command will automatically download the model and start the server.

uv pip install --pre vllm==0.10.1+gptoss \
    --extra-index-url https://wheels.vllm.ai/gpt-oss/ \
    --extra-index-url https://download.pytorch.org/whl/nightly/cu128 \
    --index-strategy unsafe-best-match

vllm serve openai/gpt-oss-120b

Learn more about how to use gpt-oss with vLLM.

Inference Examples: Transformers

Transformers

You can use gpt-oss-120b and gpt-oss-20b with Transformers. If you use the Transformers chat template, it will automatically apply the harmony response format. If you use model.generate directly, you need to apply the harmony format manually using the chat template or use our openai-harmony package.

To get started, install the necessary dependencies to setup your environment:

pip install -U transformers kernels torch 

Once, setup you can proceed to run the model by running the snippet below:

from transformers import pipeline
import torch

model_id = "openai/gpt-oss-120b"

pipe = pipeline(
    "text-generation",
    model=model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Explain quantum mechanics clearly and concisely."},
]

outputs = pipe(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])

Alternatively, you can run the model via Transformers Serve to spin up a OpenAI-compatible webserver:

transformers serve
transformers chat localhost:8000 --model-name-or-path openai/gpt-oss-120b

Learn more about how to use gpt-oss with Transformers.

Highlights (Apache 2.0, reasoning effort, MXFP4)

Highlights

  • Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk—ideal for experimentation, customization, and commercial deployment.
  • Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs.
  • Full chain-of-thought: Gain complete access to the model’s reasoning process, facilitating easier debugging and increased trust in outputs. It’s not intended to be shown to end users.
  • Fine-tunable: Fully customize models to your specific use case through parameter fine-tuning.
  • Agentic capabilities: Use the models’ native capabilities for function calling, web browsing, Python code execution, and Structured Outputs.
  • MXFP4 quantization: The models were post-trained with MXFP4 quantization of the MoE weights, making gpt-oss-120b run on a single 80GB GPU (like NVIDIA H100 or AMD MI300X) and the gpt-oss-20b model run within 16GB of memory. All evals were performed with the same MXFP4 quantization.

Architecture

Decoder Block ×36 input Embedding vocab 201K · d 2880 Sliding Window Attn Hybrid 64:8 · dₕ 64 · win 128 ×18 Full Attention Hybrid 64:8 · dₕ 64 · win 128 ×18 MoE FFN 128 experts · top-4 · dᴻ 2880 Final Norm LM Head vocab 201K output
Attention
Hybrid Attention (64:8)
MoE
128 experts · top-4 per token
Layers
36
Hidden size
2880
Context
131K tokens
RoPE θ
150K
Parameters
116830M
Active params
5130M

Source: Hugging Face config.json · GptOssForCausalLM · exact layer pattern · model repo

Type: Mixture-of-Experts (MoE) Transformer
Attention: Grouped Query Attention (GQA) with RoPE, alternating banded window (128 tokens) and dense attention, learned bias in softmax (attention sinks)
Decoder: Autoregressive decoder, Pre-LN with RMSNorm
MoE: yes (128 experts)
Routing: Linear router with top-4 expert selection, softmax weighting over selected experts
Layers 36
Context length 131K
Extended context 131K
Experts 128
Experts per token 4
Experts per token 4
Attention heads 64
KV heads 8
Hidden size 2880
Vocabulary 201K
Checkpoint size 60.8
RoPE dim 64

Training Pipeline

  1. 1
    pretraining

    Pretraining

    Trained on a text-only dataset with trillions of tokens, focusing on STEM, coding, and general knowledge. Knowledge cutoff: June 2024. Data filtered for harmful content using CBRN pre-training filters from GPT-4o. Training used NVIDIA H100 GPUs with PyTorch framework and expert-optimized Triton kernels. Required 2.1 million H100-hours. Uses Flash Attention algorithms. o200k_harmony tokenizer (BPE, 201,088 tokens).

  2. 2
    sft

    Post-Training: Distillation and SFT

    Post-training uses similar CoT RL techniques as OpenAI o3. Models trained on harmony chat format with System > Developer > User > Assistant > Tool instruction hierarchy. Training dataset covers coding, math, science, and more.

  3. 3
    rl

    Reinforcement Learning for Reasoning and Tool Use

    CoT RL techniques teach models to reason and solve problems using chain-of-thought. Variable effort reasoning training supports low, medium, high reasoning levels. Agentic tool use training includes browsing tool (search and open), Python tool (stateful Jupyter notebook), and arbitrary developer functions with structured outputs.

  4. 4
    other

    Safety Training: Deliberative Alignment

    Deliberative alignment training teaches models to refuse disallowed content, be robust to jailbreaks, and adhere to instruction hierarchy. CBRN pre-training filters applied. Safety training follows OpenAI safety policies by default.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Text-only pre-training corpus (STEM, coding, general knowledge) pretraining — —

Trend Analysis

24h Change

+0.1%

7d Change

+1.0%

Current

40,435

huggingface

downloads

+1.1%

huggingface

likes

+0.1%

ollama

downloads

+0.8%

huggingface

downloads_all_time

+0.3%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 12,500,000 pulls daily 01.09.2026
huggingface downloads_all_time 53,799,428 daily 01.09.2026
huggingface followers 40,435 daily 01.09.2026
huggingface likes 5,141 daily 01.09.2026
huggingface downloads 5,376,994 daily 01.09.2026
ollama downloads 12,400,000 pulls daily 31.08.2026
huggingface downloads_all_time 53,628,561 daily 31.08.2026
huggingface followers 40,378 daily 31.08.2026
huggingface likes 5,136 daily 31.08.2026
huggingface downloads 5,316,895 daily 31.08.2026
ollama downloads 12,400,000 pulls daily 30.08.2026
huggingface downloads_all_time 53,510,645 daily 30.08.2026
huggingface followers 40,323 daily 30.08.2026
huggingface likes 5,136 daily 30.08.2026
huggingface downloads 5,314,695 daily 30.08.2026
ollama downloads 12,400,000 pulls daily 29.08.2026
huggingface followers 40,265 daily 29.08.2026
huggingface likes 5,133 daily 29.08.2026
huggingface downloads 5,173,458 daily 29.08.2026
ollama downloads 12,300,000 pulls daily 28.08.2026
huggingface followers 40,216 daily 28.08.2026
huggingface likes 5,130 daily 28.08.2026
huggingface downloads 5,023,485 daily 28.08.2026
ollama downloads 12,300,000 pulls daily 27.08.2026
huggingface followers 40,157 daily 27.08.2026
huggingface likes 5,128 daily 27.08.2026
huggingface downloads 5,173,082 daily 27.08.2026
ollama downloads 12,200,000 pulls daily 26.08.2026
huggingface followers 40,076 daily 26.08.2026
huggingface likes 5,126 daily 26.08.2026
huggingface downloads 5,222,602 daily 26.08.2026
ollama downloads 12,200,000 pulls daily 25.08.2026
huggingface followers 40,023 daily 25.08.2026
huggingface likes 5,125 daily 25.08.2026
huggingface downloads 5,121,083 daily 25.08.2026
ollama downloads 12,200,000 pulls daily 24.08.2026
huggingface followers 39,962 daily 24.08.2026
huggingface likes 5,123 daily 24.08.2026
huggingface downloads 5,013,805 daily 24.08.2026

View full metric history →

Related Models