GPT-OSS-120B

OpenAI

Parameters

116.8B total / 5.1B active

MoE: total / active

Architecture

Mixture-of-Experts (MoE) Transformer

Released

04.08.2025

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16, U8 (MXFP4) GPT-OSS en

Input Modalities

text

Output Modalities

text

Context (native)

131,072 tokens

Context (extended)

131,072 tokens

Openness Index Score 100.0/100

About

gpt-oss-120b (openai/gpt-oss-120b) is OpenAI's larger open-weight model, released August 4, 2025 under Apache 2.0 - a 116.83B-parameter Mixture-of-Experts transformer with 5.13B activated parameters per token, a 131,072-token context, and text-only input/output. The MoE weights are post-trained with MXFP4 quantization, so the model runs on a single 80GB GPU (NVIDIA H100 or AMD MI300X); all reported evals use that same MXFP4 quantization.

It offers configurable reasoning effort (low, medium, high) tuned to latency needs, exposes its full chain-of-thought (intended for debugging and trust, not for end users), and carries native agentic capabilities: function calling, web browsing, Python code execution and Structured Outputs. It was trained on a text-only dataset of trillions of tokens focused on STEM, coding and general knowledge (knowledge cutoff June 2024), with pre-training data filtered for harmful content using CBRN pre-training filters from GPT-4o. The models are fully fine-tunable. Model card paper: arXiv 2508.10925.

Training Data Trained on a text-only dataset with trillions of tokens, with a focus on STEM, coding, and general knowledge. Knowledge cutoff: June 2024. Pre-training data filtered for harmful content using CBRN pre-training filters from GPT-4o.

Benchmark Scores

Benchmark Score Date
AIME 2024
stem_reasoning
100.00%
08.08.2025
GPQA Diamond
stem_reasoning
73.15%
08.08.2025
Humanity's Last Exam
stem_reasoning
24.24%
08.08.2025
HLE (with tools)
stem_reasoning
3.60%
08.08.2025
MMLU
knowledge
94.41%
08.08.2025
SWE-bench Verified
coding_agent
71.10%
08.08.2025
Tau-Bench Retail
general_agent
87.84%
08.08.2025
Tau-Bench Airline
general_agent
100.00%
08.08.2025
Aider Polyglot
coding_agent
100.00%
08.08.2025
MMMLU
multilingual
60.70%
08.08.2025
HealthBench Hard
general_capabilities
100.00%
08.08.2025
HealthBench Consensus
general_capabilities
100.00%
08.08.2025
CodeForces
stem_reasoning
70.01%
08.08.2025
HealthBench
general_capabilities
85.31%
08.08.2025
AIME 2025
stem_reasoning
93.46%
08.08.2025

Architecture

Decoder Block ×36 input Embedding vocab 201K · d 2880 Sliding Window Attn Hybrid 64:8 · dₕ 64 · win 128 ×18 Full Attention Hybrid 64:8 · dₕ 64 · win 128 ×18 MoE FFN 128 experts · top-4 · dᴻ 2880 Final Norm LM Head vocab 201K output
Attention
Hybrid Attention (64:8)
MoE
128 experts · top-4 per token
Layers
36
Hidden size
2880
Context
131K tokens
RoPE θ
150K
Parameters
116830M
Active params
5130M

Source: Hugging Face config.json · GptOssForCausalLM · exact layer pattern · model repo

Type: Mixture-of-Experts (MoE) Transformer
Attention: Grouped Query Attention (GQA) with RoPE, alternating banded window (128 tokens) and dense attention, learned bias in softmax (attention sinks)
Decoder: Autoregressive decoder, Pre-LN with RMSNorm
MoE: yes (128 experts)
Routing: Linear router with top-4 expert selection, softmax weighting over selected experts
Layers 36
Context length 131K
Extended context 131K
Experts 128
Experts per token 4
Experts per token 4
Attention heads 64
KV heads 8
Hidden size 2880
Vocabulary 201K
Checkpoint size 60.8
RoPE dim 64

Training Pipeline

  1. 1
    pretraining

    Pretraining

    Trained on a text-only dataset with trillions of tokens, focusing on STEM, coding, and general knowledge. Knowledge cutoff: June 2024. Data filtered for harmful content using CBRN pre-training filters from GPT-4o. Training used NVIDIA H100 GPUs with PyTorch framework and expert-optimized Triton kernels. Required 2.1 million H100-hours. Uses Flash Attention algorithms. o200k_harmony tokenizer (BPE, 201,088 tokens).

  2. 2
    sft

    Post-Training: Distillation and SFT

    Post-training uses similar CoT RL techniques as OpenAI o3. Models trained on harmony chat format with System > Developer > User > Assistant > Tool instruction hierarchy. Training dataset covers coding, math, science, and more.

  3. 3
    rl

    Reinforcement Learning for Reasoning and Tool Use

    CoT RL techniques teach models to reason and solve problems using chain-of-thought. Variable effort reasoning training supports low, medium, high reasoning levels. Agentic tool use training includes browsing tool (search and open), Python tool (stateful Jupyter notebook), and arbitrary developer functions with structured outputs.

  4. 4
    other

    Safety Training: Deliberative Alignment

    Deliberative alignment training teaches models to refuse disallowed content, be robust to jailbreaks, and adhere to instruction hierarchy. CBRN pre-training filters applied. Safety training follows OpenAI safety policies by default.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Text-only pre-training corpus (STEM, coding, general knowledge) pretraining — —

Trend Analysis

24h Change

+0.1%

7d Change

+1.0%

Current

40,435

huggingface

downloads

+1.1%

huggingface

likes

+0.1%

ollama

downloads

+0.8%

huggingface

downloads_all_time

+0.3%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 12,500,000 pulls daily 01.09.2026
huggingface downloads_all_time 53,799,428 daily 01.09.2026
huggingface followers 40,435 daily 01.09.2026
huggingface likes 5,141 daily 01.09.2026
huggingface downloads 5,376,994 daily 01.09.2026
ollama downloads 12,400,000 pulls daily 31.08.2026
huggingface downloads_all_time 53,628,561 daily 31.08.2026
huggingface followers 40,378 daily 31.08.2026
huggingface likes 5,136 daily 31.08.2026
huggingface downloads 5,316,895 daily 31.08.2026
ollama downloads 12,400,000 pulls daily 30.08.2026
huggingface downloads_all_time 53,510,645 daily 30.08.2026
huggingface followers 40,323 daily 30.08.2026
huggingface likes 5,136 daily 30.08.2026
huggingface downloads 5,314,695 daily 30.08.2026
ollama downloads 12,400,000 pulls daily 29.08.2026
huggingface followers 40,265 daily 29.08.2026
huggingface likes 5,133 daily 29.08.2026
huggingface downloads 5,173,458 daily 29.08.2026
ollama downloads 12,300,000 pulls daily 28.08.2026
huggingface followers 40,216 daily 28.08.2026
huggingface likes 5,130 daily 28.08.2026
huggingface downloads 5,023,485 daily 28.08.2026
ollama downloads 12,300,000 pulls daily 27.08.2026
huggingface followers 40,157 daily 27.08.2026
huggingface likes 5,128 daily 27.08.2026
huggingface downloads 5,173,082 daily 27.08.2026
ollama downloads 12,200,000 pulls daily 26.08.2026
huggingface followers 40,076 daily 26.08.2026
huggingface likes 5,126 daily 26.08.2026
huggingface downloads 5,222,602 daily 26.08.2026
ollama downloads 12,200,000 pulls daily 25.08.2026
huggingface followers 40,023 daily 25.08.2026
huggingface likes 5,125 daily 25.08.2026
huggingface downloads 5,121,083 daily 25.08.2026
ollama downloads 12,200,000 pulls daily 24.08.2026
huggingface followers 39,962 daily 24.08.2026
huggingface likes 5,123 daily 24.08.2026
huggingface downloads 5,013,805 daily 24.08.2026

View full metric history →

Related Models