Qwen3-Next-80B-A3B-Thinking

Qwen

Parameters

80.0B total / 3.0B active

MoE: total / active

Architecture

Mixture-of-Experts Transformer with Thinking mode

Released

10.09.2025

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Qwen3-Next en zh multi

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

1,010,000 tokens

Openness Index Score 100.0/100

About

Qwen3-Next-80B-A3B-Thinking is the first reasoning model of the Qwen3-Next series, Alibaba Qwen's next-generation foundation models built for ultimate training and inference efficiency. It is a causal Mixture-of-Experts Transformer with a hybrid attention layout that replaces standard attention by interleaving Gated DeltaNet (a gated linear-attention component) with Gated Attention in a 12 x (3 DeltaNet : 1 Gated Attention) pattern, enabling efficient context modeling for ultra-long context. The high-sparsity MoE activates only 3B of its 80B total parameters per token (512 experts, 10 routed + 1 shared, expert intermediate dimension 512, hidden dimension 2048, 48 layers), drastically reducing FLOPs per token while preserving model capacity.

The model was pretrained on 15T tokens and post-trained with GSPO (Group Sequence Policy Optimization), the RL algorithm that made the hybrid attention + high-sparsity MoE combination stable to train. Multi-Token Prediction (MTP) boosts pretraining performance and accelerates inference via speculative decoding. Qwen3-Next-80B-A3B-Thinking supports a native context of 262,144 tokens, extensible up to 1,010,000 tokens with YaRN RoPE scaling (factor 4.0). It runs in thinking mode only, targeting highly complex reasoning tasks, and serves English and Chinese plus other languages (multilingual benchmarks: MultiIF, MMLU-ProX, INCLUDE, PolyMATH). It surpasses Qwen3-30B-A3B-Thinking-2507, Qwen3-32B-Thinking and the proprietary Gemini-2.5-Flash-Thinking across multiple benchmarks. Maintained by the Qwen team (Alibaba); weights on HuggingFace and Ollama under an Apache-2.0 license.

Training Data Pre-training on 15T tokens & post-training (GSPO-based RL).

Benchmark Scores

Benchmark Score Date
MMLU-Pro
knowledge
73.23%
10.09.2025
MMLU-Redux
knowledge
86.97%
10.09.2025
GPQA
reasoning
100.00%
10.09.2025
SuperGPQA
knowledge
77.93%
10.09.2025
AIME 2025
stem_reasoning
92.53%
10.09.2025
LiveBench 241125
reasoning
100.00%
10.09.2025
LiveCodeBench v6
stem_reasoning
66.17%
10.09.2025
CFEval
coding_agent
100.00%
10.09.2025
IFEval
instruction_following
89.87%
10.09.2025
Arena-Hard v2
instruction_following
100.00%
10.09.2025
WritingBench
instruction_following
100.00%
10.09.2025
BFCL-V3
general_agent
100.00%
10.09.2025
Tau-Bench Retail
general_agent
100.00%
10.09.2025
Tau-Bench Airline
general_agent
98.21%
10.09.2025
TauBench V2 Retail
general_agent
100.00%
10.09.2025
TauBench V2 Airline
general_agent
100.00%
10.09.2025
TauBench V2 Telecom
general_agent
100.00%
10.09.2025
Multi-IF
instruction_following
100.00%
10.09.2025
MMLU-ProX
multilingual
50.32%
10.09.2025
INCLUDE
multilingual
61.24%
10.09.2025
PolyMATH
multilingual
44.61%
10.09.2025
OJBench
stem_reasoning
39.79%
10.09.2025
HMMT Feb 25
stem_reasoning
53.56%
10.09.2025

HuggingFace Evaluation Results: GPQA Diamond, LEXam

HuggingFace evaluation results for Qwen/Qwen3-Next-80B-A3B-Thinking:

Model Tree: Adapters, Finetunes, Merges, Quantizations

Model tree for Qwen/Qwen3-Next-80B-A3B-Thinking:

  • Adapters: 5 models
  • Finetunes: 16 models
  • Merges: 2 models
  • Quantizations: 49 models (llama.cpp, LM Studio, Jan, Ollama)
  • Spaces using this model: 100
  • Collection: Qwen3-Next (4 items)

Citation (Qwen Team, 2025)

Best Practices: Sampling Parameters, Output Length, Output Format, Thinking in History

Processing Ultra-Long Texts: YaRN RoPE Scaling

Agentic Use: Qwen-Agent Integration

Deployment: SGLang and vLLM Serving

Quickstart: Transformers Text Generation

Thinking-Mode-Only Behavior

Qwen3-Next-80B-A3B-Thinking supports only thinking mode. To enforce model thinking, the default chat template automatically includes (end-of-thinking marker). Therefore it is normal for the model's output to contain only without an explicit opening `` tag.

The model may generate thinking content longer than its predecessor. Qwen strongly recommends its use in highly complex reasoning tasks. In multi-turn conversations, the historical model output should only include the final output part and not the thinking content (handled by the provided Jinja2 chat template).

Performance: Benchmarks vs Qwen3-30B-A3B-Thinking-2507, Qwen3-32B, Qwen3-235B-A22B-Thinking-2507, Gemini-2.5-Flash-Thinking

Model Overview: Hybrid Attention Layout and MoE Configuration

Highlights: Hybrid Attention, High-Sparsity MoE, Stability Optimizations, MTP

Architecture

Decoder Block ×48 input Embedding vocab 152K · d 2048 Full Attention Hybrid 16:2 · dₕ 256 ×12 Linear / Recurrent Hybrid 16:2 · dₕ 256 ×36 MoE FFN 512 experts · top-10 · dᴻ 512 Final Norm LM Head vocab 152K output
Attention
Hybrid Attention (16:2)
MoE
512 experts · top-10 per token
Layers
48
Hidden size
2048
Context
262K tokens
RoPE θ
10M
Parameters
80000M
Active params
3000M

Source: Hugging Face config.json · Qwen3NextForCausalLM · exact layer pattern · model repo

Type: Hybrid MoE Causal LM (Gated DeltaNet + Gated Attention)
Attention: Hybrid Gated DeltaNet + Gated Attention (12x 3:1 layout)
Decoder: 48-layer hybrid decoder
MoE: yes (512 experts)
Routing: Top-k routing 10 routed + 1 shared expert
Layers 48
Context length 262K
Extended context 1M
Experts 512
Experts per token 11
Head dim 256
Hidden size 2048
Expert FFN dim 512
Vision No
Attn Heads Kv 2
Attn Heads Q 16
Delta Head Dim 128
Delta Qk Heads 16
Delta V Heads 32
Hybrid Layout 12 * (3 * (Gated DeltaNet -> MoE) -> 1 * (Gated Attention -> MoE))
MTP Yes
Non Embedding Params 79000M
RoPE dim 64

Training Pipeline

  1. 1
    pretraining

    Pre-training on 15T tokens

    Pretrained on 15T tokens with Multi-Token Prediction (MTP) to boost pretraining performance and accelerate inference; zero-centered and weight-decayed layernorm for stability.

  2. 2
    rl

    RL post-training with GSPO

    Post-training with Group Sequence Policy Optimization (GSPO), addressing the stability and efficiency challenges of the hybrid attention + high-sparsity MoE combination in RL training.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
GPQA (Diamond) evaluation — —
LEXam evaluation — —

Trend Analysis

24h Change

+0.2%

Current

101,994

huggingface

downloads

+2.7%

huggingface

likes

+0.2%

huggingface

downloads_all_time

+0.1%

ollama

downloads

+0.1%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 584,700 pulls daily 01.09.2026
huggingface downloads_all_time 2,465,890 daily 01.09.2026
huggingface followers 101,994 daily 01.09.2026
huggingface likes 495 daily 01.09.2026
huggingface downloads 43,899 daily 01.09.2026
ollama downloads 584,300 pulls daily 31.08.2026
huggingface downloads_all_time 2,464,304 daily 31.08.2026
huggingface followers 101,772 daily 31.08.2026
huggingface likes 494 daily 31.08.2026
huggingface downloads 42,757 daily 31.08.2026
ollama downloads 583,900 pulls daily 30.08.2026
huggingface downloads_all_time 2,463,514 daily 30.08.2026
huggingface followers 101,528 daily 30.08.2026
huggingface likes 494 daily 30.08.2026
huggingface downloads 42,925 daily 30.08.2026
ollama downloads 583,600 pulls daily 29.08.2026
ollama downloads 583,300 pulls daily 28.08.2026
ollama downloads 583,000 pulls daily 27.08.2026
ollama downloads 582,700 pulls daily 26.08.2026
ollama downloads 582,500 pulls daily 25.08.2026
ollama downloads 582,100 pulls daily 24.08.2026

View full metric history →

Related Models