Parameters
80.0B total / 3.0B active
MoE: total / active
Architecture
Mixture-of-Experts Transformer with Thinking mode
Released
10.09.2025
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
1,010,000 tokens
About
Qwen3-Next-80B-A3B-Thinking is the first reasoning model of the Qwen3-Next series, Alibaba Qwen's next-generation foundation models built for ultimate training and inference efficiency. It is a causal Mixture-of-Experts Transformer with a hybrid attention layout that replaces standard attention by interleaving Gated DeltaNet (a gated linear-attention component) with Gated Attention in a 12 x (3 DeltaNet : 1 Gated Attention) pattern, enabling efficient context modeling for ultra-long context. The high-sparsity MoE activates only 3B of its 80B total parameters per token (512 experts, 10 routed + 1 shared, expert intermediate dimension 512, hidden dimension 2048, 48 layers), drastically reducing FLOPs per token while preserving model capacity.
The model was pretrained on 15T tokens and post-trained with GSPO (Group Sequence Policy Optimization), the RL algorithm that made the hybrid attention + high-sparsity MoE combination stable to train. Multi-Token Prediction (MTP) boosts pretraining performance and accelerates inference via speculative decoding. Qwen3-Next-80B-A3B-Thinking supports a native context of 262,144 tokens, extensible up to 1,010,000 tokens with YaRN RoPE scaling (factor 4.0). It runs in thinking mode only, targeting highly complex reasoning tasks, and serves English and Chinese plus other languages (multilingual benchmarks: MultiIF, MMLU-ProX, INCLUDE, PolyMATH). It surpasses Qwen3-30B-A3B-Thinking-2507, Qwen3-32B-Thinking and the proprietary Gemini-2.5-Flash-Thinking across multiple benchmarks. Maintained by the Qwen team (Alibaba); weights on HuggingFace and Ollama under an Apache-2.0 license.
Training Data Pre-training on 15T tokens & post-training (GSPO-based RL).
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
MMLU-Pro
knowledge
|
73.23%
|
10.09.2025 |
|
MMLU-Redux
knowledge
|
86.97%
|
10.09.2025 |
|
GPQA
reasoning
|
100.00%
|
10.09.2025 |
|
SuperGPQA
knowledge
|
77.93%
|
10.09.2025 |
|
AIME 2025
stem_reasoning
|
92.53%
|
10.09.2025 |
|
LiveBench 241125
reasoning
|
100.00%
|
10.09.2025 |
|
LiveCodeBench v6
stem_reasoning
|
66.17%
|
10.09.2025 |
|
CFEval
coding_agent
|
100.00%
|
10.09.2025 |
|
IFEval
instruction_following
|
89.87%
|
10.09.2025 |
|
Arena-Hard v2
instruction_following
|
100.00%
|
10.09.2025 |
|
WritingBench
instruction_following
|
100.00%
|
10.09.2025 |
|
BFCL-V3
general_agent
|
100.00%
|
10.09.2025 |
|
Tau-Bench Retail
general_agent
|
100.00%
|
10.09.2025 |
|
Tau-Bench Airline
general_agent
|
98.21%
|
10.09.2025 |
|
TauBench V2 Retail
general_agent
|
100.00%
|
10.09.2025 |
|
TauBench V2 Airline
general_agent
|
100.00%
|
10.09.2025 |
|
TauBench V2 Telecom
general_agent
|
100.00%
|
10.09.2025 |
|
Multi-IF
instruction_following
|
100.00%
|
10.09.2025 |
|
MMLU-ProX
multilingual
|
50.32%
|
10.09.2025 |
|
INCLUDE
multilingual
|
61.24%
|
10.09.2025 |
|
PolyMATH
multilingual
|
44.61%
|
10.09.2025 |
|
OJBench
stem_reasoning
|
39.79%
|
10.09.2025 |
|
HMMT Feb 25
stem_reasoning
|
53.56%
|
10.09.2025 |
HuggingFace Evaluation Results: GPQA Diamond, LEXam
HuggingFace evaluation results for Qwen/Qwen3-Next-80B-A3B-Thinking:
- GPQA Diamond (Idavidrein/gpqa): 73.4
- LEXam Open Question (LEXam-Benchmark/LEXam, max 128B params): 43.37
- LEXam MCQ 4 Choices (max 128B params): 43.31
Model Tree: Adapters, Finetunes, Merges, Quantizations
Model tree for Qwen/Qwen3-Next-80B-A3B-Thinking:
- Adapters: 5 models
- Finetunes: 16 models
- Merges: 2 models
- Quantizations: 49 models (llama.cpp, LM Studio, Jan, Ollama)
- Spaces using this model: 100
- Collection: Qwen3-Next (4 items)
Citation (Qwen Team, 2025)
Best Practices: Sampling Parameters, Output Length, Output Format, Thinking in History
Processing Ultra-Long Texts: YaRN RoPE Scaling
Agentic Use: Qwen-Agent Integration
Deployment: SGLang and vLLM Serving
Quickstart: Transformers Text Generation
Thinking-Mode-Only Behavior
Qwen3-Next-80B-A3B-Thinking supports only thinking mode. To enforce model thinking, the default chat template automatically includes (end-of-thinking marker). Therefore it is normal for the model's output to contain only without an explicit opening `` tag.
The model may generate thinking content longer than its predecessor. Qwen strongly recommends its use in highly complex reasoning tasks. In multi-turn conversations, the historical model output should only include the final output part and not the thinking content (handled by the provided Jinja2 chat template).
Performance: Benchmarks vs Qwen3-30B-A3B-Thinking-2507, Qwen3-32B, Qwen3-235B-A22B-Thinking-2507, Gemini-2.5-Flash-Thinking
Model Overview: Hybrid Attention Layout and MoE Configuration
Highlights: Hybrid Attention, High-Sparsity MoE, Stability Optimizations, MTP
Architecture
- Attention
- Hybrid Attention (16:2)
- MoE
- 512 experts · top-10 per token
- Layers
- 48
- Hidden size
- 2048
- Context
- 262K tokens
- RoPE θ
- 10M
- Parameters
- 80000M
- Active params
- 3000M
Source: Hugging Face config.json · Qwen3NextForCausalLM · exact layer pattern · model repo
Training Pipeline
-
1
pretraining
Pre-training on 15T tokens
Pretrained on 15T tokens with Multi-Token Prediction (MTP) to boost pretraining performance and accelerate inference; zero-centered and weight-decayed layernorm for stability.
-
2
rl
RL post-training with GSPO
Post-training with Group Sequence Policy Optimization (GSPO), addressing the stability and efficiency challenges of the hybrid attention + high-sparsity MoE combination in RL training.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| GPQA (Diamond) | evaluation | — | — | |
| LEXam | evaluation | — | — |
Linked Resources
Qwen3-Next: Towards Ultimate Training & Inference Efficiency
https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd&from=research.latest-advancements-list
Qwen3 Technical Report
https://arxiv.org/abs/2505.09388
Qwen2.5-1M Technical Report
https://arxiv.org/abs/2501.15383
YaRN: Efficient Context Window Extension of Large Language Models
https://arxiv.org/abs/2309.00071
Qwen3-Next HuggingFace collection
https://huggingface.co/collections/Qwen/qwen3-next
BibTeX citation (Qwen Team, 2025)
https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking
Trend Analysis
24h Change
+0.2%
Current
101,994
downloads
+2.7%
likes
+0.2%
downloads_all_time
+0.1%
downloads
+0.1%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 584,700 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 2,465,890 | daily | 01.09.2026 |
| huggingface | followers | 101,994 | daily | 01.09.2026 |
| huggingface | likes | 495 | daily | 01.09.2026 |
| huggingface | downloads | 43,899 | daily | 01.09.2026 |
| ollama | downloads | 584,300 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 2,464,304 | daily | 31.08.2026 |
| huggingface | followers | 101,772 | daily | 31.08.2026 |
| huggingface | likes | 494 | daily | 31.08.2026 |
| huggingface | downloads | 42,757 | daily | 31.08.2026 |
| ollama | downloads | 583,900 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 2,463,514 | daily | 30.08.2026 |
| huggingface | followers | 101,528 | daily | 30.08.2026 |
| huggingface | likes | 494 | daily | 30.08.2026 |
| huggingface | downloads | 42,925 | daily | 30.08.2026 |
| ollama | downloads | 583,600 pulls | daily | 29.08.2026 |
| ollama | downloads | 583,300 pulls | daily | 28.08.2026 |
| ollama | downloads | 583,000 pulls | daily | 27.08.2026 |
| ollama | downloads | 582,700 pulls | daily | 26.08.2026 |
| ollama | downloads | 582,500 pulls | daily | 25.08.2026 |
| ollama | downloads | 582,100 pulls | daily | 24.08.2026 |