Parameters
125.0B total / 6.0B active
MoE: total / active
Architecture
Hybrid Attention (Gated DeltaNet + Qwen Sparse Attention) with MoE, N-gram Embedding, and Gated Residual
Released
26.08.2026
License
Qwen Community License 1.0
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
1,000,000 tokens
About
Qwen3.8-Flash-Next (Qwen/Qwen3.8-Flash-Next) is the first open-weight release of the Qwen3.8 architecture - a 125B-parameter unified vision-language model (text, image and video input) with 6B activated parameters plus 51B n-gram embedding and 4B MTP parameters, released August 26, 2026 under the Qwen Community License 1.0.
It introduces three architecture changes over Qwen3.5: Qwen Sparse Attention (QSA) replaces Gated Attention in the hybrid pairing (Gated DeltaNet + QSA) and selects 512-token micro-blocks rather than individual tokens (24 Q / 2 KV heads, head dim 256, MQA indexer with 4 query + 1 shared key head of dim 128, budget 512 blocks / 2048 tokens), cutting long-context latency for agentic workloads; Gated Residual modulates widened residual streams with an element-wise data-dependent read gate and per-branch scalar write gates (4 branches, bottleneck rank 320) for finer cross-layer expressiveness at low inference overhead; and N-gram Embedding scales parameters through 20M bigram/trigram embeddings at layer 2 - more offload-friendly and computation-cheaper than MoE. The 48-layer stack repeats 12x (3 Gated DeltaNet -> MoE + 1 QSA -> MoE) with 512 experts (10 routed + 1 shared per token, expert dim 640), hidden size 2560, 248K vocabulary, one Multi-Token Prediction (MTP) layer trained with multi-steps, and a 262,144-token native context extensible to 1,000,000.
Its training recipe applies the Muon and AdamW optimizers to distinct weight categories, and - guided by refitted scaling laws - eliminates batch-size warmup, starting directly at the target batch size to cut total optimizer steps while supporting larger learning rates.
Training Data Pre-training & post-training (card Model Overview); tailored recipe: Muon optimizer for specific weight categories + AdamW for others, refitted scaling laws, no batch-size warmup (card Highlights). No public token count.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
DeepSWE 1.1
coding_agent
|
76.59%
|
26.08.2026 |
|
SWE-bench Pro
coding_agent
|
78.12%
|
26.08.2026 |
|
NL2Repo-Bench
coding_agent
|
55.44%
|
26.08.2026 |
|
CoWorkBench
general_agent
|
93.51%
|
26.08.2026 |
|
JobBench
general_agent
|
73.38%
|
26.08.2026 |
|
Agents' Last Exam
general_agent
|
64.62%
|
26.08.2026 |
|
Agents' Last Exam (Score)
general_agent
|
100.00%
|
26.08.2026 |
|
Toolathlon Verified
general_agent
|
91.22%
|
26.08.2026 |
|
IFBench
instruction_following
|
97.50%
|
26.08.2026 |
|
GPQA Diamond
stem_reasoning
|
95.08%
|
26.08.2026 |
|
HLE (with tools)
stem_reasoning
|
39.41%
|
26.08.2026 |
|
LiveCodeBench v6
stem_reasoning
|
97.82%
|
26.08.2026 |
|
Claw-Eval Pass^3
coding_agent
|
86.78%
|
26.08.2026 |
|
RecreationBench
general_agent
|
100.00%
|
26.08.2026 |
|
AndroidWorld
general_agent
|
100.00%
|
26.08.2026 |
|
OSWorld 2.0 (Binary)
general_agent
|
100.00%
|
26.08.2026 |
|
OSWorld 2.0 (Partial)
general_agent
|
100.00%
|
26.08.2026 |
|
Vision2Web
vision_language
|
100.00%
|
26.08.2026 |
|
ERQA
spatial_intelligence
|
100.00%
|
26.08.2026 |
|
LVBench
video_understanding
|
100.00%
|
26.08.2026 |
|
RealWorldQA
vision_language
|
100.00%
|
26.08.2026 |
|
MathVision
vision_language
|
86.10%
|
26.08.2026 |
|
CharXiv (RQ)
document_understanding
|
73.52%
|
26.08.2026 |
|
Claw-Eval Avg
coding_agent
|
72.40%
|
26.08.2026 |
|
SWE-bench Multilingual
coding_agent
|
89.94%
|
26.08.2026 |
Benchmark Highlights
Qwen3.8-Flash-Next achieves best-in-class results across multiple benchmark categories with only 6B activated parameters:
Language Benchmarks (best results in bold):
- DeepSWE 1.1: 58.7 (vs 42.2 Qwen3.8-27B, 54.4 DeepSeek-V4-Flash)
- SWE-bench Pro: 62.5 (vs 61.7 Qwen3.8-27B)
- SWE-bench Multilingual: 81.0 (vs 77.5 Claude-Opus-4.6)
- CoWorkBench: 73.9 (vs 70.7 Qwen3.8-27B)
- JobBench: 55.7 (vs 41.3 DeepSeek-V4-Flash)
- Agents' Last Exam Score: 51.2 (vs 42.9 Qwen3.8-27B)
- Toolathlon Verified: 73.5 (vs 70.3 DeepSeek-V4-Flash)
- IFBench: 81.3 (vs 79.5 Qwen3.8-27B)
- GPQA Diamond: 91.7 (vs 91.3 Claude-Opus-4.6)
- LiveCodeBench v6: 91.9 (vs 90.6 DeepSeek-V4-Flash)
Vision Benchmarks (best results in bold):
- ClawEval-MM Pass@3: 64.4 (vs 57.4 Qwen3.8-27B)
- ClawEval-MM Average: 60.4 (vs 60.1 Qwen3.7-Plus)
- RecreationBench: 49.9 (vs 47.1 Qwen3.8-27B)
- AndroidWorld: 84.5 (vs 81.9 Qwen3.8-27B)
- OSWorld 2.0 Binary: 19.4 (tied with Qwen3.8-27B)
- OSWorld 2.0 Partial: 52.3 (vs 48.0 Qwen3.8-27B)
- Vision2Web: 64.0 (vs 62.9 Qwen3.8-27B)
- ERQA: 72.3 (vs 69.8 Qwen3.7-Plus)
- LVBench: 76.6 (vs 76.2 Qwen3.7-Plus)
- RealWorldQA: 88.5 (vs 86.9 Qwen3.7-Plus)
- MathVision: 90.6 (vs 90.3 Qwen3.7-Plus)
- CharXiv (RQ): 84.6 (vs 85.8 Qwen3.7-Plus)
Vision and Video Support
Qwen3.8-Flash-Next includes a Vision Encoder and supports image, video, and text inputs.
Input Modalities: text, image, video Output Modalities: text
The model can process:
- Images via image_url content type
- Videos via video_url content type with configurable fps sampling (default fps=2)
- Long videos with recommended video_preprocessor_config longest_edge=469,762,048 (224k video tokens)
Thinking Mode and Reasoning Effort
Qwen3.8-Flash-Next operates in thinking mode by default, generating thinking content signified by \u003cthink\u003e\n...\n\u003c/think\u003e before producing final responses.
Sampling Parameters:
- Thinking Mode: temp=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
- Instruct (non-thinking) Mode: temp=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
Reasoning Effort Levels: xhigh (default), medium, low.
Preserved Thinking: Retains thinking blocks from all historical messages for full context continuity, beneficial for agent scenarios. Can be disabled via preserve_thinking=False.
Context Length and YaRN Scaling
Qwen3.8-Flash-Next natively supports 262,144 tokens context length and is extensible up to 1,000,000 tokens using YaRN (Yet another RoPE extensioN) scaling.
YaRN configuration:
- rope_type: yarn
- rope_theta: 10,000,000
- partial_rotary_factor: 0.25
- mrope_interleaved: true
- mrope_section: [11, 11, 10]
Supported by vLLM, SGLang, and TokenSpeed inference frameworks.
Multi-Token Prediction (MTP)
Qwen3.8-Flash-Next includes a 1-layer Multi-Token Prediction (MTP) module trained with multi-steps, contributing 4B parameters to the total count.
MTP helps the model predict multiple future tokens simultaneously, improving both training efficiency and inference speed.
Tailored Training Recipe
Qwen3.8-Flash-Next uses a tailored training recipe combining the Muon and AdamW optimizers applied to specific weight categories to maximize efficiency.
Key innovations:
- Guided by refitted scaling laws
- Eliminates traditional batch-size warmups, starting directly at the target batch size
- Substantially reduces total optimizer steps
- Safely supports larger learning rates for robust convergence
The model undergoes both pre-training and post-training stages.
Mixture of Experts (MoE)
The MoE component has 512 experts with 10 routed + 1 shared activated experts per token. Expert intermediate dimension is 640.
This is a key efficiency innovation: only 6B parameters are activated per token out of 125B total, making the model extremely efficient at inference while maintaining high quality.
Gated Residual
Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
Configuration: 4 branches, bottleneck rank 320.
N-gram Embedding
N-gram Embedding provides a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams (bigrams/trigrams at layer 2), this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
Scale: 20,000,000 n-gram embeddings, contributing 51B parameters to the total 125B parameter count.
Token Embedding: 248,320 (padded).
Hybrid Attention with QSA
Qwen3.8-Flash-Next introduces Hybrid Attention combining Gated DeltaNet (linear attention) and Qwen Sparse Attention (QSA). Unlike traditional sparse attention that selects individual tokens, QSA operates at the micro-block level, significantly cutting long-context latency. This is critical for agentic workloads that dominate real-world usage.
Architecture Layout: 48 layers arranged as 12 \u00d7 (3 \u00d7 (Gated DeltaNet \u2192 MoE) \u2192 1 \u00d7 (Qwen Sparse Attention \u2192 MoE)).
Gated DeltaNet: 48 V heads, 16 QK heads, head dim 128.
Qwen Sparse Attention: 24 Q heads, 2 KV heads, head dim 256, RoPE dim 64, MQA indexer with 4 query heads and 1 shared key head (indexer head dim 128), budget of 512 blocks or 2048 tokens.
Architecture
- Attention
- Hybrid Attention (24:2)
- MoE
- 512 experts · top-10 per token
- Layers
- 48
- Hidden size
- 2560
- Context
- 262K tokens
- Parameters
- 125000M
- Active params
- 6000M
Source: Hugging Face config.json · Qwen4ExpForConditionalGeneration · exact layer pattern · model repo
Hybrid Attention with QSA (micro-block level sparse attention), Gated Residual (element-wise read gate + per-branch scalar write gate), N-gram Embedding (parameter scaling efficient for memory-constrained accelerators), Tailored Training Recipe (Muon + AdamW, no batch-size warmup)
248K
248K
Muon and AdamW (applied to specific weight categories)
Refitted scaling laws, no batch-size warmup, start at target batch size
Training Pipeline
-
1
pretraining
Pre-training with Tailored Recipe
Pre-training stage using Muon and AdamW optimizers applied to specific weight categories. Guided by refitted scaling laws, eliminates traditional batch-size warmups and starts directly at the target batch size, substantially reducing total optimizer steps while safely supporting larger learning rates for robust convergence.
-
2
other
Post-training
Post-training stage including thinking mode training with multi-token prediction (MTP). The model operates in thinking mode by default, generating thinking content before final responses. Supports reasoning effort levels (xhigh, medium, low) and preserved thinking for agent scenarios.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Muon/AdamW tailored training corpus | pretraining | — | — |
Linked Resources
Qwen3.8-Flash-Next Blog Post
https://qwen.ai/blog?id=qwen3.8-flash-next
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf
QwenLM/Qwen3.8-Flash-Next GitHub Repository
https://github.com/QwenLM/Qwen3.8-Flash-Next
Qwen Cloud - Qwen3.8-Flash Overview
https://www.qwencloud.com/models/Qwen3.8-Flash
SGLang - Qwen3.8-Flash-Next Cookbook
https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-Flash-Next
vLLM - Qwen3.8-Flash-Next Recipe
https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next
TokenSpeed - Qwen3.8-Flash-Next Recipe
https://lightseek.org/tokenspeed/recipes/models#qwen3-8-flash-next
Qwen3.8-Flash-Next Architecture Diagram
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.8-Flash-Next/architecture.png
Qwen3.8-Flash-Next HuggingFace Model Card
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
Trend Analysis
24h Change
+73.0%
Current
32,700 pulls
downloads
+31.1%
downloads_all_time
+31.1%
likes
+2.7%
followers
+0.2%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 32,700 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 207,941 | daily | 01.09.2026 |
| huggingface | followers | 101,994 | daily | 01.09.2026 |
| huggingface | likes | 4,633 | daily | 01.09.2026 |
| huggingface | downloads | 207,941 | daily | 01.09.2026 |
| ollama | downloads | 18,900 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 158,598 | daily | 31.08.2026 |
| huggingface | followers | 101,772 | daily | 31.08.2026 |
| huggingface | likes | 4,513 | daily | 31.08.2026 |
| huggingface | downloads | 158,598 | daily | 31.08.2026 |
| ollama | downloads | 10,800 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 121,976 | daily | 30.08.2026 |
| huggingface | followers | 101,528 | daily | 30.08.2026 |
| huggingface | likes | 4,378 | daily | 30.08.2026 |
| huggingface | downloads | 121,976 | daily | 30.08.2026 |
| huggingface | followers | 101,306 | daily | 29.08.2026 |
| huggingface | likes | 4,285 | daily | 29.08.2026 |
| huggingface | downloads | 52,341 | daily | 29.08.2026 |
| huggingface | followers | 101,117 | daily | 28.08.2026 |
| huggingface | likes | 4,152 | daily | 28.08.2026 |
| huggingface | downloads | 4,810 | daily | 28.08.2026 |
| huggingface | followers | 100,862 | daily | 27.08.2026 |
| huggingface | likes | 3,944 | daily | 27.08.2026 |
| huggingface | downloads | 4,810 | daily | 27.08.2026 |
| huggingface | followers | 100,537 | daily | 26.08.2026 |
| huggingface | likes | 3,587 | daily | 26.08.2026 |
| huggingface | downloads | 2,551 | daily | 26.08.2026 |