Qwen3.8-Flash-Next

Qwen Team (Alibaba)

Parameters

125.0B total / 6.0B active

MoE: total / active

Architecture

Hybrid Attention (Gated DeltaNet + Qwen Sparse Attention) with MoE, N-gram Embedding, and Gated Residual

Released

26.08.2026

License

Qwen Community License 1.0

Open Weights Commercial Use Multimodal BF16/I64 Qwen English Chinese Multilingual

Input Modalities

text image video

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 70.0/100

About

Qwen3.8-Flash-Next (Qwen/Qwen3.8-Flash-Next) is the first open-weight release of the Qwen3.8 architecture - a 125B-parameter unified vision-language model (text, image and video input) with 6B activated parameters plus 51B n-gram embedding and 4B MTP parameters, released August 26, 2026 under the Qwen Community License 1.0.

It introduces three architecture changes over Qwen3.5: Qwen Sparse Attention (QSA) replaces Gated Attention in the hybrid pairing (Gated DeltaNet + QSA) and selects 512-token micro-blocks rather than individual tokens (24 Q / 2 KV heads, head dim 256, MQA indexer with 4 query + 1 shared key head of dim 128, budget 512 blocks / 2048 tokens), cutting long-context latency for agentic workloads; Gated Residual modulates widened residual streams with an element-wise data-dependent read gate and per-branch scalar write gates (4 branches, bottleneck rank 320) for finer cross-layer expressiveness at low inference overhead; and N-gram Embedding scales parameters through 20M bigram/trigram embeddings at layer 2 - more offload-friendly and computation-cheaper than MoE. The 48-layer stack repeats 12x (3 Gated DeltaNet -> MoE + 1 QSA -> MoE) with 512 experts (10 routed + 1 shared per token, expert dim 640), hidden size 2560, 248K vocabulary, one Multi-Token Prediction (MTP) layer trained with multi-steps, and a 262,144-token native context extensible to 1,000,000.

Its training recipe applies the Muon and AdamW optimizers to distinct weight categories, and - guided by refitted scaling laws - eliminates batch-size warmup, starting directly at the target batch size to cut total optimizer steps while supporting larger learning rates.

Training Data Pre-training & post-training (card Model Overview); tailored recipe: Muon optimizer for specific weight categories + AdamW for others, refitted scaling laws, no batch-size warmup (card Highlights). No public token count.

Benchmark Scores

Benchmark Score Date
DeepSWE 1.1
coding_agent
76.59%
26.08.2026
SWE-bench Pro
coding_agent
78.12%
26.08.2026
NL2Repo-Bench
coding_agent
55.44%
26.08.2026
CoWorkBench
general_agent
93.51%
26.08.2026
JobBench
general_agent
73.38%
26.08.2026
Agents' Last Exam
general_agent
64.62%
26.08.2026
Agents' Last Exam (Score)
general_agent
100.00%
26.08.2026
Toolathlon Verified
general_agent
91.22%
26.08.2026
IFBench
instruction_following
97.50%
26.08.2026
GPQA Diamond
stem_reasoning
95.08%
26.08.2026
HLE (with tools)
stem_reasoning
39.41%
26.08.2026
LiveCodeBench v6
stem_reasoning
97.82%
26.08.2026
Claw-Eval Pass^3
coding_agent
86.78%
26.08.2026
RecreationBench
general_agent
100.00%
26.08.2026
AndroidWorld
general_agent
100.00%
26.08.2026
OSWorld 2.0 (Binary)
general_agent
100.00%
26.08.2026
OSWorld 2.0 (Partial)
general_agent
100.00%
26.08.2026
Vision2Web
vision_language
100.00%
26.08.2026
ERQA
spatial_intelligence
100.00%
26.08.2026
LVBench
video_understanding
100.00%
26.08.2026
RealWorldQA
vision_language
100.00%
26.08.2026
MathVision
vision_language
86.10%
26.08.2026
CharXiv (RQ)
document_understanding
73.52%
26.08.2026
Claw-Eval Avg
coding_agent
72.40%
26.08.2026
SWE-bench Multilingual
coding_agent
89.94%
26.08.2026

Benchmark Highlights

Qwen3.8-Flash-Next achieves best-in-class results across multiple benchmark categories with only 6B activated parameters:

Language Benchmarks (best results in bold):

  • DeepSWE 1.1: 58.7 (vs 42.2 Qwen3.8-27B, 54.4 DeepSeek-V4-Flash)
  • SWE-bench Pro: 62.5 (vs 61.7 Qwen3.8-27B)
  • SWE-bench Multilingual: 81.0 (vs 77.5 Claude-Opus-4.6)
  • CoWorkBench: 73.9 (vs 70.7 Qwen3.8-27B)
  • JobBench: 55.7 (vs 41.3 DeepSeek-V4-Flash)
  • Agents' Last Exam Score: 51.2 (vs 42.9 Qwen3.8-27B)
  • Toolathlon Verified: 73.5 (vs 70.3 DeepSeek-V4-Flash)
  • IFBench: 81.3 (vs 79.5 Qwen3.8-27B)
  • GPQA Diamond: 91.7 (vs 91.3 Claude-Opus-4.6)
  • LiveCodeBench v6: 91.9 (vs 90.6 DeepSeek-V4-Flash)

Vision Benchmarks (best results in bold):

  • ClawEval-MM Pass@3: 64.4 (vs 57.4 Qwen3.8-27B)
  • ClawEval-MM Average: 60.4 (vs 60.1 Qwen3.7-Plus)
  • RecreationBench: 49.9 (vs 47.1 Qwen3.8-27B)
  • AndroidWorld: 84.5 (vs 81.9 Qwen3.8-27B)
  • OSWorld 2.0 Binary: 19.4 (tied with Qwen3.8-27B)
  • OSWorld 2.0 Partial: 52.3 (vs 48.0 Qwen3.8-27B)
  • Vision2Web: 64.0 (vs 62.9 Qwen3.8-27B)
  • ERQA: 72.3 (vs 69.8 Qwen3.7-Plus)
  • LVBench: 76.6 (vs 76.2 Qwen3.7-Plus)
  • RealWorldQA: 88.5 (vs 86.9 Qwen3.7-Plus)
  • MathVision: 90.6 (vs 90.3 Qwen3.7-Plus)
  • CharXiv (RQ): 84.6 (vs 85.8 Qwen3.7-Plus)

Vision and Video Support

Qwen3.8-Flash-Next includes a Vision Encoder and supports image, video, and text inputs.

Input Modalities: text, image, video Output Modalities: text

The model can process:

  • Images via image_url content type
  • Videos via video_url content type with configurable fps sampling (default fps=2)
  • Long videos with recommended video_preprocessor_config longest_edge=469,762,048 (224k video tokens)

Thinking Mode and Reasoning Effort

Qwen3.8-Flash-Next operates in thinking mode by default, generating thinking content signified by \u003cthink\u003e\n...\n\u003c/think\u003e before producing final responses.

Sampling Parameters:

  • Thinking Mode: temp=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (non-thinking) Mode: temp=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Reasoning Effort Levels: xhigh (default), medium, low.

Preserved Thinking: Retains thinking blocks from all historical messages for full context continuity, beneficial for agent scenarios. Can be disabled via preserve_thinking=False.

Context Length and YaRN Scaling

Qwen3.8-Flash-Next natively supports 262,144 tokens context length and is extensible up to 1,000,000 tokens using YaRN (Yet another RoPE extensioN) scaling.

YaRN configuration:

  • rope_type: yarn
  • rope_theta: 10,000,000
  • partial_rotary_factor: 0.25
  • mrope_interleaved: true
  • mrope_section: [11, 11, 10]

Supported by vLLM, SGLang, and TokenSpeed inference frameworks.

Multi-Token Prediction (MTP)

Qwen3.8-Flash-Next includes a 1-layer Multi-Token Prediction (MTP) module trained with multi-steps, contributing 4B parameters to the total count.

MTP helps the model predict multiple future tokens simultaneously, improving both training efficiency and inference speed.

Tailored Training Recipe

Qwen3.8-Flash-Next uses a tailored training recipe combining the Muon and AdamW optimizers applied to specific weight categories to maximize efficiency.

Key innovations:

  • Guided by refitted scaling laws
  • Eliminates traditional batch-size warmups, starting directly at the target batch size
  • Substantially reduces total optimizer steps
  • Safely supports larger learning rates for robust convergence

The model undergoes both pre-training and post-training stages.

Mixture of Experts (MoE)

The MoE component has 512 experts with 10 routed + 1 shared activated experts per token. Expert intermediate dimension is 640.

This is a key efficiency innovation: only 6B parameters are activated per token out of 125B total, making the model extremely efficient at inference while maintaining high quality.

Gated Residual

Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.

Configuration: 4 branches, bottleneck rank 320.

N-gram Embedding

N-gram Embedding provides a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams (bigrams/trigrams at layer 2), this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.

Scale: 20,000,000 n-gram embeddings, contributing 51B parameters to the total 125B parameter count.

Token Embedding: 248,320 (padded).

Hybrid Attention with QSA

Qwen3.8-Flash-Next introduces Hybrid Attention combining Gated DeltaNet (linear attention) and Qwen Sparse Attention (QSA). Unlike traditional sparse attention that selects individual tokens, QSA operates at the micro-block level, significantly cutting long-context latency. This is critical for agentic workloads that dominate real-world usage.

Architecture Layout: 48 layers arranged as 12 \u00d7 (3 \u00d7 (Gated DeltaNet \u2192 MoE) \u2192 1 \u00d7 (Qwen Sparse Attention \u2192 MoE)).

Gated DeltaNet: 48 V heads, 16 QK heads, head dim 128.

Qwen Sparse Attention: 24 Q heads, 2 KV heads, head dim 256, RoPE dim 64, MQA indexer with 4 query heads and 1 shared key head (indexer head dim 128), budget of 512 blocks or 2048 tokens.

Architecture

Decoder Block ×48 input Embedding vocab 248K · d 2560 Linear / Recurrent Hybrid 24:2 · dₕ 256 ×36 Full Attention Hybrid 24:2 · dₕ 256 ×12 MoE FFN 512 experts · top-10 · dᴻ 640 Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (24:2)
MoE
512 experts · top-10 per token
Layers
48
Hidden size
2560
Context
262K tokens
Parameters
125000M
Active params
6000M

Source: Hugging Face config.json · Qwen4ExpForConditionalGeneration · exact layer pattern · model repo

Type: Hybrid Attention MoE with N-gram Embedding
Attention: Hybrid: Gated DeltaNet (linear attention) + Qwen Sparse Attention (QSA)
Decoder: Causal Language Model with Vision Encoder
MoE: yes (512 experts)
Routing: Top-10 routed + 1 shared expert
Layers 48
Total parameters 125B (6B activated + 51B n-gram + 4B MTP)
Context length 262K
Extended context 1M
Hidden size 2560
Gated Deltanet {'linear_attention_heads_v': 48, 'linear_attention_heads_qk': 16, 'head_dim': 128}
Gated Residual {'num_branches': 4, 'bottleneck_rank': 320}
Hidden Layout 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
Model size 180B params (BF16/I64)
Moe {'num_experts': 512, 'activated_experts': '10 routed + 1 shared', 'expert_intermediate_dim': 640}
MTP 1 layer, trained with multi-steps
Mtp Params 4B
N Gram Embedding 20M
N Gram Embedding Params 51B
N Gram Type bigrams/trigrams at layer 2
Qwen Sparse Attention {'attention_heads_q': 24, 'attention_heads_kv': 2, 'head_dim': 256, 'rope_dim': 64, 'indexer_structure': 'MQA with 4 Query Heads and 1 Shared Key Head', 'indexer_head_dim': 128, 'budget': '512 blocks or 2048 tokens'}
Key Innovations

Hybrid Attention with QSA (micro-block level sparse attention), Gated Residual (element-wise read gate + per-branch scalar write gate), N-gram Embedding (parameter scaling efficient for memory-constrained accelerators), Tailored Training Recipe (Muon + AdamW, no batch-size warmup)

Lm Output

248K

Token Embedding

248K

Training Optimizers

Muon and AdamW (applied to specific weight categories)

Training Recipe

Refitted scaling laws, no batch-size warmup, start at target batch size

Training Pipeline

  1. 1
    pretraining

    Pre-training with Tailored Recipe

    Pre-training stage using Muon and AdamW optimizers applied to specific weight categories. Guided by refitted scaling laws, eliminates traditional batch-size warmups and starts directly at the target batch size, substantially reducing total optimizer steps while safely supporting larger learning rates for robust convergence.

  2. 2
    other

    Post-training

    Post-training stage including thinking mode training with multi-token prediction (MTP). The model operates in thinking mode by default, generating thinking content before final responses. Supports reasoning effort levels (xhigh, medium, low) and preserved thinking for agent scenarios.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Muon/AdamW tailored training corpus pretraining — —

Trend Analysis

24h Change

+73.0%

Current

32,700 pulls

huggingface

downloads

+31.1%

huggingface

downloads_all_time

+31.1%

huggingface

likes

+2.7%

huggingface

followers

+0.2%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 32,700 pulls daily 01.09.2026
huggingface downloads_all_time 207,941 daily 01.09.2026
huggingface followers 101,994 daily 01.09.2026
huggingface likes 4,633 daily 01.09.2026
huggingface downloads 207,941 daily 01.09.2026
ollama downloads 18,900 pulls daily 31.08.2026
huggingface downloads_all_time 158,598 daily 31.08.2026
huggingface followers 101,772 daily 31.08.2026
huggingface likes 4,513 daily 31.08.2026
huggingface downloads 158,598 daily 31.08.2026
ollama downloads 10,800 pulls daily 30.08.2026
huggingface downloads_all_time 121,976 daily 30.08.2026
huggingface followers 101,528 daily 30.08.2026
huggingface likes 4,378 daily 30.08.2026
huggingface downloads 121,976 daily 30.08.2026
huggingface followers 101,306 daily 29.08.2026
huggingface likes 4,285 daily 29.08.2026
huggingface downloads 52,341 daily 29.08.2026
huggingface followers 101,117 daily 28.08.2026
huggingface likes 4,152 daily 28.08.2026
huggingface downloads 4,810 daily 28.08.2026
huggingface followers 100,862 daily 27.08.2026
huggingface likes 3,944 daily 27.08.2026
huggingface downloads 4,810 daily 27.08.2026
huggingface followers 100,537 daily 26.08.2026
huggingface likes 3,587 daily 26.08.2026
huggingface downloads 2,551 daily 26.08.2026

View full metric history →

Related Models