Qwen3.8-2.4T-A95B

Qwen

Parameters

2.4T total / 95.0B active

MoE: total / active

Architecture

92-layer hybrid MoE: 23 x (3 x (Gated DeltaNet -> MoE) + 1 x (Gated Attention -> MoE)); 512 experts (10 routed + 1 shared, dim 2048); hidden 8192; GDN 128V/16QK dim 128; GA 64Q/4KV dim 256; MTP multi-steps; ctx 262144 -> 1010000.

Released

08.08.2026

License

Qwen3.8-Max License

Open Weights Commercial Use Multimodal BF16 Qwen3.8 en zh

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

1,010,000 tokens

Openness Index Score 70.0/100

About

Qwen3.8-2.4T-A95B (Qwen/Qwen3.8-2.4T-A95B) is the flagship sparse Mixture-of-Experts model of Alibaba's Qwen3.8 generation - 2.4T total parameters with 95B activated per token - released August 8, 2026 under the Qwen3.8-Max License. Its 92-layer hybrid stack repeats 23x (3x (Gated DeltaNet -> MoE) + 1x (Gated Attention -> MoE)): 69 linear-attention layers (Gated DeltaNet, 128 V / 16 QK heads, head dim 128) and 23 full-attention layers (Gated Attention, 64 Q / 4 KV heads, head dim 256, RoPE dim 64), with 512 experts (10 routed + 1 shared per token, expert intermediate dim 2048), hidden size 8192, 248K vocabulary, Multi-Token Prediction (MTP) trained with multiple steps, and a 262,144-token native context extensible to 1,010,000 tokens.

Qwen3.8 delivers comprehensive improvements across coding, professional work, research and long-horizon agentic tasks; stronger autonomous planning and handling of environment feedback for reliable end-to-end agent execution; broader compatibility with popular harnesses and development tools; and flexible thinking control - reasoning depth tunable via reasoning_effort with reasoning context from historical messages retained via preserve_thinking.

Training Data Pre-training & Post-training

Benchmark Scores

Benchmark Score Date
Terminal Bench 2.1
coding_agent
95.57%
18.08.2026
SWE-bench Pro
coding_agent
84.62%
18.08.2026
DeepSWE 1.1
coding_agent
73.41%
18.08.2026
NL2Repo-Bench
coding_agent
71.76%
18.08.2026
FrontierSWE
coding_agent
75.17%
18.08.2026
MLS-Bench-Lite
coding_agent
61.64%
18.08.2026
PaperBench
coding_agent
100.00%
18.08.2026
AndroidBench
coding_agent
66.43%
18.08.2026
QwenSWEBench
coding_agent
84.86%
18.08.2026
QwenQoderBench
coding_agent
82.13%
18.08.2026
QwenReactBench
coding_agent
80.17%
18.08.2026
QwenSVGBench
coding_agent
82.63%
18.08.2026
CoWorkBench
general_agent
96.43%
18.08.2026
WorkSpaceBench
general_agent
56.55%
18.08.2026
JobBench
general_agent
68.40%
18.08.2026
SkillsBench
general_agent
91.97%
18.08.2026
Agents' Last Exam
general_agent
77.36%
18.08.2026
Automation-Bench
general_agent
37.50%
18.08.2026
Toolathlon Verified
general_agent
89.22%
18.08.2026
WideSearch
general_agent
91.39%
18.08.2026
HLE (with tools)
stem_reasoning
82.42%
18.08.2026
GPQA Diamond
stem_reasoning
96.79%
18.08.2026
Humanity's Last Exam
stem_reasoning
78.60%
18.08.2026
IFBench
instruction_following
100.00%
18.08.2026
$OneMillion-Bench
general_capabilities
40.68%
18.08.2026
HealthBench
general_capabilities
100.00%
18.08.2026
PLawBench
general_capabilities
100.00%
18.08.2026
PRBench-Legal
general_capabilities
100.00%
18.08.2026
PRBench-Finance
general_capabilities
100.00%
18.08.2026
MRCR v2 256K (8-needle)
long_context
91.51%
18.08.2026
LongBench v2
long_context
93.67%
18.08.2026
NL2Repo
coding_agent
70.15%
28.08.2026
ProgramBench
coding_agent
10.89%
28.08.2026
Cybergym
general_agent
80.57%
28.08.2026
ExploitGym (2h)
cybersecurity
14.00
28.08.2026
ExploitGym (6h)
cybersecurity
26.00
28.08.2026
ExploitBench
cybersecurity
8.21%
28.08.2026
GDPVal-AA v2
general_agent
94.98%
28.08.2026
SWE-bench Multilingual
coding_agent
91.83%
28.08.2026
DeepSWE
coding_agent
77.85%
28.08.2026
SWE Atlas - QnA
coding_agent
84.31%
28.08.2026
SWE Atlas - TW
coding_agent
75.27%
28.08.2026
SWE Atlas - RF
coding_agent
83.84%
28.08.2026
SWE-Marathon
coding_agent
61.22%
28.08.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
96.76%
28.08.2026
Harbor-Index
coding_agent
56.17%
28.08.2026
Hy-Backend 2.0 (Internal)
coding_agent
37.18%
28.08.2026
Hy-SWE Max Verified (Internal)
coding_agent
76.78%
28.08.2026
Hy-CompanyBench V2 (Internal)
general_agent
78.09%
28.08.2026
DRACO
general_agent
47.86%
28.08.2026
Hy-LifeSearch (Internal)
general_agent
35.10%
28.08.2026
Hy-BrowseComp-Pro2 (Internal)
general_agent
1.35%
28.08.2026
OfficeQA Pro
general_agent
96.93%
28.08.2026
Apex-Agents
general_agent
78.45%
28.08.2026
SkillsBench Avg5
coding_agent
100.00%
28.08.2026
BankerToolBench
general_agent
60.00%
28.08.2026
E-Bench (Internal)
general_agent
57.01%
28.08.2026
Hy-FinAgentBench (Internal)
domain_finance
57.04%
28.08.2026
Hy-FinmodelBench v2 (Internal)
domain_finance
63.90%
28.08.2026
CritPt (no tools)
stem_reasoning
61.20%
28.08.2026
SUPERChem
stem_reasoning
38.59%
28.08.2026
ArXivMath
stem_reasoning
55.40%
28.08.2026
HorizonMath (pass@4)
stem_reasoning
25.42%
28.08.2026
MathArena Apex 2025
stem_reasoning
64.71%
28.08.2026
BrokenArXiv
stem_reasoning
31.37%
28.08.2026
BioMysteryBench
stem_reasoning
21.98%
28.08.2026
E-Bench-Code (Internal)
coding_agent
14.74%
28.08.2026
MCP-Atlas
general_agent
94.80%
28.08.2026

Model Ecosystem and Derivatives

Model Ecosystem

Model Tree

Finetunes

Quantizations

  • 29 quantized models available
  • Supported tools: llama.cpp, LM Studio, Jan, Ollama
  • View quantizations

Spaces

  • 6 Spaces using Qwen/Qwen3.8-2.4T-A95B

Collection

Part of the Qwen3.8 collection:

Community

  • 38 discussions on HuggingFace
  • 99.9k followers
  • 18,893 downloads last month

Citation Information

Citation

If you find our work helpful, feel free to give us a cite.

@misc{qwen38,
    title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
    url = {https://qwen.ai/blog?id=qwen3.8},
    author = {{Qwen Team}},
    month = {August},
    year = {2026}
}

Source

Thinking Mode and Reasoning Control

Thinking Control

Overview

Qwen3.8-2.4T-A95B is a text-only model that requires thinking mode for all interactions. Thinking cannot be disabled. Every response automatically begins with reasoning enclosed in <think>\n...\n</think> before the final output.

reasoning_effort Levels

Reasoning depth can be tuned to balance accuracy and cost:

Level Description
xhigh (default) For complex tasks demanding thorough analysis
medium Balancing accuracy and speed
low Efficient reasoning optimizing for speed and cost

preserve_thinking

  • Enabled by default for all workloads.
  • Retains reasoning context from historical messages, providing the best out-of-the-box experience.
  • Can be passed via extra_body with chat_template_kwargs in self-hosted inference, or directly in Qwen Cloud API.

API Parameters

extra_body={
    "chat_template_kwargs": {
        "enable_thinking": True,   # on by default; should not be turned off
        "preserve_thinking": True  # on by default
    }
}
reasoning_effort="xhigh"  # xhigh by default; supported: xhigh, medium, low

Recommended Settings and Best Practices

Best Practices

Sampling Parameters

Recommended sampling parameters for optimal generation:

Parameter Value
temperature 1.0
top_p 0.95
top_k 20
min_p 0.0
presence_penalty 0.0
repetition_penalty 1.0

Notes on presence_penalty

  • Can be adjusted between 0 and 2 to reduce endless repetition.
  • Higher values may occasionally result in language mixing and slight performance decrease.
  • Support for sampling parameters varies by inference framework.

Output Length Configuration

For agentic tasks, allocate sufficient output length within the 1M context:

Output Type Max Tokens
Reasoning Content 262,144
Final Response 131,072

These settings provide necessary capacity for complex reasoning while ensuring ample space for high-quality final deliverables.

Chat Completions API Usage

API Usage

Key Characteristics

  • Text-only model: Multimodal inputs are not supported.
  • Thinking mode required: Thinking cannot be disabled. Every response automatically begins with reasoning enclosed in <think>\n...\n</think> before the final output.
  • preserve_thinking: Enabled by default for all workloads.

Chat Completions API (OpenAI-compatible)

from openai import OpenAI
## Configured by environment variables
client = OpenAI()

messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]

completion = client.chat.completions.create(
    model="Qwen/Qwen3.8-2.4T-A95B",
    messages=messages,
    extra_body={
        "chat_template_kwargs": {
            "enable_thinking": True,  # on by default; should not be turned off
            "preserve_thinking": True, # on by default
        },
    },
    reasoning_effort="xhigh",  # xhigh by default; supported: xhigh, medium, low
    stream=True,
    stream_options={"include_usage": True},
)

reasoning_content = ""
answer_content = ""
is_answering = False

for chunk in completion:
    if not chunk.choices:
        print(chunk.usage)
        continue
    delta = chunk.choices[0].delta
    if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None:
        reasoning_content += delta.reasoning_content
    if hasattr(delta, "content") and delta.content:
        is_answering = True
        answer_content += delta.content

Qwen Cloud API

When using Qwen Cloud APIs, pass extra_body directly (not nested under chat_template_kwargs):

extra_body={"enable_thinking": True, "preserve_thinking": True}

Inference Frameworks and Deployment

Deployment Options

Supported Inference Frameworks

Qwen3.8-2.4T-A95B is compatible with multiple popular inference frameworks:

SGLang

vLLM

TokenSpeed

Managed Cloud Service

For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud.

Recommendations

  • For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, vLLM, or TokenSpeed are recommended.
  • Inference efficiency and throughput vary significantly across frameworks — use the latest framework versions for optimal performance and compatibility.

Benchmark Performance Results

Benchmark Results

Coding Agent Benchmarks

Benchmark Qwen3.8-Max Qwen3.7-Max Opus 4.8 Fable 5 GPT 5.6 Sol
Terminal Bench 2.1 86.6 74.5 84.6 84.6 88.8
SWE-bench Pro 67.7 60.6 69.2 80.0 64.6
DeepSWE 1.1 56.6 21.6 59.0 70.0 73.0
NL2Repo-Bench 55.9 47.2 69.4 -- --
FrontierSWE 73.5 40.7 70.0 88.8 --
MLS-Bench-Lite 41.0 31.7 42.8 49.9 46.2
PaperBench 93.0 64.8 80.3 88.8 90.5
AndroidBench 75.1 56.5 69.8 84.5 74.0
QwenSWEBench 80.7 63.4 84.0 86.3 73.5
QwenQoderBench 58.4 36.8 62.7 63.1 53.8
QwenReactBench (Elo) 1724 1538 1694 1770 1564
QwenSVGBench (Elo) 2213 1499 1648 1690 1758

General Agent Benchmarks

Benchmark Qwen3.8-Max Qwen3.7-Max Opus 4.8 Fable 5 GPT 5.6 Sol
CoWorkBench 74.8 64.6 72.3 75.9 71.5
WorkSpaceBench 67.7 61.4 66.8 68.7 65.6
JobBench 53.4 31.3 48.4 57.4 45.4
SkillsBench 70.2 61.2 65.1 70.9 73.5
Agents' Last Exam (Pass/Score) 27.0/52.4 11.8/31.1 27.0/45.1 --/-- 30.6/53.6
Automation-Bench (Pass@1) 27.3 14.2 27.2 29.1 29.7
Toolathlon Verified (Pass@1) 72.5 49.7 76.2 77.9 74.9
WideSearch 81.9 75.2 72.9 81.2 --
HLE w/ tools 56.2 53.5 57.9 64.5 58.0

General Capabilities Benchmarks

Benchmark Qwen3.8-Max Qwen3.7-Max Opus 4.8 Fable 5 GPT 5.6 Sol
GPQA Diamond 92.6 92.4 92.0 92.6 94.1
HLE 43.6 41.4 45.7 53.3 47.2
IFBench 82.8 79.1 62.2 63.5 72.7
$OneMillion-Bench (expert) 52.5 44.4 41.8 55.9 53.8
HealthBench 60.2 54.5 52.4 -- 55.3
PLawBench 73.2 58.9 69.6 70.2 72.3
PRBench-Legal 57.6 48.5 52.7 57.6 57.6
PRBench-Finance 58.3 46.8 51.9 55.8 55.5
MRCR v2 256K (8-needle) 92.9 86.7 83.2 -- 93.8
LongBench v2 66.3 65.3 69.1 -- 67.1

HuggingFace Evaluation Results

  • DeepSWE: 56.6
  • GPQA Diamond: 92.6
  • SWE-bench Pro: 67.7
  • WildClawBench Overall: 56.2
  • Terminal Bench 2.1: 86.6
  • HLE: 43.6

Context Length Specifications

Context Length

Native Context Length

  • 262,144 tokens natively supported.

Extended Context Length

  • Extensible up to 1,010,000 tokens (approximately 1M).

Notes

  • Qwen3.8-Max (the official hosted version) offers 1M context length by default.
  • For long-context tasks, the model supports RoPE-based position scaling.
  • The extended context is particularly useful for agentic workflows requiring extensive reasoning and output space.

Model Architecture and Structure

Model Architecture

Overview

  • Type: Causal Language Model
  • Training Stage: Pre-training & Post-training
  • Total Parameters: 2.4T (2,400,000,000,000)
  • Activated Parameters: 95B (95,000,000,000)
  • Hidden Dimension: 8192
  • Token Embedding: 248,320 (Padded)
  • Number of Layers: 92
  • Tensor Type: BF16

Hidden Layout

The model uses a repeating pattern of 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)), alternating between linear attention and full attention layers.

Gated DeltaNet (Linear Attention)

  • Number of Linear Attention Heads: 128 for V and 16 for QK
  • Head Dimension: 128

Gated Attention (Full Attention)

  • Number of Attention Heads: 64 for Q and 4 for KV
  • Head Dimension: 256
  • Rotary Position Embedding Dimension: 64

Mixture of Experts (MoE)

  • Number of Experts: 512
  • Number of Activated Experts: 10 Routed + 1 Shared
  • Expert Intermediate Dimension: 2048

Multi-Token Prediction (MTP)

  • Trained with multiple steps

LM Output

  • 248,320 (Padded)

Qwen3.8 Key Highlights

Qwen3.8 Highlights

Qwen3.8 is the most capable generation in the Qwen open-model family to date, bringing a Qwen-Max-class model to open release for the first time. Built on the architectural foundation of Qwen3.5, it delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.

Core Enhancements

  • Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
  • Downstream Compatibility: Broader support for popular harnesses and development tools, making it easier to integrate into existing stacks.
  • Flexible Thinking Control: Reasoning depth can be tuned with reasoning_effort, and reasoning context from historical messages is retained via preserve_thinking.

Qwen3.8-Max

Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with additional features:

  • Vision input support
  • Non-thinking support
  • 1M context length by default
  • Official built-in tools

Blog post: Qwen3.8-Max

Architecture

Decoder Block ×92 input Embedding vocab 248K · d 8192 Linear / Recurrent Hybrid 64:4 · dₕ 256 ×69 Full Attention Hybrid 64:4 · dₕ 256 ×23 MoE FFN 512 experts · top-10 · dᴻ 2048 Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (64:4)
MoE
512 experts · top-10 per token
Layers
92
Hidden size
8192
Context
262K tokens
Parameters
2400000M
Active params
95000M

Source: Hugging Face config.json · Qwen3_5MoeForCausalLM · exact layer pattern · model repo

Type: Causal Language Model
Attention: Hybrid (Gated DeltaNet + Gated Attention)
Decoder: Autoregressive
MoE: yes (512 experts)
Routing: Top-k router with 10 routed experts + 1 shared expert
Layers 92
Context length 262K
Extended context 1M
Experts 512
Experts per token 10 Routed + 1 Shared
Hidden size 8192
Expert FFN dim 2048
Expert Intermediate Dim 2048
Gated Attention Head Dim 256
Gated Attention Kv Heads 4
Gated Attention Q Heads 64
Gated Deltanet Head Dim 128
Gated Deltanet Heads Qk 16
Gated Deltanet Heads V 128
Hidden Layout 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))
MTP trained with multiple steps
RoPE dim 64
Lm Output

248K

Token Embedding

248K

Training Pipeline

  1. 1
    pretraining

    Pre-training

    Pre-training stage: Causal Language Model trained on large-scale multilingual corpus with Multi-Token Prediction (MTP). 2.4T total parameters, 95B activated parameters, MoE architecture with 512 experts (10 routed + 1 shared).

  2. 2
    sft

    Post-training (SFT + RL)

    Post-training stage: The model is a post-trained version with reasoning capabilities. Features reasoning_effort control (xhigh/medium/low) and preserve_thinking for retaining reasoning context across messages. Thinking mode is always on and cannot be disabled.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Qwen3.8 pre-training corpus pretraining — —

Trend Analysis

24h Change

+7.5%

7d Change

+88.1%

Current

38,782

huggingface

followers

+0.2%

huggingface

downloads_all_time

+7.5%

huggingface

likes

+0.0%

ollama

downloads

+0.0%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 1,300,000 pulls daily 01.09.2026
huggingface downloads_all_time 38,782 daily 01.09.2026
huggingface followers 101,994 daily 01.09.2026
huggingface likes 1,189 daily 01.09.2026
huggingface downloads 38,782 daily 01.09.2026
ollama downloads 1,300,000 pulls daily 31.08.2026
huggingface downloads_all_time 36,081 daily 31.08.2026
huggingface followers 101,772 daily 31.08.2026
huggingface likes 1,189 daily 31.08.2026
huggingface downloads 36,081 daily 31.08.2026
ollama downloads 1,200,000 pulls daily 30.08.2026
huggingface downloads_all_time 35,309 daily 30.08.2026
huggingface followers 101,528 daily 30.08.2026
huggingface likes 1,183 daily 30.08.2026
huggingface downloads 35,309 daily 30.08.2026
ollama downloads 1,100,000 pulls daily 29.08.2026
huggingface followers 101,306 daily 29.08.2026
huggingface likes 1,182 daily 29.08.2026
huggingface downloads 27,374 daily 29.08.2026
ollama downloads 1,000,000 pulls daily 28.08.2026
huggingface followers 101,117 daily 28.08.2026
huggingface likes 1,179 daily 28.08.2026
huggingface downloads 21,924 daily 28.08.2026
ollama downloads 967,200 pulls daily 27.08.2026
huggingface followers 100,862 daily 27.08.2026
huggingface likes 1,176 daily 27.08.2026
huggingface downloads 21,924 daily 27.08.2026
ollama downloads 893,300 pulls daily 26.08.2026
huggingface followers 100,537 daily 26.08.2026
huggingface likes 1,169 daily 26.08.2026
huggingface downloads 21,340 daily 26.08.2026
ollama downloads 810,300 pulls daily 25.08.2026
huggingface followers 100,133 daily 25.08.2026
huggingface likes 1,160 daily 25.08.2026
huggingface downloads 20,616 daily 25.08.2026
ollama downloads 726,100 pulls daily 24.08.2026
huggingface followers 99,896 daily 24.08.2026
huggingface likes 1,155 daily 24.08.2026
huggingface downloads 18,893 daily 24.08.2026
huggingface followers 99,669 daily 23.08.2026
huggingface likes 1,152 daily 23.08.2026
huggingface downloads 18,115 daily 23.08.2026
huggingface followers 99,461 daily 22.08.2026
huggingface likes 1,146 daily 22.08.2026
huggingface downloads 17,386 daily 22.08.2026
huggingface followers 99,274 daily 21.08.2026
huggingface likes 1,139 daily 21.08.2026
huggingface downloads 15,702 daily 21.08.2026
huggingface followers 99,037 daily 20.08.2026
huggingface likes 1,121 daily 20.08.2026

View full metric history →

Related Models