Qwen3.5-397B-A17B
Qwen Team (Alibaba Cloud)
Parameters
—
Architecture
—
Released
—
License
—
About
No detailed description available for this model.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
SWE-bench Verified
coding_agent
|
86.93%
|
22.04.2026 |
|
SWE-bench Pro
coding_agent
|
63.62%
|
22.04.2026 |
|
SWE-bench Multilingual
coding_agent
|
76.09%
|
22.04.2026 |
|
Terminal-Bench 2.0
coding_agent
|
56.28%
|
22.04.2026 |
|
SkillsBench Avg5
coding_agent
|
41.09%
|
22.04.2026 |
|
QwenWebBench
coding_agent
|
37.28%
|
22.04.2026 |
|
NL2Repo
coding_agent
|
33.69%
|
22.04.2026 |
|
Claw-Eval Avg
coding_agent
|
85.32%
|
22.04.2026 |
|
Claw-Eval Pass^3
coding_agent
|
50.88%
|
22.04.2026 |
|
QwenClawBench
coding_agent
|
96.73%
|
22.04.2026 |
|
MMLU-Redux
knowledge
|
97.06%
|
22.04.2026 |
|
SuperGPQA
knowledge
|
99.55%
|
22.04.2026 |
|
GPQA Diamond
stem_reasoning
|
88.84%
|
22.04.2026 |
|
Humanity's Last Exam
stem_reasoning
|
50.38%
|
22.04.2026 |
|
LiveCodeBench v6
stem_reasoning
|
86.49%
|
22.04.2026 |
|
HMMT Feb 25
stem_reasoning
|
100.00%
|
22.04.2026 |
|
HMMT Nov 25
stem_reasoning
|
80.71%
|
22.04.2026 |
|
HMMT Feb 26
stem_reasoning
|
87.31%
|
22.04.2026 |
|
IMOAnswerBench
stem_reasoning
|
82.81%
|
22.04.2026 |
|
AIME 26
stem_reasoning
|
92.47%
|
22.04.2026 |
|
MMMU
vision_language
|
100.00%
|
22.04.2026 |
|
MMMU-Pro
vision_language
|
88.78%
|
22.04.2026 |
|
DynaMath
vision_language
|
84.27%
|
22.04.2026 |
|
RealWorldQA
vision_language
|
89.30%
|
22.04.2026 |
|
MMStar
vision_language
|
100.00%
|
22.04.2026 |
|
SimpleVQA
vision_language
|
100.00%
|
22.04.2026 |
|
CC-OCR
document_understanding
|
100.00%
|
22.04.2026 |
|
ERQA
spatial_intelligence
|
84.76%
|
22.04.2026 |
|
CountBench
spatial_intelligence
|
92.31%
|
22.04.2026 |
|
RefCOCO (avg)
spatial_intelligence
|
95.45%
|
22.04.2026 |
|
VideoMME (w sub.)
video_understanding
|
79.67%
|
22.04.2026 |
|
VideoMMMU
video_understanding
|
76.81%
|
22.04.2026 |
|
MLVU
video_understanding
|
95.86%
|
22.04.2026 |
|
MVBench
video_understanding
|
100.00%
|
22.04.2026 |
|
V-Star
vision_language
|
93.47%
|
22.04.2026 |
|
AgentWorldBench MCP
general_agent
|
88.20%
|
24.06.2026 |
|
AgentWorldBench Search
general_agent
|
55.86%
|
24.06.2026 |
|
AgentWorldBench Terminal
general_agent
|
77.90%
|
24.06.2026 |
|
AgentWorldBench SWE
general_agent
|
86.96%
|
24.06.2026 |
|
AgentWorldBench Web
general_agent
|
19.18%
|
24.06.2026 |
|
AgentWorldBench OS
general_agent
|
32.88%
|
24.06.2026 |
|
AgentWorldBench Overall
general_agent
|
68.47%
|
24.06.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
47.94%
|
25.06.2026 |
|
Terminal-Bench 2.1 (Claude Code)
coding_agent
|
59.88%
|
25.06.2026 |
|
SWE Atlas - QnA
coding_agent
|
20.44%
|
25.06.2026 |
|
SWE Atlas - RF
coding_agent
|
25.31%
|
25.06.2026 |
|
SWE Atlas - TW
coding_agent
|
21.93%
|
25.06.2026 |
|
BrowseComp
general_agent
|
42.45%
|
04.06.2026 |
|
Vals.ai Financial Agent 1.1 (without web search)
general_agent
|
100.00%
|
04.06.2026 |
|
HLE (with tools)
stem_reasoning
|
65.68%
|
04.06.2026 |
|
Vals.ai Financial Agent 1.1 (with web search)
general_agent
|
72.03%
|
04.06.2026 |
|
CritPt (no tools)
stem_reasoning
|
5.68%
|
04.06.2026 |
|
GDPVal
general_agent
|
34.60
|
04.06.2026 |
|
IOI 2025
stem_reasoning
|
441.30
|
04.06.2026 |
|
OmniScience Accuracy
knowledge
|
58.56%
|
04.06.2026 |
|
ProfBench (Search)
general_agent
|
50.36%
|
04.06.2026 |
|
IMOAnswerBench (with tools)
stem_reasoning
|
50.56%
|
04.06.2026 |
|
Apex-Shortlist (no tools)
stem_reasoning
|
53.02%
|
04.06.2026 |
|
OmniScience Non-Hallucination
knowledge
|
6.06%
|
04.06.2026 |
|
IFBench (prompt loose)
instruction_following
|
80.51%
|
04.06.2026 |
|
PinchBench
general_agent
|
65.69%
|
04.06.2026 |
|
TauBench V3 Airline
general_agent
|
11.43%
|
04.06.2026 |
|
Apex-Shortlist (with tools)
stem_reasoning
|
24.57%
|
04.06.2026 |
|
Multi-Challenge
instruction_following
|
99.07%
|
04.06.2026 |
|
TauBench V3 Retail
general_agent
|
90.32%
|
04.06.2026 |
|
AA-LCR
long_context
|
85.38%
|
04.06.2026 |
|
RULER (1M)
long_context
|
34.29%
|
04.06.2026 |
|
TauBench V3 Telecom
general_agent
|
96.55%
|
04.06.2026 |
|
TauBench V3 Banking
general_agent
|
39.32%
|
04.06.2026 |
|
SciCode (subtask)
stem_reasoning
|
47.09%
|
04.06.2026 |
|
Longbench v2 (≤ 1M)
long_context
|
100.00%
|
04.06.2026 |
|
MMLU-ProX
multilingual
|
100.00%
|
04.06.2026 |
|
TauBench V3 Average
general_agent
|
64.47%
|
04.06.2026 |
|
WMT24++ (en→xx)
multilingual
|
100.00%
|
04.06.2026 |
|
DeepSWE
coding_agent
|
1.38%
|
19.08.2026 |
|
Frontier-Bench v0.1
coding_agent
|
1.40
|
19.08.2026 |
|
MCP-Atlas
general_agent
|
81.67%
|
19.08.2026 |
|
Toolathlon Verified
general_agent
|
20.96%
|
19.08.2026 |
|
WideSearch
general_agent
|
75.93%
|
19.08.2026 |
|
WildClawBench
coding_agent
|
97.35%
|
19.08.2026 |
|
C-Eval
knowledge
|
99.08%
|
22.04.2026 |
|
CharXiv (RQ)
document_understanding
|
58.50%
|
22.04.2026 |
|
AgentWorldBench Android
general_agent
|
26.77%
|
24.06.2026 |
|
MMLU-Pro
knowledge
|
89.68%
|
22.04.2026 |