Qwen3.6-27B vs Qwen3.6-35B-A3B

New Comparison

Scores are fetched live from the current database on every visit — nothing is cached.

Benchmark Scores Comparison

Qwen3.6-35B-A3B Qwen3.6-27B

Qwen3.6-35B-A3B

35.0B params

Score: 100.0

Qwen3.6-27B

27.0B params

Score: 100.0

Coverage matrix

2 models × 103 benchmarks — hover a benchmark name for its category

Full Coverage Matrix
Shading = score relative to the best score among the selected models: Low Mid High Missing (—) The totals row shows how many of the 103 benchmarks each model has scores for.
Benchmark Qwen3.6-35B-A3B Qwen Team (Alibaba Cloud) Qwen3.6-27B Qwen Team (Alibaba Cloud)
HMMT Feb 25 90.7 93.8
DynaMath 82.8 85.6
GPQA Diamond 86 87.8
IFBench (prompt loose) — 69.1
MathVista (mini) 86.4 87.4
MMBench EN-DEV v1.1 92.8 92.3
LiveCodeBench v6 80.4 83.9
SWE-bench Pro 49.5 53.5
RealWorldQA 85.3 84.1
EmbSpatialBench 84.3 84.6
AA-LCR — 73.3
MathVision — 85.1
AIME 26 92.7 94.1
AndroidWorld — 70.3
Agents' Last Exam — 10.6
Toolathlon Verified 41.7 —
RecreationBench — 29.8
VITA-Bench 35.6 —
TAU3-Bench 67.2 —
ScreenSpot Pro — 76.1
WMDP (Bio) — 84.8
MMStar 80.7 81.4
DeepSearch QA — 71.1
SWE Atlas - TW 13.3 —
WildClawBench 68.7 43.2
SWE Atlas - QnA 15.5 —
AI2D_TEST 92.7 —
ParseBench Text Formatting 58.3 —
NL2Repo 29.4 36.2
HPCT — 48.7
MMMU-Pro 75.3 75.8
QwenClawBench 52.6 53.4
BabyVision — 28.9
Claw-Eval Avg 68.7 72.4
VideoMMMU 83.7 84.4
HMMT Nov 25 89.1 90.7
SWE-MM — 25.7
MMLU-Pro 85.2 86.2
SuperGPQA 64.7 66
OSWorld-Verified — 63.9
HLE (with tools) 28.9 —
LVBench 71.4 —
CharXiv (RQ) 78 78.4
BrowseComp 62 —
MMLU-Redux 93.3 93.5
SkillsBench — 46.6
WebArena-Verified — 48.8
ParseBench Text Content 90.7 —
CoWorkBench — 61
Lab Bench (ProtocolQA) — 69.1
MCPMark 37 —
SWE-bench Multilingual 67.2 71.3
MLVU 86.2 86.6
Beam128K — 63
IFBench — 70.8
SWE-bench Verified 73.4 77.2
IMOAnswerBench 78.9 80.8
Humanity's Last Exam 21.4 24
CountBench 96.1 97.8
WMDP (Chem) — 74.8
ERQA 61.8 62.5
VideoMME (w/o sub.) 82.5 —
Terminal-Bench 2.0 51.5 59.3
DeepPlanning 25.9 —
QwenSWEBench — 49.3
QwenWebBench 1397 1487
ODInW13 50.8 —
VlmsAreBlind 96.6 97
Terminal-Bench 2.1 (Terminus-2) 52.5 63.4
DeepSWE 0 —
OmniDocBench 1.5 89.9 89.4
HallusionBench 69.8 —
OCRBench 90 89.4
HMMT Feb 26 83.6 84.3
CC-OCR 81.9 81.2
RefCOCO (avg) 92 92.5
JobBench — 21.8
Artificial Analysis Intelligence Index 32 —
SciCode (subtask) — 39.8
Tool Decathlon 26.9 —
Frontier-Bench v0.1 1.4 —
ParseBench Mean 44.1 —
ZEROBench_sub 34.4 —
C-Eval 90 91.4
WideSearch 60.1 —
SimpleVQA 58.9 56.1
VideoMME (w sub.) 86.6 87.7
SkillsBench Avg5 28.7 48.2
Vision2Web — 45
TauBench V3 Banking — 16.7
Claw-Eval Pass^3 50 60.6
MCP-Atlas 62.8 62.5
SWE Atlas - RF 11.4 —
V-Star 90.1 94.7
DeepSWE 1.1 — 13.3
Terminal-Bench 2.1 (Claude Code) 49.2 —
MBCT — 45.9
VCT — 33.7
Gaia2 — 40
RefSpatialBench 64.3 70
MMMU 81.7 82.9
GDPVal-AA v2 — 1141
MVBench 74.6 75.5
Covered 73/103 78/103

Fully comparable benchmarks

Scored by all 2 selected models (48)

Benchmark Category Qwen3.6-35B-A3B Qwen3.6-27B
HMMT Feb 25 stem_reasoning 90.7 93.8
DynaMath vision_language 82.8 85.6
GPQA Diamond stem_reasoning 86 87.8
MathVista (mini) vision_language 86.4 87.4
MMBench EN-DEV v1.1 vision_language 92.8 92.3
LiveCodeBench v6 stem_reasoning 80.4 83.9
SWE-bench Pro coding_agent 49.5 53.5
RealWorldQA vision_language 85.3 84.1
EmbSpatialBench spatial_intelligence 84.3 84.6
AIME 26 stem_reasoning 92.7 94.1
MMStar vision_language 80.7 81.4
WildClawBench coding_agent 68.7 43.2
NL2Repo coding_agent 29.4 36.2
MMMU-Pro vision_language 75.3 75.8
QwenClawBench coding_agent 52.6 53.4
Claw-Eval Avg coding_agent 68.7 72.4
VideoMMMU video_understanding 83.7 84.4
HMMT Nov 25 stem_reasoning 89.1 90.7
MMLU-Pro knowledge 85.2 86.2
SuperGPQA knowledge 64.7 66
CharXiv (RQ) document_understanding 78 78.4
MMLU-Redux knowledge 93.3 93.5
SWE-bench Multilingual coding_agent 67.2 71.3
MLVU video_understanding 86.2 86.6
SWE-bench Verified coding_agent 73.4 77.2
IMOAnswerBench stem_reasoning 78.9 80.8
Humanity's Last Exam stem_reasoning 21.4 24
CountBench spatial_intelligence 96.1 97.8
ERQA spatial_intelligence 61.8 62.5
Terminal-Bench 2.0 coding_agent 51.5 59.3
QwenWebBench coding_agent 1397 1487
VlmsAreBlind vision_language 96.6 97
Terminal-Bench 2.1 (Terminus-2) coding_agent 52.5 63.4
OmniDocBench 1.5 document_understanding 89.9 89.4
OCRBench document_understanding 90 89.4
HMMT Feb 26 stem_reasoning 83.6 84.3
CC-OCR document_understanding 81.9 81.2
RefCOCO (avg) spatial_intelligence 92 92.5
C-Eval knowledge 90 91.4
SimpleVQA vision_language 58.9 56.1
VideoMME (w sub.) video_understanding 86.6 87.7
SkillsBench Avg5 coding_agent 28.7 48.2
Claw-Eval Pass^3 coding_agent 50 60.6
MCP-Atlas general_agent 62.8 62.5
V-Star vision_language 90.1 94.7
RefSpatialBench spatial_intelligence 64.3 70
MMMU vision_language 81.7 82.9
MVBench video_understanding 74.6 75.5

Partial coverage benchmarks

Scored by at least one but not all selected models (55) — missing scores show as —

Benchmark Category Qwen3.6-35B-A3B Qwen3.6-27B
IFBench (prompt loose) 1/2 models instruction_following — 69.1
AA-LCR 1/2 models long_context — 73.3
MathVision 1/2 models vision_language — 85.1
AndroidWorld 1/2 models general_agent — 70.3
Agents' Last Exam 1/2 models general_agent — 10.6
Toolathlon Verified 1/2 models general_agent 41.7 —
RecreationBench 1/2 models general_agent — 29.8
VITA-Bench 1/2 models general_agent 35.6 —
TAU3-Bench 1/2 models general_agent 67.2 —
ScreenSpot Pro 1/2 models general_agent — 76.1
WMDP (Bio) 1/2 models safety — 84.8
DeepSearch QA 1/2 models general_agent — 71.1
SWE Atlas - TW 1/2 models coding_agent 13.3 —
SWE Atlas - QnA 1/2 models coding_agent 15.5 —
AI2D_TEST 1/2 models document_understanding 92.7 —
ParseBench Text Formatting 1/2 models document_understanding 58.3 —
HPCT 1/2 models safety — 48.7
BabyVision 1/2 models vision_language — 28.9
SWE-MM 1/2 models coding_agent — 25.7
OSWorld-Verified 1/2 models general_agent — 63.9
HLE (with tools) 1/2 models stem_reasoning 28.9 —
LVBench 1/2 models video_understanding 71.4 —
BrowseComp 1/2 models general_agent 62 —
SkillsBench 1/2 models general_agent — 46.6
WebArena-Verified 1/2 models general_agent — 48.8
ParseBench Text Content 1/2 models document_understanding 90.7 —
CoWorkBench 1/2 models general_agent — 61
Lab Bench (ProtocolQA) 1/2 models safety — 69.1
MCPMark 1/2 models general_agent 37 —
Beam128K 1/2 models long_context — 63
IFBench 1/2 models instruction_following — 70.8
WMDP (Chem) 1/2 models safety — 74.8
VideoMME (w/o sub.) 1/2 models video_understanding 82.5 —
DeepPlanning 1/2 models general_agent 25.9 —
QwenSWEBench 1/2 models coding_agent — 49.3
ODInW13 1/2 models spatial_intelligence 50.8 —
DeepSWE 1/2 models coding_agent 0 —
HallusionBench 1/2 models vision_language 69.8 —
JobBench 1/2 models general_agent — 21.8
Artificial Analysis Intelligence Index 1/2 models composite 32 —
SciCode (subtask) 1/2 models stem_reasoning — 39.8
Tool Decathlon 1/2 models general_agent 26.9 —
Frontier-Bench v0.1 1/2 models coding_agent 1.4 —
ParseBench Mean 1/2 models document_understanding 44.1 —
ZEROBench_sub 1/2 models vision_language 34.4 —
WideSearch 1/2 models general_agent 60.1 —
Vision2Web 1/2 models vision_language — 45
TauBench V3 Banking 1/2 models general_agent — 16.7
SWE Atlas - RF 1/2 models coding_agent 11.4 —
DeepSWE 1.1 1/2 models coding_agent — 13.3
Terminal-Bench 2.1 (Claude Code) 1/2 models coding_agent 49.2 —
MBCT 1/2 models safety — 45.9
VCT 1/2 models safety — 33.7
Gaia2 1/2 models general_agent — 40
GDPVal-AA v2 1/2 models general_agent — 1141