Qwen3.6-35B-A3B vs Qwen3.8-27B
New ComparisonScores are fetched live from the current database on every visit — nothing is cached.
Benchmark Scores Comparison
Qwen3.6-35B-A3B
Qwen3.8-27B
Qwen3.6-35B-A3B
35.0B params
Score: 100.0
Qwen3.8-27B
27.0B params
Score: 100.0
Coverage matrix
2 models × 92 benchmarks — hover a benchmark name for its category
Shading = score relative to the best score among the selected models:
Low
Mid
High
Missing (—)
The totals row shows how many of the 92 benchmarks each model has scores for.
| Benchmark | Qwen3.6-35B-A3B Qwen Team (Alibaba Cloud) | Qwen3.8-27B Qwen Team (Alibaba Cloud) |
|---|---|---|
| HMMT Feb 25 | 90.7 | — |
| DynaMath | 82.8 | — |
| GPQA Diamond | 86 | 89.2 |
| IFBench (prompt loose) | — | 79.5 |
| MathVista (mini) | 86.4 | — |
| MMBench EN-DEV v1.1 | 92.8 | — |
| LiveCodeBench v6 | 80.4 | 90.3 |
| SWE-bench Pro | 49.5 | 61.7 |
| RealWorldQA | 85.3 | 85.9 |
| Agents' Last Exam (Score) | — | 42.9 |
| EmbSpatialBench | 84.3 | — |
| MathVision | — | 90 |
| AIME 26 | 92.7 | — |
| AndroidWorld | — | 81.9 |
| Agents' Last Exam | — | 20.4 |
| Toolathlon Verified | 41.7 | 67.1 |
| RecreationBench | — | 47.1 |
| VITA-Bench | 35.6 | — |
| TAU3-Bench | 67.2 | — |
| MMStar | 80.7 | — |
| SWE Atlas - TW | 13.3 | — |
| WildClawBench | 68.7 | — |
| SWE Atlas - QnA | 15.5 | — |
| AI2D_TEST | 92.7 | — |
| ParseBench Text Formatting | 58.3 | — |
| NL2Repo | 29.4 | 42.3 |
| MMMU-Pro | 75.3 | — |
| QwenClawBench | 52.6 | — |
| BabyVision | — | 65.7 |
| Claw-Eval Avg | 68.7 | 56.9 |
| VideoMMMU | 83.7 | — |
| HMMT Nov 25 | 89.1 | — |
| SWE-MM | — | 38.6 |
| MMLU-Pro | 85.2 | — |
| SuperGPQA | 64.7 | — |
| OSWorld-Verified | — | 84.3 |
| HLE (with tools) | 28.9 | 30.8 |
| LVBench | 71.4 | 72.4 |
| CharXiv (RQ) | 78 | 83.7 |
| BrowseComp | 62 | — |
| MMLU-Redux | 93.3 | — |
| NL2Repo-Bench | — | 42.3 |
| WebArena-Verified | — | 64.8 |
| ParseBench Text Content | 90.7 | — |
| CoWorkBench | — | 70.7 |
| MCPMark | 37 | — |
| SWE-bench Multilingual | 67.2 | 73.8 |
| MLVU | 86.2 | — |
| IFBench | — | 79.5 |
| SWE-bench Verified | 73.4 | — |
| IMOAnswerBench | 78.9 | — |
| Humanity's Last Exam | 21.4 | 30.8 |
| CountBench | 96.1 | — |
| ERQA | 61.8 | 65.5 |
| VideoMME (w/o sub.) | 82.5 | — |
| Terminal-Bench 2.0 | 51.5 | — |
| DeepPlanning | 25.9 | — |
| QwenSWEBench | — | 79 |
| QwenWebBench | 1397 | — |
| ODInW13 | 50.8 | — |
| VlmsAreBlind | 96.6 | — |
| Terminal-Bench 2.1 (Terminus-2) | 52.5 | 73 |
| DeepSWE | 0 | — |
| OmniDocBench 1.5 | 89.9 | 91.1 |
| HallusionBench | 69.8 | — |
| OCRBench | 90 | — |
| HMMT Feb 26 | 83.6 | — |
| CC-OCR | 81.9 | — |
| RefCOCO (avg) | 92 | — |
| JobBench | — | 33.4 |
| Artificial Analysis Intelligence Index | 32 | — |
| Tool Decathlon | 26.9 | — |
| Frontier-Bench v0.1 | 1.4 | — |
| OSWorld 2.0 (Partial) | — | 48 |
| ParseBench Mean | 44.1 | — |
| OSWorld 2.0 (Binary) | — | 19.4 |
| ZEROBench_sub | 34.4 | — |
| C-Eval | 90 | — |
| WideSearch | 60.1 | — |
| SimpleVQA | 58.9 | — |
| VideoMME (w sub.) | 86.6 | — |
| SkillsBench Avg5 | 28.7 | — |
| Vision2Web | — | 62.9 |
| Claw-Eval Pass^3 | 50 | 57.4 |
| MCP-Atlas | 62.8 | — |
| SWE Atlas - RF | 11.4 | — |
| V-Star | 90.1 | — |
| DeepSWE 1.1 | — | 42.2 |
| Terminal-Bench 2.1 (Claude Code) | 49.2 | — |
| RefSpatialBench | 64.3 | — |
| MMMU | 81.7 | — |
| MVBench | 74.6 | — |
| Covered | 73/92 | 35/92 |
Fully comparable benchmarks
Scored by all 2 selected models (16)
| Benchmark | Category | Qwen3.6-35B-A3B | Qwen3.8-27B |
|---|---|---|---|
| GPQA Diamond | stem_reasoning | 86 | 89.2 |
| LiveCodeBench v6 | stem_reasoning | 80.4 | 90.3 |
| SWE-bench Pro | coding_agent | 49.5 | 61.7 |
| RealWorldQA | vision_language | 85.3 | 85.9 |
| Toolathlon Verified | general_agent | 41.7 | 67.1 |
| NL2Repo | coding_agent | 29.4 | 42.3 |
| Claw-Eval Avg | coding_agent | 68.7 | 56.9 |
| HLE (with tools) | stem_reasoning | 28.9 | 30.8 |
| LVBench | video_understanding | 71.4 | 72.4 |
| CharXiv (RQ) | document_understanding | 78 | 83.7 |
| SWE-bench Multilingual | coding_agent | 67.2 | 73.8 |
| Humanity's Last Exam | stem_reasoning | 21.4 | 30.8 |
| ERQA | spatial_intelligence | 61.8 | 65.5 |
| Terminal-Bench 2.1 (Terminus-2) | coding_agent | 52.5 | 73 |
| OmniDocBench 1.5 | document_understanding | 89.9 | 91.1 |
| Claw-Eval Pass^3 | coding_agent | 50 | 57.4 |
Partial coverage benchmarks
Scored by at least one but not all selected models (76) — missing scores show as —
| Benchmark | Category | Qwen3.6-35B-A3B | Qwen3.8-27B |
|---|---|---|---|
| HMMT Feb 25 1/2 models | stem_reasoning | 90.7 | — |
| DynaMath 1/2 models | vision_language | 82.8 | — |
| IFBench (prompt loose) 1/2 models | instruction_following | — | 79.5 |
| MathVista (mini) 1/2 models | vision_language | 86.4 | — |
| MMBench EN-DEV v1.1 1/2 models | vision_language | 92.8 | — |
| Agents' Last Exam (Score) 1/2 models | general_agent | — | 42.9 |
| EmbSpatialBench 1/2 models | spatial_intelligence | 84.3 | — |
| MathVision 1/2 models | vision_language | — | 90 |
| AIME 26 1/2 models | stem_reasoning | 92.7 | — |
| AndroidWorld 1/2 models | general_agent | — | 81.9 |
| Agents' Last Exam 1/2 models | general_agent | — | 20.4 |
| RecreationBench 1/2 models | general_agent | — | 47.1 |
| VITA-Bench 1/2 models | general_agent | 35.6 | — |
| TAU3-Bench 1/2 models | general_agent | 67.2 | — |
| MMStar 1/2 models | vision_language | 80.7 | — |
| SWE Atlas - TW 1/2 models | coding_agent | 13.3 | — |
| WildClawBench 1/2 models | coding_agent | 68.7 | — |
| SWE Atlas - QnA 1/2 models | coding_agent | 15.5 | — |
| AI2D_TEST 1/2 models | document_understanding | 92.7 | — |
| ParseBench Text Formatting 1/2 models | document_understanding | 58.3 | — |
| MMMU-Pro 1/2 models | vision_language | 75.3 | — |
| QwenClawBench 1/2 models | coding_agent | 52.6 | — |
| BabyVision 1/2 models | vision_language | — | 65.7 |
| VideoMMMU 1/2 models | video_understanding | 83.7 | — |
| HMMT Nov 25 1/2 models | stem_reasoning | 89.1 | — |
| SWE-MM 1/2 models | coding_agent | — | 38.6 |
| MMLU-Pro 1/2 models | knowledge | 85.2 | — |
| SuperGPQA 1/2 models | knowledge | 64.7 | — |
| OSWorld-Verified 1/2 models | general_agent | — | 84.3 |
| BrowseComp 1/2 models | general_agent | 62 | — |
| MMLU-Redux 1/2 models | knowledge | 93.3 | — |
| NL2Repo-Bench 1/2 models | coding_agent | — | 42.3 |
| WebArena-Verified 1/2 models | general_agent | — | 64.8 |
| ParseBench Text Content 1/2 models | document_understanding | 90.7 | — |
| CoWorkBench 1/2 models | general_agent | — | 70.7 |
| MCPMark 1/2 models | general_agent | 37 | — |
| MLVU 1/2 models | video_understanding | 86.2 | — |
| IFBench 1/2 models | instruction_following | — | 79.5 |
| SWE-bench Verified 1/2 models | coding_agent | 73.4 | — |
| IMOAnswerBench 1/2 models | stem_reasoning | 78.9 | — |
| CountBench 1/2 models | spatial_intelligence | 96.1 | — |
| VideoMME (w/o sub.) 1/2 models | video_understanding | 82.5 | — |
| Terminal-Bench 2.0 1/2 models | coding_agent | 51.5 | — |
| DeepPlanning 1/2 models | general_agent | 25.9 | — |
| QwenSWEBench 1/2 models | coding_agent | — | 79 |
| QwenWebBench 1/2 models | coding_agent | 1397 | — |
| ODInW13 1/2 models | spatial_intelligence | 50.8 | — |
| VlmsAreBlind 1/2 models | vision_language | 96.6 | — |
| DeepSWE 1/2 models | coding_agent | 0 | — |
| HallusionBench 1/2 models | vision_language | 69.8 | — |
| OCRBench 1/2 models | document_understanding | 90 | — |
| HMMT Feb 26 1/2 models | stem_reasoning | 83.6 | — |
| CC-OCR 1/2 models | document_understanding | 81.9 | — |
| RefCOCO (avg) 1/2 models | spatial_intelligence | 92 | — |
| JobBench 1/2 models | general_agent | — | 33.4 |
| Artificial Analysis Intelligence Index 1/2 models | composite | 32 | — |
| Tool Decathlon 1/2 models | general_agent | 26.9 | — |
| Frontier-Bench v0.1 1/2 models | coding_agent | 1.4 | — |
| OSWorld 2.0 (Partial) 1/2 models | general_agent | — | 48 |
| ParseBench Mean 1/2 models | document_understanding | 44.1 | — |
| OSWorld 2.0 (Binary) 1/2 models | general_agent | — | 19.4 |
| ZEROBench_sub 1/2 models | vision_language | 34.4 | — |
| C-Eval 1/2 models | knowledge | 90 | — |
| WideSearch 1/2 models | general_agent | 60.1 | — |
| SimpleVQA 1/2 models | vision_language | 58.9 | — |
| VideoMME (w sub.) 1/2 models | video_understanding | 86.6 | — |
| SkillsBench Avg5 1/2 models | coding_agent | 28.7 | — |
| Vision2Web 1/2 models | vision_language | — | 62.9 |
| MCP-Atlas 1/2 models | general_agent | 62.8 | — |
| SWE Atlas - RF 1/2 models | coding_agent | 11.4 | — |
| V-Star 1/2 models | vision_language | 90.1 | — |
| DeepSWE 1.1 1/2 models | coding_agent | — | 42.2 |
| Terminal-Bench 2.1 (Claude Code) 1/2 models | coding_agent | 49.2 | — |
| RefSpatialBench 1/2 models | spatial_intelligence | 64.3 | — |
| MMMU 1/2 models | vision_language | 81.7 | — |
| MVBench 1/2 models | video_understanding | 74.6 | — |