Qwen3.6-27B vs Gemma3-27B vs Qwen3.8-27B
New ComparisonScores are fetched live from the current database on every visit — nothing is cached.
Benchmark Scores Comparison
Qwen3.6-27B
Gemma3-27B
Qwen3.8-27B
Qwen3.6-27B
27.0B params
Score: 100.0
Gemma3-27B
27.0B params
Score: 100.0
Qwen3.8-27B
27.0B params
Score: 100.0
Coverage matrix
3 models × 90 benchmarks — hover a benchmark name for its category
Shading = score relative to the best score among the selected models:
Low
Mid
High
Missing (—)
The totals row shows how many of the 90 benchmarks each model has scores for.
| Benchmark | Gemma3-27B Google | Qwen3.6-27B Qwen Team (Alibaba Cloud) | Qwen3.8-27B Qwen Team (Alibaba Cloud) |
|---|---|---|---|
| HMMT Feb 25 | — | 93.8 | — |
| DynaMath | — | 85.6 | — |
| GPQA Diamond | 42.4 | 87.8 | 89.2 |
| IFBench (prompt loose) | — | 69.1 | 79.5 |
| MathVista (mini) | — | 87.4 | — |
| MMBench EN-DEV v1.1 | — | 92.3 | — |
| LiveCodeBench v6 | 29.1 | 83.9 | 90.3 |
| SWE-bench Pro | — | 53.5 | 61.7 |
| RealWorldQA | — | 84.1 | 85.9 |
| Agents' Last Exam (Score) | — | — | 42.9 |
| EmbSpatialBench | — | 84.6 | — |
| AA-LCR | — | 73.3 | — |
| MathVision | 46 | 85.1 | 90 |
| AIME 26 | 20.8 | 94.1 | — |
| AndroidWorld | — | 70.3 | 81.9 |
| Agents' Last Exam | — | 10.6 | 20.4 |
| Toolathlon Verified | — | — | 67.1 |
| RecreationBench | — | 29.8 | 47.1 |
| ScreenSpot Pro | — | 76.1 | — |
| WMDP (Bio) | — | 84.8 | — |
| MMStar | — | 81.4 | — |
| DeepSearch QA | — | 71.1 | — |
| WildClawBench | — | 43.2 | — |
| NL2Repo | — | 36.2 | 42.3 |
| HPCT | — | 48.7 | — |
| MMMU-Pro | 49.7 | 75.8 | — |
| QwenClawBench | — | 53.4 | — |
| BabyVision | — | 28.9 | 65.7 |
| MMMLU | 70.7 | — | — |
| Claw-Eval Avg | — | 72.4 | 56.9 |
| VideoMMMU | — | 84.4 | — |
| HMMT Nov 25 | — | 90.7 | — |
| SWE-MM | — | 25.7 | 38.6 |
| MMLU-Pro | 67.6 | 86.2 | — |
| CodeForces | 110 | — | — |
| SuperGPQA | — | 66 | — |
| OSWorld-Verified | — | 63.9 | 84.3 |
| HLE (with tools) | — | — | 30.8 |
| LVBench | — | — | 72.4 |
| CharXiv (RQ) | — | 78.4 | 83.7 |
| MRCR v2 128K (8-needle) | 13.5 | — | — |
| MMLU-Redux | — | 93.5 | — |
| NL2Repo-Bench | — | — | 42.3 |
| SkillsBench | — | 46.6 | — |
| WebArena-Verified | — | 48.8 | 64.8 |
| CoWorkBench | — | 61 | 70.7 |
| Lab Bench (ProtocolQA) | — | 69.1 | — |
| SWE-bench Multilingual | — | 71.3 | 73.8 |
| MLVU | — | 86.6 | — |
| Beam128K | — | 63 | — |
| IFBench | — | 70.8 | 79.5 |
| SWE-bench Verified | — | 77.2 | — |
| IMOAnswerBench | — | 80.8 | — |
| Humanity's Last Exam | — | 24 | 30.8 |
| CountBench | — | 97.8 | — |
| WMDP (Chem) | — | 74.8 | — |
| ERQA | — | 62.5 | 65.5 |
| Terminal-Bench 2.0 | — | 59.3 | — |
| QwenSWEBench | — | 49.3 | 79 |
| QwenWebBench | — | 1487 | — |
| TAU2-Bench | 16.2 | — | — |
| VlmsAreBlind | — | 97 | — |
| Terminal-Bench 2.1 (Terminus-2) | — | 63.4 | 73 |
| OmniDocBench 1.5 | 0.365 | 89.4 | 91.1 |
| OCRBench | — | 89.4 | — |
| HMMT Feb 26 | — | 84.3 | — |
| CC-OCR | — | 81.2 | — |
| RefCOCO (avg) | — | 92.5 | — |
| JobBench | — | 21.8 | 33.4 |
| SciCode (subtask) | — | 39.8 | — |
| OSWorld 2.0 (Partial) | — | — | 48 |
| OSWorld 2.0 (Binary) | — | — | 19.4 |
| C-Eval | — | 91.4 | — |
| SimpleVQA | — | 56.1 | — |
| VideoMME (w sub.) | — | 87.7 | — |
| SkillsBench Avg5 | — | 48.2 | — |
| Vision2Web | — | 45 | 62.9 |
| TauBench V3 Banking | — | 16.7 | — |
| Claw-Eval Pass^3 | — | 60.6 | 57.4 |
| MCP-Atlas | — | 62.5 | — |
| V-Star | — | 94.7 | — |
| DeepSWE 1.1 | — | 13.3 | 42.2 |
| MBCT | — | 45.9 | — |
| BigBench Extra Hard | 19.3 | — | — |
| VCT | — | 33.7 | — |
| Gaia2 | — | 40 | — |
| RefSpatialBench | — | 70 | — |
| MMMU | — | 82.9 | — |
| GDPVal-AA v2 | — | 1141 | — |
| MVBench | — | 75.5 | — |
| Covered | 12/90 | 78/90 | 35/90 |
Fully comparable benchmarks
Scored by all 3 selected models (4)
| Benchmark | Category | Qwen3.6-27B | Gemma3-27B | Qwen3.8-27B |
|---|---|---|---|---|
| GPQA Diamond | stem_reasoning | 87.8 | 42.4 | 89.2 |
| LiveCodeBench v6 | stem_reasoning | 83.9 | 29.1 | 90.3 |
| MathVision | vision_language | 85.1 | 46 | 90 |
| OmniDocBench 1.5 | document_understanding | 89.4 | 0.365 | 91.1 |
Partial coverage benchmarks
Scored by at least one but not all selected models (86) — missing scores show as —
| Benchmark | Category | Qwen3.6-27B | Gemma3-27B | Qwen3.8-27B |
|---|---|---|---|---|
| HMMT Feb 25 1/3 models | stem_reasoning | 93.8 | — | — |
| DynaMath 1/3 models | vision_language | 85.6 | — | — |
| IFBench (prompt loose) 2/3 models | instruction_following | 69.1 | — | 79.5 |
| MathVista (mini) 1/3 models | vision_language | 87.4 | — | — |
| MMBench EN-DEV v1.1 1/3 models | vision_language | 92.3 | — | — |
| SWE-bench Pro 2/3 models | coding_agent | 53.5 | — | 61.7 |
| RealWorldQA 2/3 models | vision_language | 84.1 | — | 85.9 |
| Agents' Last Exam (Score) 1/3 models | general_agent | — | — | 42.9 |
| EmbSpatialBench 1/3 models | spatial_intelligence | 84.6 | — | — |
| AA-LCR 1/3 models | long_context | 73.3 | — | — |
| AIME 26 2/3 models | stem_reasoning | 94.1 | 20.8 | — |
| AndroidWorld 2/3 models | general_agent | 70.3 | — | 81.9 |
| Agents' Last Exam 2/3 models | general_agent | 10.6 | — | 20.4 |
| Toolathlon Verified 1/3 models | general_agent | — | — | 67.1 |
| RecreationBench 2/3 models | general_agent | 29.8 | — | 47.1 |
| ScreenSpot Pro 1/3 models | general_agent | 76.1 | — | — |
| WMDP (Bio) 1/3 models | safety | 84.8 | — | — |
| MMStar 1/3 models | vision_language | 81.4 | — | — |
| DeepSearch QA 1/3 models | general_agent | 71.1 | — | — |
| WildClawBench 1/3 models | coding_agent | 43.2 | — | — |
| NL2Repo 2/3 models | coding_agent | 36.2 | — | 42.3 |
| HPCT 1/3 models | safety | 48.7 | — | — |
| MMMU-Pro 2/3 models | vision_language | 75.8 | 49.7 | — |
| QwenClawBench 1/3 models | coding_agent | 53.4 | — | — |
| BabyVision 2/3 models | vision_language | 28.9 | — | 65.7 |
| MMMLU 1/3 models | multilingual | — | 70.7 | — |
| Claw-Eval Avg 2/3 models | coding_agent | 72.4 | — | 56.9 |
| VideoMMMU 1/3 models | video_understanding | 84.4 | — | — |
| HMMT Nov 25 1/3 models | stem_reasoning | 90.7 | — | — |
| SWE-MM 2/3 models | coding_agent | 25.7 | — | 38.6 |
| MMLU-Pro 2/3 models | knowledge | 86.2 | 67.6 | — |
| CodeForces 1/3 models | stem_reasoning | — | 110 | — |
| SuperGPQA 1/3 models | knowledge | 66 | — | — |
| OSWorld-Verified 2/3 models | general_agent | 63.9 | — | 84.3 |
| HLE (with tools) 1/3 models | stem_reasoning | — | — | 30.8 |
| LVBench 1/3 models | video_understanding | — | — | 72.4 |
| CharXiv (RQ) 2/3 models | document_understanding | 78.4 | — | 83.7 |
| MRCR v2 128K (8-needle) 1/3 models | long_context | — | 13.5 | — |
| MMLU-Redux 1/3 models | knowledge | 93.5 | — | — |
| NL2Repo-Bench 1/3 models | coding_agent | — | — | 42.3 |
| SkillsBench 1/3 models | general_agent | 46.6 | — | — |
| WebArena-Verified 2/3 models | general_agent | 48.8 | — | 64.8 |
| CoWorkBench 2/3 models | general_agent | 61 | — | 70.7 |
| Lab Bench (ProtocolQA) 1/3 models | safety | 69.1 | — | — |
| SWE-bench Multilingual 2/3 models | coding_agent | 71.3 | — | 73.8 |
| MLVU 1/3 models | video_understanding | 86.6 | — | — |
| Beam128K 1/3 models | long_context | 63 | — | — |
| IFBench 2/3 models | instruction_following | 70.8 | — | 79.5 |
| SWE-bench Verified 1/3 models | coding_agent | 77.2 | — | — |
| IMOAnswerBench 1/3 models | stem_reasoning | 80.8 | — | — |
| Humanity's Last Exam 2/3 models | stem_reasoning | 24 | — | 30.8 |
| CountBench 1/3 models | spatial_intelligence | 97.8 | — | — |
| WMDP (Chem) 1/3 models | safety | 74.8 | — | — |
| ERQA 2/3 models | spatial_intelligence | 62.5 | — | 65.5 |
| Terminal-Bench 2.0 1/3 models | coding_agent | 59.3 | — | — |
| QwenSWEBench 2/3 models | coding_agent | 49.3 | — | 79 |
| QwenWebBench 1/3 models | coding_agent | 1487 | — | — |
| TAU2-Bench 1/3 models | general_agent | — | 16.2 | — |
| VlmsAreBlind 1/3 models | vision_language | 97 | — | — |
| Terminal-Bench 2.1 (Terminus-2) 2/3 models | coding_agent | 63.4 | — | 73 |
| OCRBench 1/3 models | document_understanding | 89.4 | — | — |
| HMMT Feb 26 1/3 models | stem_reasoning | 84.3 | — | — |
| CC-OCR 1/3 models | document_understanding | 81.2 | — | — |
| RefCOCO (avg) 1/3 models | spatial_intelligence | 92.5 | — | — |
| JobBench 2/3 models | general_agent | 21.8 | — | 33.4 |
| SciCode (subtask) 1/3 models | stem_reasoning | 39.8 | — | — |
| OSWorld 2.0 (Partial) 1/3 models | general_agent | — | — | 48 |
| OSWorld 2.0 (Binary) 1/3 models | general_agent | — | — | 19.4 |
| C-Eval 1/3 models | knowledge | 91.4 | — | — |
| SimpleVQA 1/3 models | vision_language | 56.1 | — | — |
| VideoMME (w sub.) 1/3 models | video_understanding | 87.7 | — | — |
| SkillsBench Avg5 1/3 models | coding_agent | 48.2 | — | — |
| Vision2Web 2/3 models | vision_language | 45 | — | 62.9 |
| TauBench V3 Banking 1/3 models | general_agent | 16.7 | — | — |
| Claw-Eval Pass^3 2/3 models | coding_agent | 60.6 | — | 57.4 |
| MCP-Atlas 1/3 models | general_agent | 62.5 | — | — |
| V-Star 1/3 models | vision_language | 94.7 | — | — |
| DeepSWE 1.1 2/3 models | coding_agent | 13.3 | — | 42.2 |
| MBCT 1/3 models | safety | 45.9 | — | — |
| BigBench Extra Hard 1/3 models | stem_reasoning | — | 19.3 | — |
| VCT 1/3 models | safety | 33.7 | — | — |
| Gaia2 1/3 models | general_agent | 40 | — | — |
| RefSpatialBench 1/3 models | spatial_intelligence | 70 | — | — |
| MMMU 1/3 models | vision_language | 82.9 | — | — |
| GDPVal-AA v2 1/3 models | general_agent | 1141 | — | — |
| MVBench 1/3 models | video_understanding | 75.5 | — | — |