Qwen3.8-Flash-Next vs Qwen3.8-27B
New ComparisonScores are fetched live from the current database on every visit — nothing is cached.
Benchmark Scores Comparison
Qwen3.8-Flash-Next
Qwen3.8-27B
Qwen3.8-Flash-Next
125.0B params
Score: 70.0
Qwen3.8-27B
27.0B params
Score: 100.0
Coverage matrix
2 models × 35 benchmarks — hover a benchmark name for its category
Shading = score relative to the best score among the selected models:
Low
Mid
High
Missing (—)
The totals row shows how many of the 35 benchmarks each model has scores for.
| Benchmark | Qwen3.8-27B Qwen Team (Alibaba Cloud) | Qwen3.8-Flash-Next Qwen Team (Alibaba) |
|---|---|---|
| GPQA Diamond | 89.2 | 91.7 |
| IFBench (prompt loose) | 79.5 | — |
| LiveCodeBench v6 | 90.3 | 91.9 |
| SWE-bench Pro | 61.7 | 62.5 |
| RealWorldQA | 85.9 | 88.5 |
| Agents' Last Exam (Score) | 42.9 | 51.2 |
| MathVision | 90 | 90.6 |
| AndroidWorld | 81.9 | 84.5 |
| Agents' Last Exam | 20.4 | 24.3 |
| Toolathlon Verified | 67.1 | 73.5 |
| RecreationBench | 47.1 | 49.9 |
| NL2Repo | 42.3 | — |
| BabyVision | 65.7 | — |
| Claw-Eval Avg | 56.9 | 60.4 |
| SWE-MM | 38.6 | — |
| OSWorld-Verified | 84.3 | — |
| HLE (with tools) | 30.8 | 35.9 |
| LVBench | 72.4 | 76.6 |
| CharXiv (RQ) | 83.7 | 84.6 |
| NL2Repo-Bench | 42.3 | 48.1 |
| WebArena-Verified | 64.8 | — |
| CoWorkBench | 70.7 | 73.9 |
| SWE-bench Multilingual | 73.8 | 81 |
| IFBench | 79.5 | 81.3 |
| Humanity's Last Exam | 30.8 | — |
| ERQA | 65.5 | 72.3 |
| QwenSWEBench | 79 | — |
| Terminal-Bench 2.1 (Terminus-2) | 73 | — |
| OmniDocBench 1.5 | 91.1 | — |
| JobBench | 33.4 | 55.7 |
| OSWorld 2.0 (Partial) | 48 | 52.3 |
| OSWorld 2.0 (Binary) | 19.4 | 19.4 |
| Vision2Web | 62.9 | 64 |
| Claw-Eval Pass^3 | 57.4 | 64.4 |
| DeepSWE 1.1 | 42.2 | 58.7 |
| Covered | 35/35 | 25/35 |
Fully comparable benchmarks
Scored by all 2 selected models (25)
| Benchmark | Category | Qwen3.8-Flash-Next | Qwen3.8-27B |
|---|---|---|---|
| GPQA Diamond | stem_reasoning | 91.7 | 89.2 |
| LiveCodeBench v6 | stem_reasoning | 91.9 | 90.3 |
| SWE-bench Pro | coding_agent | 62.5 | 61.7 |
| RealWorldQA | vision_language | 88.5 | 85.9 |
| Agents' Last Exam (Score) | general_agent | 51.2 | 42.9 |
| MathVision | vision_language | 90.6 | 90 |
| AndroidWorld | general_agent | 84.5 | 81.9 |
| Agents' Last Exam | general_agent | 24.3 | 20.4 |
| Toolathlon Verified | general_agent | 73.5 | 67.1 |
| RecreationBench | general_agent | 49.9 | 47.1 |
| Claw-Eval Avg | coding_agent | 60.4 | 56.9 |
| HLE (with tools) | stem_reasoning | 35.9 | 30.8 |
| LVBench | video_understanding | 76.6 | 72.4 |
| CharXiv (RQ) | document_understanding | 84.6 | 83.7 |
| NL2Repo-Bench | coding_agent | 48.1 | 42.3 |
| CoWorkBench | general_agent | 73.9 | 70.7 |
| SWE-bench Multilingual | coding_agent | 81 | 73.8 |
| IFBench | instruction_following | 81.3 | 79.5 |
| ERQA | spatial_intelligence | 72.3 | 65.5 |
| JobBench | general_agent | 55.7 | 33.4 |
| OSWorld 2.0 (Partial) | general_agent | 52.3 | 48 |
| OSWorld 2.0 (Binary) | general_agent | 19.4 | 19.4 |
| Vision2Web | vision_language | 64 | 62.9 |
| Claw-Eval Pass^3 | coding_agent | 64.4 | 57.4 |
| DeepSWE 1.1 | coding_agent | 58.7 | 42.2 |
Partial coverage benchmarks
Scored by at least one but not all selected models (10) — missing scores show as —
| Benchmark | Category | Qwen3.8-Flash-Next | Qwen3.8-27B |
|---|---|---|---|
| IFBench (prompt loose) 1/2 models | instruction_following | — | 79.5 |
| NL2Repo 1/2 models | coding_agent | — | 42.3 |
| BabyVision 1/2 models | vision_language | — | 65.7 |
| SWE-MM 1/2 models | coding_agent | — | 38.6 |
| OSWorld-Verified 1/2 models | general_agent | — | 84.3 |
| WebArena-Verified 1/2 models | general_agent | — | 64.8 |
| Humanity's Last Exam 1/2 models | stem_reasoning | — | 30.8 |
| QwenSWEBench 1/2 models | coding_agent | — | 79 |
| Terminal-Bench 2.1 (Terminus-2) 1/2 models | coding_agent | — | 73 |
| OmniDocBench 1.5 1/2 models | document_understanding | — | 91.1 |