Qwen3.8-Flash-Next vs Qwen3.8-27B

New Comparison

Scores are fetched live from the current database on every visit — nothing is cached.

Benchmark Scores Comparison

Qwen3.8-Flash-Next Qwen3.8-27B

Qwen3.8-Flash-Next

125.0B params

Score: 70.0

Qwen3.8-27B

27.0B params

Score: 100.0

Coverage matrix

2 models × 35 benchmarks — hover a benchmark name for its category

Full Coverage Matrix
Shading = score relative to the best score among the selected models: Low Mid High Missing (—) The totals row shows how many of the 35 benchmarks each model has scores for.
Benchmark Qwen3.8-27B Qwen Team (Alibaba Cloud) Qwen3.8-Flash-Next Qwen Team (Alibaba)
GPQA Diamond 89.2 91.7
IFBench (prompt loose) 79.5 —
LiveCodeBench v6 90.3 91.9
SWE-bench Pro 61.7 62.5
RealWorldQA 85.9 88.5
Agents' Last Exam (Score) 42.9 51.2
MathVision 90 90.6
AndroidWorld 81.9 84.5
Agents' Last Exam 20.4 24.3
Toolathlon Verified 67.1 73.5
RecreationBench 47.1 49.9
NL2Repo 42.3 —
BabyVision 65.7 —
Claw-Eval Avg 56.9 60.4
SWE-MM 38.6 —
OSWorld-Verified 84.3 —
HLE (with tools) 30.8 35.9
LVBench 72.4 76.6
CharXiv (RQ) 83.7 84.6
NL2Repo-Bench 42.3 48.1
WebArena-Verified 64.8 —
CoWorkBench 70.7 73.9
SWE-bench Multilingual 73.8 81
IFBench 79.5 81.3
Humanity's Last Exam 30.8 —
ERQA 65.5 72.3
QwenSWEBench 79 —
Terminal-Bench 2.1 (Terminus-2) 73 —
OmniDocBench 1.5 91.1 —
JobBench 33.4 55.7
OSWorld 2.0 (Partial) 48 52.3
OSWorld 2.0 (Binary) 19.4 19.4
Vision2Web 62.9 64
Claw-Eval Pass^3 57.4 64.4
DeepSWE 1.1 42.2 58.7
Covered 35/35 25/35

Fully comparable benchmarks

Scored by all 2 selected models (25)

Benchmark Category Qwen3.8-Flash-Next Qwen3.8-27B
GPQA Diamond stem_reasoning 91.7 89.2
LiveCodeBench v6 stem_reasoning 91.9 90.3
SWE-bench Pro coding_agent 62.5 61.7
RealWorldQA vision_language 88.5 85.9
Agents' Last Exam (Score) general_agent 51.2 42.9
MathVision vision_language 90.6 90
AndroidWorld general_agent 84.5 81.9
Agents' Last Exam general_agent 24.3 20.4
Toolathlon Verified general_agent 73.5 67.1
RecreationBench general_agent 49.9 47.1
Claw-Eval Avg coding_agent 60.4 56.9
HLE (with tools) stem_reasoning 35.9 30.8
LVBench video_understanding 76.6 72.4
CharXiv (RQ) document_understanding 84.6 83.7
NL2Repo-Bench coding_agent 48.1 42.3
CoWorkBench general_agent 73.9 70.7
SWE-bench Multilingual coding_agent 81 73.8
IFBench instruction_following 81.3 79.5
ERQA spatial_intelligence 72.3 65.5
JobBench general_agent 55.7 33.4
OSWorld 2.0 (Partial) general_agent 52.3 48
OSWorld 2.0 (Binary) general_agent 19.4 19.4
Vision2Web vision_language 64 62.9
Claw-Eval Pass^3 coding_agent 64.4 57.4
DeepSWE 1.1 coding_agent 58.7 42.2

Partial coverage benchmarks

Scored by at least one but not all selected models (10) — missing scores show as —

Benchmark Category Qwen3.8-Flash-Next Qwen3.8-27B
IFBench (prompt loose) 1/2 models instruction_following — 79.5
NL2Repo 1/2 models coding_agent — 42.3
BabyVision 1/2 models vision_language — 65.7
SWE-MM 1/2 models coding_agent — 38.6
OSWorld-Verified 1/2 models general_agent — 84.3
WebArena-Verified 1/2 models general_agent — 64.8
Humanity's Last Exam 1/2 models stem_reasoning — 30.8
QwenSWEBench 1/2 models coding_agent — 79
Terminal-Bench 2.1 (Terminus-2) 1/2 models coding_agent — 73
OmniDocBench 1.5 1/2 models document_understanding — 91.1