GLM-5.2 vs GLM-5.3-Flash

New Comparison

Scores are fetched live from the current database on every visit — nothing is cached.

Benchmark Scores Comparison

GLM-5.2 GLM-5.3-Flash

GLM-5.2

753.0B params

Score: 100.0

GLM-5.3-Flash

320.0B params

Score: 100.0

Coverage matrix

2 models × 37 benchmarks — hover a benchmark name for its category

Full Coverage Matrix
Shading = score relative to the best score among the selected models: Low Mid High Missing (—) The totals row shows how many of the 37 benchmarks each model has scores for.
Benchmark GLM-5.2 Zhipu AI GLM-5.3-Flash Zhipu AI
GPQA Diamond 91.2 —
Automation-Bench 12.9 48.8
SWE-bench Pro 62.1 —
AIME 26 99.2 —
Agents' Last Exam 23.8 26.3
Toolathlon Verified 59.9 —
WildClawBench 54.2 —
NL2Repo 48.9 —
PostTrainBench 34.3 —
DSBench-FullStack 61.8 —
Terminal-Bench 3.0 4.6 —
HMMT Nov 25 94.4 —
ExploitGym (6h) 39 —
HellaSwag — 87.1
HLE (with tools) 54.7 55.3
FrontierSWE 74.4 —
CritPt (no tools) 20.9 —
MMLU — 88.1
IMOAnswerBench 91 —
Humanity's Last Exam 40.5 —
DSBench-Hard 54.5 —
BIG-Bench Hard — 86.6
ExploitGym (2h) 29 —
Terminal-Bench 2.1 (Terminus-2) 81 —
DeepSWE 46.2 —
ExploitBench 24.4 —
HMMT Feb 26 92.5 —
Terminal Bench 2.1 81 84.3
Artificial Analysis Intelligence Index — 57
Tool Decathlon 48.2 —
ProgramBench 63.7 —
MCP-Atlas 76.8 —
SWE-Marathon 13 —
Terminal-Bench 2.1 (Best Reported Harness) 82.7 —
Cybergym 77.2 —
DeepSWE 1.1 46.2 63.4
GDPVal-AA v2 1504 1773
Covered 33/37 10/37

Fully comparable benchmarks

Scored by all 2 selected models (6)

Benchmark Category GLM-5.2 GLM-5.3-Flash
Automation-Bench general_agent 12.9 48.8
Agents' Last Exam general_agent 23.8 26.3
HLE (with tools) stem_reasoning 54.7 55.3
Terminal Bench 2.1 coding_agent 81 84.3
DeepSWE 1.1 coding_agent 46.2 63.4
GDPVal-AA v2 general_agent 1504 1773

Partial coverage benchmarks

Scored by at least one but not all selected models (31) — missing scores show as —

Benchmark Category GLM-5.2 GLM-5.3-Flash
GPQA Diamond 1/2 models stem_reasoning 91.2 —
SWE-bench Pro 1/2 models coding_agent 62.1 —
AIME 26 1/2 models stem_reasoning 99.2 —
Toolathlon Verified 1/2 models general_agent 59.9 —
WildClawBench 1/2 models coding_agent 54.2 —
NL2Repo 1/2 models coding_agent 48.9 —
PostTrainBench 1/2 models coding_agent 34.3 —
DSBench-FullStack 1/2 models coding_agent 61.8 —
Terminal-Bench 3.0 1/2 models coding_agent 4.6 —
HMMT Nov 25 1/2 models stem_reasoning 94.4 —
ExploitGym (6h) 1/2 models cybersecurity 39 —
HellaSwag 1/2 models reasoning — 87.1
FrontierSWE 1/2 models coding_agent 74.4 —
CritPt (no tools) 1/2 models stem_reasoning 20.9 —
MMLU 1/2 models knowledge — 88.1
IMOAnswerBench 1/2 models stem_reasoning 91 —
Humanity's Last Exam 1/2 models stem_reasoning 40.5 —
DSBench-Hard 1/2 models coding_agent 54.5 —
BIG-Bench Hard 1/2 models reasoning — 86.6
ExploitGym (2h) 1/2 models cybersecurity 29 —
Terminal-Bench 2.1 (Terminus-2) 1/2 models coding_agent 81 —
DeepSWE 1/2 models coding_agent 46.2 —
ExploitBench 1/2 models cybersecurity 24.4 —
HMMT Feb 26 1/2 models stem_reasoning 92.5 —
Artificial Analysis Intelligence Index 1/2 models composite — 57
Tool Decathlon 1/2 models general_agent 48.2 —
ProgramBench 1/2 models coding_agent 63.7 —
MCP-Atlas 1/2 models general_agent 76.8 —
SWE-Marathon 1/2 models coding_agent 13 —
Terminal-Bench 2.1 (Best Reported Harness) 1/2 models coding_agent 82.7 —
Cybergym 1/2 models general_agent 77.2 —