GLM-5.2 vs GLM-5.3 vs GLM-5.3-Flash

New Comparison

Scores are fetched live from the current database on every visit — nothing is cached.

Benchmark Scores Comparison

GLM-5.2 GLM-5.3 GLM-5.3-Flash

GLM-5.2

753.0B params

Score: 100.0

GLM-5.3

753.3B params

Score: 70.0

GLM-5.3-Flash

320.0B params

Score: 100.0

Coverage matrix

3 models × 62 benchmarks — hover a benchmark name for its category

Full Coverage Matrix
Shading = score relative to the best score among the selected models: Low Mid High Missing (—) The totals row shows how many of the 62 benchmarks each model has scores for.
Benchmark GLM-5.2 Zhipu AI GLM-5.3 Zhipu AI GLM-5.3-Flash Zhipu AI
GPQA Diamond 91.2 91.7 —
Hy-LifeSearch (Internal) — 49.2 —
Automation-Bench 12.9 48.2 48.8
SWE-bench Pro 62.1 64.6 —
BioMysteryBench — 69 —
AIME 26 99.2 — —
Agents' Last Exam 23.8 28.5 26.3
Toolathlon Verified 59.9 73 —
Hy-Backend 2.0 (Internal) — 41.9 —
Hy-SWE Max Verified (Internal) — 67.2 —
Hy-BrowseComp-Pro2 (Internal) — 48.4 —
OfficeQA Pro — 66.2 —
SWE Atlas - TW — 49.6 —
WildClawBench 54.2 — —
SWE Atlas - QnA — 55.8 —
NL2Repo 48.9 58 —
PostTrainBench 34.3 39.8 —
DSBench-FullStack 61.8 — —
Terminal-Bench 3.0 4.6 28.3 —
HMMT Nov 25 94.4 — —
Apex-Agents — 38.1 —
WorkSpaceBench — 68.2 —
BankerToolBench — 77.8 —
E-Bench-Code (Internal) — 66.5 —
ExploitGym (6h) 39 130 —
HellaSwag — — 87.1
HLE (with tools) 54.7 62.5 55.3
FrontierSWE 74.4 78.1 —
CritPt (no tools) 20.9 19.1 —
DRACO — 78.1 —
SWE-bench Multilingual — 81.3 —
MMLU — — 88.1
Hy-FinAgentBench (Internal) — 80.4 —
IMOAnswerBench 91 — —
Humanity's Last Exam 40.5 42.3 —
DSBench-Hard 54.5 — —
BIG-Bench Hard — — 86.6
ExploitGym (2h) 29 105 —
SUPERChem — 58.5 —
Terminal-Bench 2.1 (Terminus-2) 81 88.2 —
DeepSWE 46.2 66.9 —
ExploitBench 24.4 54.4 —
HMMT Feb 26 92.5 — —
Hy-FinmodelBench v2 (Internal) — 57.8 —
Terminal Bench 2.1 81 88.2 84.3
JobBench — 58.2 —
Artificial Analysis Intelligence Index — — 57
$OneMillion-Bench — 64.5 —
Tool Decathlon 48.2 — —
WideSearch — 83.2 —
SkillsBench Avg5 — 63.3 —
E-Bench (Internal) — 71.4 —
ProgramBench 63.7 19 —
MCP-Atlas 76.8 81.9 —
SWE-Marathon 13 42.5 —
Terminal-Bench 2.1 (Best Reported Harness) 82.7 — —
Cybergym 77.2 84.5 —
SWE Atlas - RF — 51.9 —
DeepSWE 1.1 46.2 66.9 63.4
Hy-CompanyBench V2 (Internal) — 64.5 —
Harbor-Index — 42.5 —
GDPVal-AA v2 1504 1769 1773
Covered 33/62 49/62 10/62

Fully comparable benchmarks

Scored by all 3 selected models (6)

Benchmark Category GLM-5.2 GLM-5.3 GLM-5.3-Flash
Automation-Bench general_agent 12.9 48.2 48.8
Agents' Last Exam general_agent 23.8 28.5 26.3
HLE (with tools) stem_reasoning 54.7 62.5 55.3
Terminal Bench 2.1 coding_agent 81 88.2 84.3
DeepSWE 1.1 coding_agent 46.2 66.9 63.4
GDPVal-AA v2 general_agent 1504 1769 1773

Partial coverage benchmarks

Scored by at least one but not all selected models (56) — missing scores show as —

Benchmark Category GLM-5.2 GLM-5.3 GLM-5.3-Flash
GPQA Diamond 2/3 models stem_reasoning 91.2 91.7 —
Hy-LifeSearch (Internal) 1/3 models general_agent — 49.2 —
SWE-bench Pro 2/3 models coding_agent 62.1 64.6 —
BioMysteryBench 1/3 models stem_reasoning — 69 —
AIME 26 1/3 models stem_reasoning 99.2 — —
Toolathlon Verified 2/3 models general_agent 59.9 73 —
Hy-Backend 2.0 (Internal) 1/3 models coding_agent — 41.9 —
Hy-SWE Max Verified (Internal) 1/3 models coding_agent — 67.2 —
Hy-BrowseComp-Pro2 (Internal) 1/3 models general_agent — 48.4 —
OfficeQA Pro 1/3 models general_agent — 66.2 —
SWE Atlas - TW 1/3 models coding_agent — 49.6 —
WildClawBench 1/3 models coding_agent 54.2 — —
SWE Atlas - QnA 1/3 models coding_agent — 55.8 —
NL2Repo 2/3 models coding_agent 48.9 58 —
PostTrainBench 2/3 models coding_agent 34.3 39.8 —
DSBench-FullStack 1/3 models coding_agent 61.8 — —
Terminal-Bench 3.0 2/3 models coding_agent 4.6 28.3 —
HMMT Nov 25 1/3 models stem_reasoning 94.4 — —
Apex-Agents 1/3 models general_agent — 38.1 —
WorkSpaceBench 1/3 models general_agent — 68.2 —
BankerToolBench 1/3 models general_agent — 77.8 —
E-Bench-Code (Internal) 1/3 models coding_agent — 66.5 —
ExploitGym (6h) 2/3 models cybersecurity 39 130 —
HellaSwag 1/3 models reasoning — — 87.1
FrontierSWE 2/3 models coding_agent 74.4 78.1 —
CritPt (no tools) 2/3 models stem_reasoning 20.9 19.1 —
DRACO 1/3 models general_agent — 78.1 —
SWE-bench Multilingual 1/3 models coding_agent — 81.3 —
MMLU 1/3 models knowledge — — 88.1
Hy-FinAgentBench (Internal) 1/3 models domain_finance — 80.4 —
IMOAnswerBench 1/3 models stem_reasoning 91 — —
Humanity's Last Exam 2/3 models stem_reasoning 40.5 42.3 —
DSBench-Hard 1/3 models coding_agent 54.5 — —
BIG-Bench Hard 1/3 models reasoning — — 86.6
ExploitGym (2h) 2/3 models cybersecurity 29 105 —
SUPERChem 1/3 models stem_reasoning — 58.5 —
Terminal-Bench 2.1 (Terminus-2) 2/3 models coding_agent 81 88.2 —
DeepSWE 2/3 models coding_agent 46.2 66.9 —
ExploitBench 2/3 models cybersecurity 24.4 54.4 —
HMMT Feb 26 1/3 models stem_reasoning 92.5 — —
Hy-FinmodelBench v2 (Internal) 1/3 models domain_finance — 57.8 —
JobBench 1/3 models general_agent — 58.2 —
Artificial Analysis Intelligence Index 1/3 models composite — — 57
$OneMillion-Bench 1/3 models general_capabilities — 64.5 —
Tool Decathlon 1/3 models general_agent 48.2 — —
WideSearch 1/3 models general_agent — 83.2 —
SkillsBench Avg5 1/3 models coding_agent — 63.3 —
E-Bench (Internal) 1/3 models general_agent — 71.4 —
ProgramBench 2/3 models coding_agent 63.7 19 —
MCP-Atlas 2/3 models general_agent 76.8 81.9 —
SWE-Marathon 2/3 models coding_agent 13 42.5 —
Terminal-Bench 2.1 (Best Reported Harness) 1/3 models coding_agent 82.7 — —
Cybergym 2/3 models general_agent 77.2 84.5 —
SWE Atlas - RF 1/3 models coding_agent — 51.9 —
Hy-CompanyBench V2 (Internal) 1/3 models general_agent — 64.5 —
Harbor-Index 1/3 models coding_agent — 42.5 —