GLM-5.2 vs GLM-5.3 vs GLM-5.3-Flash
New ComparisonScores are fetched live from the current database on every visit — nothing is cached.
Benchmark Scores Comparison
GLM-5.2
GLM-5.3
GLM-5.3-Flash
GLM-5.2
753.0B params
Score: 100.0
GLM-5.3
753.3B params
Score: 70.0
GLM-5.3-Flash
320.0B params
Score: 100.0
Coverage matrix
3 models × 62 benchmarks — hover a benchmark name for its category
Shading = score relative to the best score among the selected models:
Low
Mid
High
Missing (—)
The totals row shows how many of the 62 benchmarks each model has scores for.
| Benchmark | GLM-5.2 Zhipu AI | GLM-5.3 Zhipu AI | GLM-5.3-Flash Zhipu AI |
|---|---|---|---|
| GPQA Diamond | 91.2 | 91.7 | — |
| Hy-LifeSearch (Internal) | — | 49.2 | — |
| Automation-Bench | 12.9 | 48.2 | 48.8 |
| SWE-bench Pro | 62.1 | 64.6 | — |
| BioMysteryBench | — | 69 | — |
| AIME 26 | 99.2 | — | — |
| Agents' Last Exam | 23.8 | 28.5 | 26.3 |
| Toolathlon Verified | 59.9 | 73 | — |
| Hy-Backend 2.0 (Internal) | — | 41.9 | — |
| Hy-SWE Max Verified (Internal) | — | 67.2 | — |
| Hy-BrowseComp-Pro2 (Internal) | — | 48.4 | — |
| OfficeQA Pro | — | 66.2 | — |
| SWE Atlas - TW | — | 49.6 | — |
| WildClawBench | 54.2 | — | — |
| SWE Atlas - QnA | — | 55.8 | — |
| NL2Repo | 48.9 | 58 | — |
| PostTrainBench | 34.3 | 39.8 | — |
| DSBench-FullStack | 61.8 | — | — |
| Terminal-Bench 3.0 | 4.6 | 28.3 | — |
| HMMT Nov 25 | 94.4 | — | — |
| Apex-Agents | — | 38.1 | — |
| WorkSpaceBench | — | 68.2 | — |
| BankerToolBench | — | 77.8 | — |
| E-Bench-Code (Internal) | — | 66.5 | — |
| ExploitGym (6h) | 39 | 130 | — |
| HellaSwag | — | — | 87.1 |
| HLE (with tools) | 54.7 | 62.5 | 55.3 |
| FrontierSWE | 74.4 | 78.1 | — |
| CritPt (no tools) | 20.9 | 19.1 | — |
| DRACO | — | 78.1 | — |
| SWE-bench Multilingual | — | 81.3 | — |
| MMLU | — | — | 88.1 |
| Hy-FinAgentBench (Internal) | — | 80.4 | — |
| IMOAnswerBench | 91 | — | — |
| Humanity's Last Exam | 40.5 | 42.3 | — |
| DSBench-Hard | 54.5 | — | — |
| BIG-Bench Hard | — | — | 86.6 |
| ExploitGym (2h) | 29 | 105 | — |
| SUPERChem | — | 58.5 | — |
| Terminal-Bench 2.1 (Terminus-2) | 81 | 88.2 | — |
| DeepSWE | 46.2 | 66.9 | — |
| ExploitBench | 24.4 | 54.4 | — |
| HMMT Feb 26 | 92.5 | — | — |
| Hy-FinmodelBench v2 (Internal) | — | 57.8 | — |
| Terminal Bench 2.1 | 81 | 88.2 | 84.3 |
| JobBench | — | 58.2 | — |
| Artificial Analysis Intelligence Index | — | — | 57 |
| $OneMillion-Bench | — | 64.5 | — |
| Tool Decathlon | 48.2 | — | — |
| WideSearch | — | 83.2 | — |
| SkillsBench Avg5 | — | 63.3 | — |
| E-Bench (Internal) | — | 71.4 | — |
| ProgramBench | 63.7 | 19 | — |
| MCP-Atlas | 76.8 | 81.9 | — |
| SWE-Marathon | 13 | 42.5 | — |
| Terminal-Bench 2.1 (Best Reported Harness) | 82.7 | — | — |
| Cybergym | 77.2 | 84.5 | — |
| SWE Atlas - RF | — | 51.9 | — |
| DeepSWE 1.1 | 46.2 | 66.9 | 63.4 |
| Hy-CompanyBench V2 (Internal) | — | 64.5 | — |
| Harbor-Index | — | 42.5 | — |
| GDPVal-AA v2 | 1504 | 1769 | 1773 |
| Covered | 33/62 | 49/62 | 10/62 |
Fully comparable benchmarks
Scored by all 3 selected models (6)
| Benchmark | Category | GLM-5.2 | GLM-5.3 | GLM-5.3-Flash |
|---|---|---|---|---|
| Automation-Bench | general_agent | 12.9 | 48.2 | 48.8 |
| Agents' Last Exam | general_agent | 23.8 | 28.5 | 26.3 |
| HLE (with tools) | stem_reasoning | 54.7 | 62.5 | 55.3 |
| Terminal Bench 2.1 | coding_agent | 81 | 88.2 | 84.3 |
| DeepSWE 1.1 | coding_agent | 46.2 | 66.9 | 63.4 |
| GDPVal-AA v2 | general_agent | 1504 | 1769 | 1773 |
Partial coverage benchmarks
Scored by at least one but not all selected models (56) — missing scores show as —
| Benchmark | Category | GLM-5.2 | GLM-5.3 | GLM-5.3-Flash |
|---|---|---|---|---|
| GPQA Diamond 2/3 models | stem_reasoning | 91.2 | 91.7 | — |
| Hy-LifeSearch (Internal) 1/3 models | general_agent | — | 49.2 | — |
| SWE-bench Pro 2/3 models | coding_agent | 62.1 | 64.6 | — |
| BioMysteryBench 1/3 models | stem_reasoning | — | 69 | — |
| AIME 26 1/3 models | stem_reasoning | 99.2 | — | — |
| Toolathlon Verified 2/3 models | general_agent | 59.9 | 73 | — |
| Hy-Backend 2.0 (Internal) 1/3 models | coding_agent | — | 41.9 | — |
| Hy-SWE Max Verified (Internal) 1/3 models | coding_agent | — | 67.2 | — |
| Hy-BrowseComp-Pro2 (Internal) 1/3 models | general_agent | — | 48.4 | — |
| OfficeQA Pro 1/3 models | general_agent | — | 66.2 | — |
| SWE Atlas - TW 1/3 models | coding_agent | — | 49.6 | — |
| WildClawBench 1/3 models | coding_agent | 54.2 | — | — |
| SWE Atlas - QnA 1/3 models | coding_agent | — | 55.8 | — |
| NL2Repo 2/3 models | coding_agent | 48.9 | 58 | — |
| PostTrainBench 2/3 models | coding_agent | 34.3 | 39.8 | — |
| DSBench-FullStack 1/3 models | coding_agent | 61.8 | — | — |
| Terminal-Bench 3.0 2/3 models | coding_agent | 4.6 | 28.3 | — |
| HMMT Nov 25 1/3 models | stem_reasoning | 94.4 | — | — |
| Apex-Agents 1/3 models | general_agent | — | 38.1 | — |
| WorkSpaceBench 1/3 models | general_agent | — | 68.2 | — |
| BankerToolBench 1/3 models | general_agent | — | 77.8 | — |
| E-Bench-Code (Internal) 1/3 models | coding_agent | — | 66.5 | — |
| ExploitGym (6h) 2/3 models | cybersecurity | 39 | 130 | — |
| HellaSwag 1/3 models | reasoning | — | — | 87.1 |
| FrontierSWE 2/3 models | coding_agent | 74.4 | 78.1 | — |
| CritPt (no tools) 2/3 models | stem_reasoning | 20.9 | 19.1 | — |
| DRACO 1/3 models | general_agent | — | 78.1 | — |
| SWE-bench Multilingual 1/3 models | coding_agent | — | 81.3 | — |
| MMLU 1/3 models | knowledge | — | — | 88.1 |
| Hy-FinAgentBench (Internal) 1/3 models | domain_finance | — | 80.4 | — |
| IMOAnswerBench 1/3 models | stem_reasoning | 91 | — | — |
| Humanity's Last Exam 2/3 models | stem_reasoning | 40.5 | 42.3 | — |
| DSBench-Hard 1/3 models | coding_agent | 54.5 | — | — |
| BIG-Bench Hard 1/3 models | reasoning | — | — | 86.6 |
| ExploitGym (2h) 2/3 models | cybersecurity | 29 | 105 | — |
| SUPERChem 1/3 models | stem_reasoning | — | 58.5 | — |
| Terminal-Bench 2.1 (Terminus-2) 2/3 models | coding_agent | 81 | 88.2 | — |
| DeepSWE 2/3 models | coding_agent | 46.2 | 66.9 | — |
| ExploitBench 2/3 models | cybersecurity | 24.4 | 54.4 | — |
| HMMT Feb 26 1/3 models | stem_reasoning | 92.5 | — | — |
| Hy-FinmodelBench v2 (Internal) 1/3 models | domain_finance | — | 57.8 | — |
| JobBench 1/3 models | general_agent | — | 58.2 | — |
| Artificial Analysis Intelligence Index 1/3 models | composite | — | — | 57 |
| $OneMillion-Bench 1/3 models | general_capabilities | — | 64.5 | — |
| Tool Decathlon 1/3 models | general_agent | 48.2 | — | — |
| WideSearch 1/3 models | general_agent | — | 83.2 | — |
| SkillsBench Avg5 1/3 models | coding_agent | — | 63.3 | — |
| E-Bench (Internal) 1/3 models | general_agent | — | 71.4 | — |
| ProgramBench 2/3 models | coding_agent | 63.7 | 19 | — |
| MCP-Atlas 2/3 models | general_agent | 76.8 | 81.9 | — |
| SWE-Marathon 2/3 models | coding_agent | 13 | 42.5 | — |
| Terminal-Bench 2.1 (Best Reported Harness) 1/3 models | coding_agent | 82.7 | — | — |
| Cybergym 2/3 models | general_agent | 77.2 | 84.5 | — |
| SWE Atlas - RF 1/3 models | coding_agent | — | 51.9 | — |
| Hy-CompanyBench V2 (Internal) 1/3 models | general_agent | — | 64.5 | — |
| Harbor-Index 1/3 models | coding_agent | — | 42.5 | — |