ExploitBench

Exploitation benchmark: average coverage score over 41 tasks across 3 revisions (union of capabilities across revisions per task), max 300 interaction rounds, Claude Code 2.1.207 with max reasoning effort, no web tools, domain whitelist. Reported in the GLM-5.3 model card.

Category

Cybersecurity

Max Score

100.0

Score Type

percent

Active

Yes

Model Scores

Scores for ExploitBench

Model Score Date Verified
GLM-5.2
24.40
28.08.2026 Verified
Kimi K3
14.55%
28.08.2026 Verified
Qwen3.8-2.4T-A95B
8.21%
28.08.2026 Verified
Opus-4.8
29.10%
28.08.2026 Verified
Fable-5
100.00%
28.08.2026 Verified
GPT-5.6 Sol
97.20%
28.08.2026 Verified
GLM-5.3
55.97%
28.08.2026 Verified