ExploitBench
Exploitation benchmark: average coverage score over 41 tasks across 3 revisions (union of capabilities across revisions per task), max 300 interaction rounds, Claude Code 2.1.207 with max reasoning effort, no web tools, domain whitelist. Reported in the GLM-5.3 model card.
Category
Cybersecurity
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for ExploitBench
| Model | Score | Date | Verified |
|---|---|---|---|
| GLM-5.2 |
24.40
|
28.08.2026 | Verified |
| Kimi K3 |
14.55%
|
28.08.2026 | Verified |
| Qwen3.8-2.4T-A95B |
8.21%
|
28.08.2026 | Verified |
| Opus-4.8 |
29.10%
|
28.08.2026 | Verified |
| Fable-5 |
100.00%
|
28.08.2026 | Verified |
| GPT-5.6 Sol |
97.20%
|
28.08.2026 | Verified |
| GLM-5.3 |
55.97%
|
28.08.2026 | Verified |