ExploitGym (6h)
ExploitGym exploitation benchmark, single-run Pass@1 on 869 tasks under a 6-hour timeout budget (API inference time rescaled by per-model TPS from Artificial Analysis plus non-API overhead), evaluated in Claude Code 2.1.207 with max reasoning effort, no web tools, and a domain whitelist. Reported in the GLM-5.3 model card.
Category
Cybersecurity
Max Score
869.0
Score Type
count
Active
Yes
Model Scores
Scores for ExploitGym (6h)
| Model | Score | Date | Verified |
|---|---|---|---|
| GLM-5.3 |
38.95%
|
28.08.2026 | Verified |
| GLM-5.2 |
4.87%
|
28.08.2026 | Verified |
| Kimi K3 |
16.48%
|
28.08.2026 | Verified |
| Qwen3.8-2.4T-A95B |
26.00
|
28.08.2026 | Verified |
| Opus-4.8 |
35.21%
|
28.08.2026 | Verified |
| Fable-5 |
82.77%
|
28.08.2026 | Verified |
| GPT-5.6 Sol |
100.00%
|
28.08.2026 | Verified |