ExploitGym (6h)

ExploitGym exploitation benchmark, single-run Pass@1 on 869 tasks under a 6-hour timeout budget (API inference time rescaled by per-model TPS from Artificial Analysis plus non-API overhead), evaluated in Claude Code 2.1.207 with max reasoning effort, no web tools, and a domain whitelist. Reported in the GLM-5.3 model card.

Category

Cybersecurity

Max Score

869.0

Score Type

count

Active

Yes

Model Scores

Scores for ExploitGym (6h)

Model Score Date Verified
GLM-5.3
38.95%
28.08.2026 Verified
GLM-5.2
4.87%
28.08.2026 Verified
Kimi K3
16.48%
28.08.2026 Verified
Qwen3.8-2.4T-A95B
26.00
28.08.2026 Verified
Opus-4.8
35.21%
28.08.2026 Verified
Fable-5
82.77%
28.08.2026 Verified
GPT-5.6 Sol
100.00%
28.08.2026 Verified