SWE-Marathon
SWE-Marathon - long-horizon software engineering benchmark. Evaluated by Abundant AI with 1M context length, max effort level, and 128K maximum output tokens.
Category
Agentic Coding
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for SWE-Marathon
| Model | Score | Date | Verified |
|---|---|---|---|
| GLM-5.2 |
24.49%
|
17.06.2026 | Verified |
| Gemini 3.1 Pro |
6.12%
|
17.06.2026 | Unverified |
| GLM-5.1 |
1.00
|
17.06.2026 | Unverified |
| GPT-5.5 |
22.45%
|
17.06.2026 | Unverified |
| Claude Opus 4.8 |
51.02%
|
17.06.2026 | Unverified |
| GLM-5.3 |
84.69%
|
28.08.2026 | Verified |
| Kimi K3 |
96.12%
|
28.08.2026 | Verified |
| Opus-4.8 |
97.55%
|
28.08.2026 | Verified |
| Fable-5 |
65.51%
|
28.08.2026 | Verified |
| GPT-5.6 Sol |
84.69%
|
28.08.2026 | Verified |
| Hy3 |
8.16%
|
28.08.2026 | Verified |
| Hy4 preview |
63.06%
|
28.08.2026 | Verified |
| DeepSeek V4 Pro |
36.73%
|
28.08.2026 | Verified |
| Qwen3.8-2.4T-A95B |
61.22%
|
28.08.2026 | Verified |
| Claude Opus 5 |
100.00%
|
28.08.2026 | Verified |