FrontierSWE
FrontierSWE - agentic coding benchmark. MEAN@5 scores from the official FrontierSWE leaderboard. Dominance scores recomputed from raw scores using the official evaluation script.
Category
Agentic Coding
Max Score
—
Score Type
percent
Active
Yes
Model Scores
Scores for FrontierSWE
| Model | Score | Date | Verified |
|---|---|---|---|
| Qwen3.8-2.4T-A95B |
75.17%
|
18.08.2026 | Verified |
| Claude Opus 4.8 |
69.26%
|
18.08.2026 | Unverified |
| Qwen3.7-Max |
19.76%
|
18.08.2026 | Unverified |
| Kimi K3 |
88.18%
|
18.08.2026 | Verified |
| DeepSeek-V4-Pro |
29.00
|
17.06.2026 | Unverified |
| GLM-5.1 |
2.53%
|
17.06.2026 | Unverified |
| Opus-4.8 |
63.34%
|
28.08.2026 | Verified |
| Fable-5 |
100.00%
|
28.08.2026 | Verified |
| GLM-5.2 |
76.69%
|
17.06.2026 | Verified |
| Gemini 3.1 Pro |
17.91%
|
17.06.2026 | Unverified |
| GPT-5.5 |
73.65%
|
17.06.2026 | Unverified |
| GLM-5.3 |
82.94%
|
28.08.2026 | Verified |