MLS-Bench-Lite
MLS-Bench-Lite - agentic coding benchmark. Evaluated with Claude Code using a 5-hour timeout and max_tokens=131,072. Scores from the official leaderboard.
Category
Agentic Coding
Max Score
—
Score Type
percent
Active
Yes
Model Scores
Scores for MLS-Bench-Lite
| Model | Score | Date | Verified |
|---|---|---|---|
| Qwen3.8-2.4T-A95B |
61.64%
|
18.08.2026 | Verified |
| Fable-5 |
100.00%
|
18.08.2026 | Unverified |
| GPT-5.6 Sol |
84.05%
|
18.08.2026 | Unverified |
| Claude Opus 4.8 |
69.40%
|
18.08.2026 | Unverified |
| Qwen3.7-Max |
21.55%
|
18.08.2026 | Unverified |
| Kimi K3 |
93.10%
|
18.08.2026 | Verified |
| Kimi K2.6 |
26.70
|
23.08.2026 | Verified |
| Kimi K2.7 Code |
36.21%
|
23.08.2026 | Verified |
| GPT-5.5 |
37.93%
|
23.08.2026 | Verified |