Terminal-Bench 3.0
Terminal-Bench 3.0 agentic terminal tasks, evaluated with the Claude Code 2.1.207 harness (reasoning effort=max, 400K context, 128K max output), avg@3 over three rollouts per task scored by the task's official verifier. Open-source SOTA for GLM-5.3 per the GLM-5.3 model card.
Category
Agentic Coding
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for Terminal-Bench 3.0
| Model | Score | Date | Verified |
|---|---|---|---|
| Kimi K3 |
42.67%
|
28.08.2026 | Verified |
| GLM-5.3 |
79.00%
|
28.08.2026 | Verified |
| GLM-5.2 |
4.60
|
28.08.2026 | Verified |
| Opus-4.8 |
55.00%
|
28.08.2026 | Verified |
| Fable-5 |
97.00%
|
28.08.2026 | Verified |
| GPT-5.6 Sol |
100.00%
|
28.08.2026 | Verified |
| DeepSeek-V4.1-Flash |
84.67%
|
10.09.2026 | Verified |