Terminal-Bench 2.1 (Claude Code)
Terminal-Bench 2.1 evaluated using Claude Code 2.1.126, parser=json, temp=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.
Category
Agentic Coding
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for Terminal-Bench 2.1 (Claude Code)
| Model | Score | Date | Verified |
|---|---|---|---|
| Ornith-1.0-9B |
43.75%
|
25.06.2026 | Verified |
| Qwen3.5-9B |
18.90
|
25.06.2026 | Unverified |
| Qwen3.5-35B-A3B |
40.32%
|
25.06.2026 | Unverified |
| Ornith-1.0-35B |
88.51%
|
25.06.2026 | Verified |
| Qwen3.6-35B-A3B |
61.09%
|
25.06.2026 | Unverified |
| Qwen3.5-397B-A17B |
59.88%
|
25.06.2026 | Unverified |
| Ornith-1.5-35B-A3B |
100.00%
|
19.08.2026 | Verified |
| Ornith-1.0-35B-A3B |
88.51%
|
19.08.2026 | Unverified |