TAU3-Bench
TAU3-Bench - agentic tool use benchmark. Uses official user model (gpt-5.2, low reasoning effort) + default BM25 retrieval.
Category
General Agents
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for TAU3-Bench
| Model | Score | Date | Verified |
|---|---|---|---|
| Qwen3.6-35B-A3B |
97.49%
|
16.04.2026 | Verified |
| Gemma4-31B-it |
97.93%
|
16.04.2026 | Unverified |
| Qwen3.5-35B-A3B |
100.00%
|
16.04.2026 | Unverified |
| Gemma4-26B-A4B |
85.38%
|
16.04.2026 | Unverified |
| Qwen3.5-27B |
99.26%
|
16.04.2026 | Unverified |
| Spark-X2.5-4B |
43.13%
|
— | Verified |
| Spark-X2.5-1.7B |
27.92%
|
— | Verified |
| Qwen3.5-9B |
11.96%
|
— | Verified |
| Qwen3.5-4B |
8.12%
|
— | Verified |
| Qwen3.5-2B |
4.28%
|
— | Verified |
| Gemma4-12B |
17.87%
|
— | Verified |
| Gemma4-E4B |
13.15%
|
— | Verified |
| Gemma4-E2B |
11.23%
|
— | Verified |
| MiniCPM5-2B |
28.95%
|
— | Verified |
| granite-4.2-3B |
6.50%
|
— | Verified |
| Nemotron-3-Nano-4B |
1.20
|
— | Verified |
| LFM2.5-8B-A1B |
3.25%
|
— | Verified |
| LFM2.5-2.6B |
8.86%
|
— | Verified |