IFEval
IFEval - Instruction Following Evaluation benchmark for measuring LLM instruction-following capabilities.
Category
Instruction Following
Max Score
—
Score Type
percent
Active
Yes
Model Scores
Scores for IFEval
| Model | Score | Date | Verified |
|---|---|---|---|
| Qwen3.5-4B |
91.36%
|
01.02.2026 | Verified |
| Qwen3.5-122B-A10B |
97.34%
|
01.02.2026 | Verified |
| Qwen3.5-35B-A3B |
94.85%
|
01.02.2026 | Unverified |
| Qwen3.5-27B |
100.00%
|
01.02.2026 | Unverified |
| Qwen3-235B-A22B |
88.04%
|
01.02.2026 | Unverified |
| GPT-5-mini 2025-08-07 |
98.17%
|
01.02.2026 | Unverified |
| Gemini-2.5-Flash-Thinking |
91.36%
|
31.07.2025 | Verified |
| Qwen3-30B-A3B |
85.88%
|
29.04.2025 | Verified |
| Qwen3-30B-A3B-Thinking-2507 |
89.87%
|
31.07.2025 | Verified |
| Spark-X2.5-4B |
96.68%
|
— | Verified |
| Spark-X2.5-1.7B |
90.86%
|
— | Verified |
| Qwen3.5-2B |
72.76%
|
— | Verified |
| Gemma4-E4B |
17.44%
|
— | Verified |
| MiniCPM5-2B |
86.21%
|
— | Verified |
| LFM2.5-2.6B |
97.34%
|
— | Verified |
| granite-4.2-3B |
97.84%
|
— | Verified |
| LFM2.5-8B-A1B |
93.02%
|
— | Verified |
| Qwen3-Next-80B-A3B-Thinking |
89.87%
|
10.09.2025 | Verified |
| Qwen3-VL-30B-A3B |
84.72%
|
04.10.2025 | Verified |
| Qwen3.5-9B |
94.19%
|
— | Verified |
| Gemma4-12B |
99.67%
|
— | Verified |
| Gemma4-E2B |
34.80
|
— | Verified |
| Nemotron-3-Nano-4B |
88.37%
|
— | Verified |