NoLiMa

NoLiMa - long context benchmark evaluating retrieval and reasoning without explicit long-context matching.

Category

Long Context

Max Score

100.0

Score Type

percent

Active

Yes

Model Scores

Scores for NoLiMa

Model Score Date Verified
LFM2.5-2.6B
0.30%
— Verified
Qwen3.5-2B
24.56%
— Verified
Qwen3.5-4B
63.61%
— Verified
granite-4.2-3B
6.80%
— Verified
Nemotron-3-Nano-4B
0.89%
— Verified
Gemma4-E4B
2.66%
— Verified
LFM2.5-8B-A1B
0.50
— Verified
MiniCPM5-2B
100.00%
— Verified
Gemma4-E2B
5.03%
— Verified