NL2Repo-Bench
NL2Repo-Bench - repository generation from natural language. Evaluated with Claude Code harness. Bash commands that access the specific repository are disabled to prevent reward hacking.
Category
Agentic Coding
Max Score
—
Score Type
percent
Active
Yes
Model Scores
Scores for NL2Repo-Bench
| Model | Score | Date | Verified |
|---|---|---|---|
| GLM 5.1 Thinking |
40.59%
|
11.06.2026 | Unverified |
| GPT 5.5 |
65.48%
|
11.06.2026 | Unverified |
| Kimi K2.6 |
44.35%
|
11.06.2026 | Unverified |
| MiniMax-M2.7 |
28.03%
|
11.06.2026 | Unverified |
| Qwen3.8-2.4T-A95B |
71.76%
|
18.08.2026 | Verified |
| Claude Opus 4.8 |
100.00%
|
18.08.2026 | Unverified |
| Qwen3.7-Max |
53.56%
|
18.08.2026 | Unverified |
| DeepSeek V4 Pro |
29.08%
|
11.06.2026 | Unverified |
| Claude Opus 4.7 |
72.59%
|
11.06.2026 | Unverified |
| MiniMax-M3 |
42.89%
|
11.06.2026 | Unverified |
| Gemini 3.1 Pro |
21.60
|
11.06.2026 | Unverified |
| Qwen3.8-Flash-Next |
55.44%
|
26.08.2026 | Verified |
| Qwen3.8-27B |
43.31%
|
26.08.2026 | Unverified |
| Qwen3.7-Plus |
40.79%
|
26.08.2026 | Unverified |
| DeepSeek-V4-Flash-0731 |
68.20%
|
26.08.2026 | Unverified |
| Claude-Opus-4.6 (Max) |
54.39%
|
26.08.2026 | Unverified |
| DeepSeek-V4.1-Flash |
88.70%
|
10.09.2026 | Verified |