Program Bench
Evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and its documentation. It spans 200 tasks, from small CLI tools to large systems like FFmpeg and SQLite. Submissions are judged against over 248,000 fuzz-generated behavioral tests. In each task, the agent is given an executable and its documentation, but no source code, decompilation, or internet access. It must choose its own implementation language, build the full program from scratch, and pass a behavioral test suite comparing its output against the original binary.
Category
Agentic Coding
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for Program Bench
| Model | Score | Date | Verified |
|---|---|---|---|
| Kimi K2.6 |
48.30
|
23.08.2026 | Verified |
| Kimi K2.7 Code |
25.48%
|
23.08.2026 | Verified |
| GPT-5.5 |
100.00%
|
23.08.2026 | Verified |
| Claude Opus 4.8 |
74.52%
|
23.08.2026 | Verified |