Program Bench

Evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and its documentation. It spans 200 tasks, from small CLI tools to large systems like FFmpeg and SQLite. Submissions are judged against over 248,000 fuzz-generated behavioral tests. In each task, the agent is given an executable and its documentation, but no source code, decompilation, or internet access. It must choose its own implementation language, build the full program from scratch, and pass a behavioral test suite comparing its output against the original binary.

Category

Agentic Coding

Max Score

100.0

Score Type

percent

Active

Yes

Model Scores

Scores for Program Bench

Model Score Date Verified
Kimi K2.6
48.30
23.08.2026 Verified
Kimi K2.7 Code
25.48%
23.08.2026 Verified
GPT-5.5
100.00%
23.08.2026 Verified
Claude Opus 4.8
74.52%
23.08.2026 Verified