all the models — AI benchmark observatory
← Benchmarks

Program Bench

Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200 tasks from small CLI tools to large systems such as FFmpeg and SQLite, with submissions judged against more than 248,000 fuzz-generated behavioral tests.

id program-bench · max 1 · 11 models reported

#ModelScore
1Claude Opus 5.5
Anthropic
0.91
2Kimi K3
Moonshot AI · open
0.78
Program Bench Leaderboard · all the models