ProgramBench Vetted tests agents rebuilding runnable binaries, with Claude Opus 5 resolving 14%
Hacker News opinions
A benchmark can be hard for the wrong reasons. What matters here is whether its score measures reverse-engineering skill rather than missing information or exploitable shortcuts.