all the models — AI benchmark observatory
← Benchmarks

Internal Research Debugging Evaluation

The Internal Research Debugging Evaluation measures whether models can debug 41 real bugs from internal OpenAI research experiments (plus alignment-auditing tasks), where the original solutions took experienced researchers hours to days. Passing corresponds to providing assistance that would unblock the user, including partial root-cause explanations or fixes.

id internal-research-debugging-evaluation · max 1 · 3 models reported

#ModelScore

No scores for this benchmark yet.

Internal Research Debugging Evaluation Leaderboard · all the models