← Benchmarks
Internal Research Debugging Evaluation
The Internal Research Debugging Evaluation measures whether models can debug 41 real bugs from internal OpenAI research experiments (plus alignment-auditing tasks), where the original solutions took experienced researchers hours to days. Passing corresponds to providing assistance that would unblock the user, including partial root-cause explanations or fixes.
id internal-research-debugging-evaluation · max 1 · 3 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
