all the models — AI benchmark observatory
← Benchmarks

CruxEval-O

CruxEval-O is the output prediction task of the CRUXEval benchmark, designed to evaluate code reasoning, understanding, and execution capabilities. It consists of 800 Python functions (3-13 lines) where models must predict the output given a function and input. The benchmark tests fundamental code execution reasoning abilities and goes beyond simple code generation to assess deeper understanding of program behavior.

id cruxeval-o · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

CruxEval-O Leaderboard · all the models