all the models — AI benchmark observatory
← Benchmarks

HumanEval+

Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional correctness, detecting previously undetected wrong code

id humaneval+ · max 1 · 10 models reported

#ModelScore
1Phi 4 Reasoning
Microsoft · open
0.93
2Phi 4 Reasoning Plus
Microsoft · open
0.92