← Benchmarks
HumanEval+
Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional correctness, detecting previously undetected wrong code
id humaneval+ · max 1 · 10 models reported
| # | Model | Score |
|---|---|---|
| 1 | Phi 4 Reasoning Microsoft · open | 0.93 |
| 2 | Phi 4 Reasoning Plus Microsoft · open | 0.92 |
