← Benchmarks
HumanEval Plus
Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code functional correctness, detecting previously undetected wrong code
id humaneval-plus · max 1 · 1 models reported
| # | Model | Score |
|---|---|---|
| 1 | Mistral Small 3.2 24B Instruct Mistral AI · open | 0.93 |
