← Benchmarks
EvalPlus
A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to better assess functional correctness of generated code, including HumanEval+ with 80x more test cases
id evalplus · max 100 · 4 models reported
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2 Base Moonshot AI · open | 0.80 |
| 2 | Qwen2 72B Instruct Alibaba Cloud / Qwen Team · open | 0.79 |
| 3 | Qwen3 235B A22B Alibaba Cloud / Qwen Team · open | 0.78 |
