all the models — AI benchmark observatory
← Benchmarks

EvalPlus

A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to better assess functional correctness of generated code, including HumanEval+ with 80x more test cases

id evalplus · max 100 · 4 models reported

#ModelScore
1Kimi K2 Base
Moonshot AI · open
0.80
2Qwen2 72B Instruct
Alibaba Cloud / Qwen Team · open
0.79
3Qwen3 235B A22B
Alibaba Cloud / Qwen Team · open
0.78