all the models — AI benchmark observatory
← Benchmarks

MultiPL-E

MultiPL-E is a scalable and extensible system for translating unit test-driven code generation benchmarks to multiple programming languages. It extends HumanEval and MBPP Python benchmarks to 18 additional programming languages, enabling evaluation of neural code generation models across diverse programming paradigms and language features.

id multipl-e · max 1 · 13 models reported

#ModelScore

No scores for this benchmark yet.