← Benchmarks
HumanEval-Mul
A multilingual variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics
id humaneval-mul · max 1 · 2 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
