all the models — AI benchmark observatory
← Benchmarks

HumanEval-Average

A variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

id humaneval-average · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.