all the models — AI benchmark observatory
← Benchmarks

HumanEvalFIM-Average

Average evaluation of HumanEval Fill-in-the-Middle benchmark variants (single-line, multi-line, random-span) for assessing code infilling capabilities of language models

id humanevalfim-average · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

HumanEvalFIM-Average Leaderboard · all the models