all the models — AI benchmark observatory
← Benchmarks

Instruct HumanEval

Instruction-based variant of HumanEval benchmark for evaluating large language models' code generation capabilities with functional correctness using pass@k metric on programming problems

id instruct-humaneval · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

Instruct HumanEval Leaderboard · all the models