← Benchmarks
Instruct HumanEval
Instruction-based variant of HumanEval benchmark for evaluating large language models' code generation capabilities with functional correctness using pass@k metric on programming problems
id instruct-humaneval · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
