← Benchmarks
BigCodeBench-Full
A comprehensive benchmark that evaluates large language models' ability to solve complex, practical programming tasks via code generation. Contains 1,140 fine-grained tasks across 7 domains using function calls from 139 libraries. Challenges LLMs to invoke multiple function calls as tools and handle complex instructions for realistic software engineering and general-purpose reasoning tasks.
id bigcodebench-full · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
