all the models — AI benchmark observatory
← Benchmarks

BigCodeBench-Full

A comprehensive benchmark that evaluates large language models' ability to solve complex, practical programming tasks via code generation. Contains 1,140 fine-grained tasks across 7 domains using function calls from 139 libraries. Challenges LLMs to invoke multiple function calls as tools and handle complex instructions for realistic software engineering and general-purpose reasoning tasks.

id bigcodebench-full · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

BigCodeBench-Full Leaderboard · all the models