all the models — AI benchmark observatory
← Benchmarks

BIG-Bench Hard

BIG-Bench Hard (BBH) is a subset of 23 challenging BIG-Bench tasks selected because prior language model evaluations did not outperform average human-rater performance. The benchmark contains 6,511 evaluation examples testing various forms of multi-step reasoning including arithmetic, logical reasoning (Boolean expressions, logical deduction), geometric reasoning, temporal reasoning, and language understanding. Tasks require capabilities such as causal judgment, object counting, navigation, pattern recognition, and complex problem solving.

id big-bench-hard · max 1 · 21 models reported

#ModelScore
1Claude 3.5 Sonnet
Anthropic
0.93
2Claude 3.5 Sonnet
Anthropic
0.93
3Gemini 1.5 Pro
Google
0.89
BIG-Bench Hard Leaderboard · all the models