all the models — AI benchmark observatory
← Benchmarks

BoolQ

BoolQ is a reading comprehension dataset for yes/no questions containing 15,942 naturally occurring examples. Each example consists of a question, passage, and boolean answer, where questions are generated in unprompted and unconstrained settings. The dataset challenges models with complex, non-factoid information requiring entailment-like inference to solve.

id boolq · max 1 · 11 models reported

#ModelScore

No scores for this benchmark yet.