← Benchmarks
FlenQA
Flexible Length Question Answering dataset for evaluating the impact of input length on reasoning performance of language models, featuring True/False questions embedded in contexts of varying lengths (250-3000 tokens) across three reasoning tasks: Monotone Relations, People In Rooms, and simplified Ruletaker
id flenqa · max 1 · 2 models reported
| # | Model | Score |
|---|---|---|
| 1 | Phi 4 Reasoning Plus Microsoft · open | 0.98 |
| 2 | Phi 4 Reasoning Microsoft · open | 0.98 |
