all the models — AI benchmark observatory
← Benchmarks

FlenQA

Flexible Length Question Answering dataset for evaluating the impact of input length on reasoning performance of language models, featuring True/False questions embedded in contexts of varying lengths (250-3000 tokens) across three reasoning tasks: Monotone Relations, People In Rooms, and simplified Ruletaker

id flenqa · max 1 · 2 models reported

#ModelScore
1Phi 4 Reasoning Plus
Microsoft · open
0.98
2Phi 4 Reasoning
Microsoft · open
0.98