all the models — AI benchmark observatory
← Benchmarks

OfficeQA

OfficeQA is a Databricks public benchmark that evaluates end-to-end grounded reasoning over a large corpus of historical U.S. Treasury Bulletin documents. Models must locate relevant tables across the corpus and perform precise numerical reasoning over them.

id officeqa · max 1 · 1 models reported

#ModelScore
1Claude Opus 5.5
Anthropic
0.79
OfficeQA Leaderboard · all the models