all the models — AI benchmark observatory
← Benchmarks

CorpusQA 1M

CorpusQA 1M is a long-context question answering benchmark designed to evaluate models at approximately 1 million token contexts. Models are scored on accuracy when retrieving and reasoning over information distributed across an extremely long input corpus.

id corpusqa-1m · max 1 · 3 models reported

#ModelScore

No scores for this benchmark yet.