← Benchmarks
MRCR 1M (pointwise)
MRCR 1M (pointwise) is a variant of the Multi-Round Coreference Resolution benchmark that uses pointwise evaluation for ultra-long contexts (~1M tokens). This version evaluates each response independently rather than comparatively, testing models' absolute performance on long-context reasoning tasks.
id mrcr-1m-(pointwise) · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
