← Benchmarks
MMLU-redux-2.0
A curated version of the MMLU benchmark featuring manually re-annotated 5,700 questions across 57 subjects to identify and correct errors in the original dataset. Addresses the 6.49% error rate found in MMLU and provides more reliable evaluation metrics for language models.
id mmlu-redux-2.0 · max 1 · 1 models reported
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2 Base Moonshot AI · open | 0.90 |
