← Benchmarks
ARC-C
The AI2 Reasoning Challenge (ARC) Challenge Set is a multiple-choice question-answering benchmark containing grade-school level science questions that require advanced reasoning capabilities. ARC-C specifically contains questions that were answered incorrectly by both retrieval-based and word co-occurrence algorithms, making it a particularly challenging subset designed to test commonsense reasoning abilities in AI systems.
id arc-c · max 1 · 35 models reported
| # | Model | Score |
|---|---|---|
| 1 | MiMo-V2.5-Pro Xiaomi · open | 0.97 |
| 2 | Llama 3.1 405B Instruct Meta · open | 0.97 |
| 3 | Claude 3 Opus Anthropic | 0.96 |
| 4 | Llama 3.1 70B Instruct Meta · open | 0.95 |
| 5 | Nova Pro Amazon | 0.95 |
| 6 | Claude 3 Sonnet Anthropic | 0.93 |
| 7 | Jamba 1.5 Large AI21 Labs · open | 0.93 |
| 8 | Nova Lite Amazon | 0.92 |
