all the models — AI benchmark observatory
← Benchmarks

ARC-C

The AI2 Reasoning Challenge (ARC) Challenge Set is a multiple-choice question-answering benchmark containing grade-school level science questions that require advanced reasoning capabilities. ARC-C specifically contains questions that were answered incorrectly by both retrieval-based and word co-occurrence algorithms, making it a particularly challenging subset designed to test commonsense reasoning abilities in AI systems.

id arc-c · max 1 · 35 models reported

#ModelScore
1MiMo-V2.5-Pro
Xiaomi · open
0.97
2Llama 3.1 405B Instruct
Meta · open
0.97
3Claude 3 Opus
Anthropic
0.96
4Llama 3.1 70B Instruct
Meta · open
0.95
5Nova Pro
Amazon
0.95
6Claude 3 Sonnet
Anthropic
0.93
7Jamba 1.5 Large
AI21 Labs · open
0.93
8Nova Lite
Amazon
0.92