all the models — AI benchmark observatory
← Benchmarks

CharXiv-R

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

id charxiv-r · max 1 · 58 models reported

#ModelScore
1Claude Mythos Preview
Anthropic
0.93
2Kimi K3
Moonshot AI · open
0.91
3Claude Opus 4.7
Anthropic
0.91
4Qwen3.8 Flash
Alibaba Cloud / Qwen Team
0.91
5Qwen3.8-Flash-Next
Alibaba Cloud / Qwen Team · open
0.91
6Qwen3.8-27B
Alibaba Cloud / Qwen Team · open
0.90
7Claude Opus 4.8
Anthropic
0.90
8Gemini 3.6 Flash
Google
0.89
9GLM-5.3-Flash
Zhipu AI · open
0.89
10Gemini 3.7 Flash
Google
0.89
11Muse Spark 1.1
Meta
0.88
12Claude Sonnet 5
Anthropic
0.88
13Kimi K2.6
Moonshot AI · open
0.87
14Muse Spark
Meta
0.86
15Seed 2.1 Pro
ByteDance
0.86
16Gemini 3.8 Flash
Google
0.86
17Qwen3.7-Plus
Alibaba Cloud / Qwen Team
0.86