← Benchmarks
CharXiv-R
CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.
id charxiv-r · max 1 · 58 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude Mythos Preview Anthropic | 0.93 |
| 2 | Kimi K3 Moonshot AI · open | 0.91 |
| 3 | Claude Opus 4.7 Anthropic | 0.91 |
| 4 | Qwen3.8 Flash Alibaba Cloud / Qwen Team | 0.91 |
| 5 | Qwen3.8-Flash-Next Alibaba Cloud / Qwen Team · open | 0.91 |
| 6 | Qwen3.8-27B Alibaba Cloud / Qwen Team · open | 0.90 |
| 7 | Claude Opus 4.8 Anthropic | 0.90 |
| 8 | Gemini 3.6 Flash Google | 0.89 |
| 9 | GLM-5.3-Flash Zhipu AI · open | 0.89 |
| 10 | Gemini 3.7 Flash Google | 0.89 |
| 11 | Muse Spark 1.1 Meta | 0.88 |
| 12 | Claude Sonnet 5 Anthropic | 0.88 |
| 13 | Kimi K2.6 Moonshot AI · open | 0.87 |
| 14 | Muse Spark Meta | 0.86 |
| 15 | Seed 2.1 Pro ByteDance | 0.86 |
| 16 | Gemini 3.8 Flash Google | 0.86 |
| 17 | Qwen3.7-Plus Alibaba Cloud / Qwen Team | 0.86 |
