← Benchmarks
CharXiv-D
CharXiv-D is the descriptive questions subset of the CharXiv benchmark, designed to assess multimodal large language models' ability to extract basic information from scientific charts. It contains descriptive questions covering information extraction, enumeration, pattern recognition, and counting across 2,323 diverse charts from arXiv papers, all curated and verified by human experts.
id charxiv-d · max 1 · 18 models reported
| # | Model | Score |
|---|---|---|
| 1 | Seed 2.1 Pro ByteDance | 0.95 |
| 2 | Seed 2.1 Turbo ByteDance | 0.95 |
| 3 | Seed 2.0 Mini ByteDance | 0.92 |
| 4 | Qwen3 VL 32B Instruct Alibaba Cloud / Qwen Team · open | 0.91 |
| 5 | Qwen3 VL 32B Thinking Alibaba Cloud / Qwen Team · open | 0.90 |
| 6 | GPT-4.5 OpenAI | 0.90 |
| 7 | GPT-4.1 mini OpenAI | 0.88 |
| 8 | Command A+ Cohere · open | 0.88 |
| 9 | GPT-4.1 OpenAI | 0.88 |
| 10 | Qwen3 VL 30B A3B Thinking Alibaba Cloud / Qwen Team · open | 0.87 |
| 11 | Qwen3 VL 8B Thinking Alibaba Cloud / Qwen Team · open | 0.86 |
| 12 | Qwen3 VL 30B A3B Instruct Alibaba Cloud / Qwen Team · open | 0.85 |
| 13 | GPT-4o OpenAI | 0.85 |
