← Benchmarks
ChartQA
ChartQA is a large-scale benchmark comprising 9.6K human-written questions and 23.1K questions generated from human-written chart summaries, designed to evaluate models' abilities in visual and logical reasoning over charts.
id chartqa · max 1 · 26 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.5 Sonnet Anthropic | 0.91 |
| 2 | Llama 4 Maverick Meta · open | 0.90 |
| 3 | Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team · open | 0.90 |
| 4 | Nova Pro Amazon | 0.89 |
| 5 | Llama 4 Scout Meta · open | 0.89 |
| 6 | Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team · open | 0.88 |
| 7 | Pixtral Large Mistral AI · open | 0.88 |
| 8 | Mistral Small 3.2 24B Instruct Mistral AI · open | 0.87 |
| 9 | Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team · open | 0.87 |
| 10 | Nova Lite Amazon | 0.87 |
| 11 | DeepSeek VL2 DeepSeek · open | 0.86 |
| 12 | GPT-4o OpenAI | 0.86 |
| 13 | Llama 3.2 90B Instruct Meta · open | 0.85 |
| 14 | Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team · open | 0.85 |
