← Benchmarks
DocVQA
A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.
id docvqa · max 1 · 28 models reported
| # | Model | Score |
|---|---|---|
| 1 | Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team · open | 0.96 |
| 2 | Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team · open | 0.96 |
| 3 | Claude 3.5 Sonnet Anthropic | 0.95 |
| 4 | Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team · open | 0.95 |
| 5 | Mistral Small 3.2 24B Instruct Mistral AI · open | 0.95 |
| 6 | Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team · open | 0.95 |
| 7 | Llama 4 Maverick Meta · open | 0.94 |
| 8 | Llama 4 Scout Meta · open | 0.94 |
| 9 | Grok-2 xAI | 0.94 |
| 10 | Nova Pro Amazon | 0.94 |
| 11 | DeepSeek VL2 DeepSeek · open | 0.93 |
| 12 | Pixtral Large Mistral AI · open | 0.93 |
| 13 | Grok-2 mini xAI | 0.93 |
| 14 | Phi-4-multimodal-instruct Microsoft · open | 0.93 |
| 15 | GPT-4o OpenAI | 0.93 |
| 16 | Nova Lite Amazon | 0.92 |
| 17 | DeepSeek VL2 Small DeepSeek · open | 0.92 |
| 18 | North Micro Vision Instruct Cohere · open | 0.92 |
| 19 | LFM2.5-VL-3B Liquid AI · open | 0.91 |
| 20 | Pixtral-12B Mistral AI · open | 0.91 |
| 21 | Llama 3.2 90B Instruct Meta · open | 0.90 |
| 22 | DeepSeek VL2 Tiny DeepSeek · open | 0.89 |
| 23 | Llama 3.2 11B Instruct Meta · open | 0.88 |
| 24 | Gemma 3 12B Google · open | 0.87 |
| 25 | Gemma 3 27B Google · open | 0.87 |
| 26 | Grok-1.5 xAI | 0.86 |
| 27 | Grok-1.5V xAI | 0.86 |
