all the models — AI benchmark observatory
← Benchmarks

DocVQA

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.

id docvqa · max 1 · 28 models reported

#ModelScore
1Qwen2.5 VL 72B Instruct
Alibaba Cloud / Qwen Team · open
0.96
2Qwen2.5 VL 7B Instruct
Alibaba Cloud / Qwen Team · open
0.96
3Claude 3.5 Sonnet
Anthropic
0.95
4Qwen2.5-Omni-7B
Alibaba Cloud / Qwen Team · open
0.95
5Mistral Small 3.2 24B Instruct
Mistral AI · open
0.95
6Qwen2.5 VL 32B Instruct
Alibaba Cloud / Qwen Team · open
0.95
7Llama 4 Maverick
Meta · open
0.94
8Llama 4 Scout
Meta · open
0.94
9Grok-2
xAI
0.94
10Nova Pro
Amazon
0.94
11DeepSeek VL2
DeepSeek · open
0.93
12Pixtral Large
Mistral AI · open
0.93
13Grok-2 mini
xAI
0.93
14Phi-4-multimodal-instruct
Microsoft · open
0.93
15GPT-4o
OpenAI
0.93
16Nova Lite
Amazon
0.92
17DeepSeek VL2 Small
DeepSeek · open
0.92
18North Micro Vision Instruct
Cohere · open
0.92
19LFM2.5-VL-3B
Liquid AI · open
0.91
20Pixtral-12B
Mistral AI · open
0.91
21Llama 3.2 90B Instruct
Meta · open
0.90
22DeepSeek VL2 Tiny
DeepSeek · open
0.89
23Llama 3.2 11B Instruct
Meta · open
0.88
24Gemma 3 12B
Google · open
0.87
25Gemma 3 27B
Google · open
0.87
26Grok-1.5
xAI
0.86
27Grok-1.5V
xAI
0.86