all the models — AI benchmark observatory
← Benchmarks

MathVision

MATH-Vision is a dataset designed to measure multimodal mathematical reasoning capabilities. It focuses on evaluating how well models can solve mathematical problems that require both visual understanding and mathematical reasoning, bridging the gap between visual and mathematical domains.

id mathvision · max 1 · 37 models reported

#ModelScore
1Kimi K3
Moonshot AI · open
0.98
2Qwen3.8 Flash
Alibaba Cloud / Qwen Team
0.96
3Qwen3.8-Flash-Next
Alibaba Cloud / Qwen Team · open
0.96
4Qwen3.8-27B
Alibaba Cloud / Qwen Team · open
0.95
5Seed 2.1 Pro
ByteDance
0.94
6Kimi K2.6
Moonshot AI · open
0.93
7Seed 2.1 Turbo
ByteDance
0.93
8Qwen3.7-Plus
Alibaba Cloud / Qwen Team
0.90
9Qwen3.6 Plus
Alibaba Cloud / Qwen Team
0.88
10Qwen3.5-122B-A10B
Alibaba Cloud / Qwen Team · open
0.86
11Qwen3.5-27B
Alibaba Cloud / Qwen Team · open
0.86
12Gemma 4 31B
Google · open
0.86