all the models — AI benchmark observatory
← Benchmarks

MathVista

MathVista evaluates mathematical reasoning of foundation models in visual contexts. It consists of 6,141 examples derived from 28 existing multimodal datasets and 3 newly created datasets (IQTest, FunctionQA, and PaperQA), combining challenges from diverse mathematical and visual tasks to assess models' ability to understand complex figures and perform rigorous reasoning.

id mathvista · max 1 · 40 models reported

#ModelScore
1Seed 2.1 Pro
ByteDance
0.91
2Seed 2.1 Turbo
ByteDance
0.91
3Seed 1.8
ByteDance
0.88
4o3
OpenAI
0.87