← Benchmarks
MathVista
MathVista evaluates mathematical reasoning of foundation models in visual contexts. It consists of 6,141 examples derived from 28 existing multimodal datasets and 3 newly created datasets (IQTest, FunctionQA, and PaperQA), combining challenges from diverse mathematical and visual tasks to assess models' ability to understand complex figures and perform rigorous reasoning.
id mathvista · max 1 · 40 models reported
| # | Model | Score |
|---|---|---|
| 1 | Seed 2.1 Pro ByteDance | 0.91 |
| 2 | Seed 2.1 Turbo ByteDance | 0.91 |
| 3 | Seed 1.8 ByteDance | 0.88 |
| 4 | o3 OpenAI | 0.87 |
