all the models — AI benchmark observatory
← Benchmarks

PolyMATH

Polymath is a challenging multi-modal mathematical reasoning benchmark designed to evaluate the general cognitive reasoning abilities of Multi-modal Large Language Models (MLLMs). The benchmark comprises 5,000 manually collected high-quality images of cognitive textual and visual challenges across 10 distinct categories, including pattern recognition, spatial reasoning, and relative reasoning.

id polymath · max 1 · 23 models reported

#ModelScore
1Qwen3.7 Max
Alibaba Cloud / Qwen Team
0.86
PolyMATH Leaderboard · all the models