all the models — AI benchmark observatory
← Benchmarks

PolyMath-en

PolyMath is a multilingual mathematical reasoning benchmark covering 18 languages and 4 difficulty levels from easy to hard, ensuring difficulty comprehensiveness, language diversity, and high-quality translation. The benchmark evaluates mathematical reasoning capabilities of large language models across diverse linguistic contexts, making it a highly discriminative multilingual mathematical benchmark.

id polymath-en · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.