all the models — AI benchmark observatory
← Benchmarks

ArXivMath

ArXivMath is a final-answer benchmark of research-level mathematics maintained by MathArena. Problems are extracted monthly from recent arXiv paper abstracts, then filtered through automated and manual checks to ensure they are self-contained, non-trivial, and verifiable. Because problems are drawn from active research, the benchmark is more realistic and more closely connected to mathematical research than contest or olympiad benchmarks.

id arxivmath · max 1 · 4 models reported

#ModelScore
1Claude Opus 5.5
Anthropic
0.91
ArXivMath Leaderboard · all the models