all the models — AI benchmark observatory
← Benchmarks

ArXivMath (with tools)

ArXivMath final-answer accuracy with a code-execution sandbox (no internet), distinct from the without-tools setting.

id arxivmath-with-tools · max 1 · 1 models reported

#ModelScore
1Claude Sonnet 5.5
Anthropic
0.95
ArXivMath (with tools) Leaderboard · all the models