all the models — AI benchmark observatory
← Benchmarks

MATH-500

MATH-500 is a subset of the MATH dataset containing 500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels across seven mathematical subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus.

id math-500 · max 1 · 33 models reported

#ModelScore
1LongCat-Flash-Thinking
Meituan · open
0.99
2Sarvam-105B
Sarvam AI · open
0.99
3GLM-4.5
Zhipu AI · open
0.98
4GLM-4.5-Air
Zhipu AI · open
0.98
5Nemotron Nano 9B v2
NVIDIA · open
0.98
6Kimi K2 Instruct
Moonshot AI · open
0.97
7Kimi K2-Instruct-0905
Moonshot AI · open
0.97
8Llama 3.1 Nemotron Ultra 253B v1
NVIDIA · open
0.97
9Sarvam-30B
Sarvam AI · open
0.97
10LongCat-Flash-Lite
Meituan · open
0.97
11MiniMax M1 80K
MiniMax · open
0.97
12Qwen3 14B
Alibaba Cloud / Qwen Team · open
0.97
13Llama-3.3 Nemotron Super 49B v1
NVIDIA · open
0.97
14LongCat-Flash-Chat
Meituan · open
0.96
15Claude 3.7 Sonnet
Anthropic
0.96
16Kimi-k1.5
Moonshot AI
0.96
17MiniMax M1 40K
MiniMax · open
0.96
18DeepSeek R1 Zero
DeepSeek · open
0.96
19Llama 3.1 Nemotron Nano 8B V1
NVIDIA · open
0.95
20Phi 4 Mini Reasoning
Microsoft · open
0.95
21DeepSeek R1 Distill Llama 70B
DeepSeek · open
0.94
22DeepSeek R1 Distill Qwen 32B
DeepSeek · open
0.94
23DeepSeek-V3 0324
DeepSeek · open
0.94
24DeepSeek R1 Distill Qwen 14B
DeepSeek · open
0.94
25DeepSeek R1 Distill Qwen 7B
DeepSeek · open
0.93
26QwQ-32B
Alibaba Cloud / Qwen Team · open
0.91
27QwQ-32B-Preview
Alibaba Cloud / Qwen Team · open
0.91
28DeepSeek-V3
DeepSeek · open
0.90
29o1-mini
OpenAI
0.90
30DeepSeek R1 Distill Llama 8B
DeepSeek · open
0.89