all the models — AI benchmark observatory
← Benchmarks

AIME 2026

All 30 problems from the 2026 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

id aime-2026 · max 1 · 26 models reported

#ModelScore
1GLM-5.2
Zhipu AI · open
0.99
2Inkling
Thinking Machines Lab · open
0.97
3Sakana Namazu
Sakana AI
0.97
4Kimi K2.6
Moonshot AI · open
0.96
5Inkling-Small
Thinking Machines Lab · open
0.95
6GLM-5.1
Zhipu AI · open
0.95
7Qwen3.6 Plus
Alibaba Cloud / Qwen Team
0.95
8Solar Pro 4
Upstage
0.95
9Muse Glimmer-30B
Meta · open
0.95
10MAI-Thinking-1
Microsoft
0.94
11Seed 2.0 Pro
ByteDance
0.94
12Qwen3.6-27B
Alibaba Cloud / Qwen Team · open
0.94
13Qwen3 Max Thinking
Alibaba Cloud / Qwen Team
0.93
14Ling 3.0 Flash
InclusionAI
0.93
15Qwen3.6-35B-A3B
Alibaba Cloud / Qwen Team · open
0.93
16EXAONE 4.5 33B
LG AI Research
0.93
17MAI-Code-1-Flash
Microsoft
0.93
18Qwen3.5-397B-A17B
Alibaba Cloud / Qwen Team · open
0.91
19Gemma 4 31B
Google · open
0.89
AIME 2026 Leaderboard · all the models