← Benchmarks
AIME 2026
All 30 problems from the 2026 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.
id aime-2026 · max 1 · 26 models reported
| # | Model | Score |
|---|---|---|
| 1 | GLM-5.2 Zhipu AI · open | 0.99 |
| 2 | Inkling Thinking Machines Lab · open | 0.97 |
| 3 | Sakana Namazu Sakana AI | 0.97 |
| 4 | Kimi K2.6 Moonshot AI · open | 0.96 |
| 5 | Inkling-Small Thinking Machines Lab · open | 0.95 |
| 6 | GLM-5.1 Zhipu AI · open | 0.95 |
| 7 | Qwen3.6 Plus Alibaba Cloud / Qwen Team | 0.95 |
| 8 | Solar Pro 4 Upstage | 0.95 |
| 9 | Muse Glimmer-30B Meta · open | 0.95 |
| 10 | MAI-Thinking-1 Microsoft | 0.94 |
| 11 | Seed 2.0 Pro ByteDance | 0.94 |
| 12 | Qwen3.6-27B Alibaba Cloud / Qwen Team · open | 0.94 |
| 13 | Qwen3 Max Thinking Alibaba Cloud / Qwen Team | 0.93 |
| 14 | Ling 3.0 Flash InclusionAI | 0.93 |
| 15 | Qwen3.6-35B-A3B Alibaba Cloud / Qwen Team · open | 0.93 |
| 16 | EXAONE 4.5 33B LG AI Research | 0.93 |
| 17 | MAI-Code-1-Flash Microsoft | 0.93 |
| 18 | Qwen3.5-397B-A17B Alibaba Cloud / Qwen Team · open | 0.91 |
| 19 | Gemma 4 31B Google · open | 0.89 |
