all the models — AI benchmark observatory
← Benchmarks

AIME

American Invitational Mathematics Examination (AIME) benchmark for evaluating mathematical reasoning capabilities of large language models. Contains 30 challenging mathematical problems from AIME 2024 competition that require multi-step reasoning and advanced mathematical insight. Each problem has an integer answer between 000-999.

id aime · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.

AIME Leaderboard · all the models