all the models — AI benchmark observatory
← Models

o1

OpenAI · proprietary · o1-2024-12-17

A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.

GSM8kMATHGPQA PhysicsMMLUMGSMHumanEval

Benchmark scores

BenchmarkScore
GSM8k0.97
MATH0.96
GPQA Physics0.93
MMLU0.92
MGSM0.89
HumanEval0.88
GPQA0.78
MMMU0.78
AIME 20240.74
SWE-Bench Verified0.41

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.