← Models
o1
OpenAI · proprietary · o1-2024-12-17
A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.
Benchmark scores
| Benchmark | Score |
|---|---|
| GSM8k | 0.97 |
| MATH | 0.96 |
| GPQA Physics | 0.93 |
| MMLU | 0.92 |
| MGSM | 0.89 |
| HumanEval | 0.88 |
| GPQA | 0.78 |
| MMMU | 0.78 |
| AIME 2024 | 0.74 |
| SWE-Bench Verified | 0.41 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
