all the models — AI benchmark observatory
← Models

Grok-3

xAI · proprietary · grok-3

Grok 3, launched by xAI on February 17, 2025, is an advanced AI model with significantly enhanced capabilities compared to Grok 2, boasting an order of magnitude increase in performance. Trained on a vast dataset that includes legal documents among others, and utilizing a massive compute infrastructure with around 200,000 GPUs in a Memphis data center, Grok 3's training used ten times more compute than its predecessor. It features specialized models like Grok 3 Reasoning and Grok 3 Mini Reasoning for complex problem-solving, and it excels in benchmarks like AIME for mathematics and GPQA for PhD-level science.

AIME 2024AIME 2025GPQALiveCodeBencMMMU

Benchmark scores

BenchmarkScore
AIME 20240.93
AIME 20250.93
GPQA0.85
LiveCodeBench0.79
MMMU0.78

Pricing

  • xAI$3.00 / $15.00

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Grok-3 Benchmarks · all the models