Grok-4 Heavy
xAI · proprietary · grok-4-heavy
Grok 4 Heavy is the multi-agent version of Grok 4, released alongside the standard model in summer 2025. This system spawns multiple Grok 4 agents in parallel that work independently on problems and then collaborate by comparing their solutions, similar to a study group. The agents share insights and tricks they discover, with the system intelligently combining their work rather than simply using majority voting. Grok 4 Heavy uses approximately 10x more test-time compute than regular Grok 4, enabling it to solve significantly more complex problems. On the Humanities Last Exam, it achieves over 50% accuracy on text-only problems, and it scored a perfect result on the AIME 2025 mathematics competition. The system represents a major advancement in multi-agent AI collaboration and reasoning capabilities.
Benchmark scores
| Benchmark | Score |
|---|---|
| AIME 2025 | 1.00 |
| HMMT25 | 0.97 |
| GPQA | 0.88 |
| LiveCodeBench | 0.79 |
| Humanity's Last Exam | 0.51 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
