← Benchmarks
ARC-AGI
The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.
id arc-agi · max 1 · 11 models reported
| # | Model | Score |
|---|---|---|
| 1 | GPT-6 Astra OpenAI | 0.98 |
| 2 | GPT-5.5 OpenAI | 0.95 |
| 3 | GPT-5.4 OpenAI | 0.94 |
| 4 | GPT-5.2 Pro OpenAI | 0.91 |
| 5 | o3 OpenAI | 0.88 |
| 6 | GPT-5.2 OpenAI | 0.86 |
