Abstraction and Reasoning Corpus for AGI v2 (ARC-AGI-2)
A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.
Successor track
ARC-AGI-3 evaluates interactive agentic reasoning rather than ARC-AGI-2's static grid puzzles. The two score tables are not directly comparable.
View ARC-AGI-3GPT-6 Astra leads the ARC-AGI-2 leaderboard on BenchLM's September 2026 update with 95%, ahead of GPT-5.6 Sol (92.5%) and Claude Opus 5 (90.4%), across 23 models.
Top models on ARC-AGI-2 — September 10, 2026
As of September 10, 2026, GPT-6 Astra leads the ARC-AGI-2 leaderboard with 95% , followed by GPT-5.6 Sol (92.5%) and Claude Opus 5 (90.4%).
GPT-6 Astra
OpenAI
GPT-5.6 Sol
OpenAI
Claude Opus 5
Anthropic
23 modelsReasoning25% of category scoreCurrentUpdated September 10, 2026
Leaderboard (23 models)
ScoreAccording to BenchLM.ai, GPT-6 Astra leads the ARC-AGI-2 benchmark with a score of 95%, followed by GPT-5.6 Sol (92.5%) and Claude Opus 5 (90.4%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.
23 models have been evaluated on ARC-AGI-2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, ARC-AGI-2 contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.
About ARC-AGI-2
Year
2025
Tasks
Visual pattern completion and abstract reasoning
Format
Grid transformation puzzles with novel rules
Difficulty
Expert-level — hardest public reasoning benchmark
ARC-AGI-2 extends the original ARC benchmark with harder puzzles designed to test genuine fluid intelligence. Four major AI labs (Anthropic, Google, OpenAI, xAI) now report their model performance on this benchmark. Average individual human performance is 66%, the human panel completion rate is 100%, and the grand prize threshold is greater than 85%. Top frontier models reach 75-85 in BenchLM's tracked data, making it one of the few benchmarks that still separates current reasoning systems.
BenchLM freshness & provenance
Version
ARC-AGI 2
Refresh cadence
Static
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does ARC-AGI-2 measure?
A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.
Which model scores highest on ARC-AGI-2?
GPT-6 Astra by OpenAI currently leads with a score of 95% on ARC-AGI-2.
How many models are evaluated on ARC-AGI-2?
23 AI models have been evaluated on ARC-AGI-2 on BenchLM.
Compare Top Models on ARC-AGI-2
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.