# Abstraction and Reasoning Corpus for AGI v2 (ARC-AGI-2)

> A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

Canonical page: https://benchlm.ai/benchmarks/arc-agi-2

- Category: [Reasoning](/reasoning)
- Last updated: September 10, 2026

## About ARC-AGI-2

- Year: 2025
- Tasks: Visual pattern completion and abstract reasoning
- Format: Grid transformation puzzles with novel rules
- Difficulty: Expert-level — hardest public reasoning benchmark
- Paper: [ARC-AGI-2: A Harder General Intelligence Benchmark](https://arcprize.org/arc-agi/2/)

ARC-AGI-2 extends the original ARC benchmark with harder puzzles designed to test genuine fluid intelligence. Four major AI labs (Anthropic, Google, OpenAI, xAI) now report their model performance on this benchmark. Average individual human performance is 66%, the human panel completion rate is 100%, and the grand prize threshold is greater than 85%. Top frontier models reach 75-85 in BenchLM's tracked data, making it one of the few benchmarks that still separates current reasoning systems.

ARC-AGI-2 is currently weighted in BenchLM's scoring formula. The Reasoning category carries 17% of the overall score, and ARC-AGI-2 contributes 25% of that category score.

## Leaderboard (22 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 95% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 92.5% |
| 3 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 90.4% |
| 4 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 90% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 85% |
| 6 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 83.9% |
| 7 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 83.3% |
| 8 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 81.4% |
| 9 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 77.1% |
| 10 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 75.8% |
| 11 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 74.0% |
| 12 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 72.1% |
| 13 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 72.1% |
| 14 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 59.5% |
| 15 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 53.3% |
| 16 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 52.9% |
| 17 | [Grok 4.5](/models/grok-4-5) | xAI | 52.6% |
| 18 | [Gemini 3 Pro Deep Think](/models/gemini-3-pro-deep-think) | Google | 45.1% |
| 19 | [Muse Spark](/models/muse-spark) | Meta | 42.5% |
| 20 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 40.1% |
| 21 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 31.1% |
| 22 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 13.6% |

## FAQ

### What does ARC-AGI-2 measure?

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

### Which model scores highest on ARC-AGI-2?

GPT-6 Astra by OpenAI currently leads with a score of 95% on ARC-AGI-2.

### How many models are evaluated on ARC-AGI-2?

22 AI models have been evaluated on ARC-AGI-2 on BenchLM.

### Does ARC-AGI-2 affect BenchLM's overall score?

Yes. ARC-AGI-2 is a weighted benchmark inside the Reasoning category, which carries 17% of BenchLM's overall score. ARC-AGI-2 itself contributes 25% of that category score.

## Compare Top Models on ARC-AGI-2

- [GPT-6 Astra vs GPT-5.6 Sol](/compare/gpt-5-6-sol-vs-gpt-6-astra)
- [GPT-5.6 Sol vs Claude Opus 5](/compare/claude-opus-5-vs-gpt-5-6-sol)
- [Claude Opus 5 vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-claude-opus-5)
- [Claude Fable 5.1 vs GPT-5.5](/compare/claude-fable-5-1-vs-gpt-5-5)
