Skip to main content

Benchmark profile

Abstraction and Reasoning Corpus for AGI v2 (ARC-AGI-2)

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

Data verified

GPT-5.6 Sol leads the ARC-AGI-2 leaderboard on BenchLM's July 2026 update with 92.5%, ahead of Claude Opus 5 (90.4%) and GPT-5.5 (85%), across 19 tracked models.

Top models on ARC-AGI-2 — July 29, 2026

As of July 29, 2026, GPT-5.6 Sol leads the ARC-AGI-2 leaderboard with 92.5% , followed by Claude Opus 5 (90.4%) and GPT-5.5 (85%).

19 modelsReasoning31% of category scoreCurrentUpdated July 29, 2026

Leaderboard (19 models)

Score
1
GPT-5.6 SolOpenAI · Closed
92.5%
2
Claude Opus 5Anthropic · Closed
90.4%
3
GPT-5.5OpenAI · Closed
85%
4
GPT-5.6 TerraOpenAI · Closed
83.9%
5
GPT-5.4 ProOpenAI · Closed
83.3%
6
Gemini 3.1 ProGoogle · Closed
77.1%
7
Claude Opus 4.7 (Adaptive)Anthropic · Closed
75.8%
8
GPT-5.4OpenAI · Closed
74.0%
9
Gemini 3.5 FlashGoogle · Closed
72.1%
10
Claude Opus 4.8Anthropic · Closed
72.1%
11
GPT-5.6 LunaOpenAI · Closed
59.5%
12
GPT-5.2 ProOpenAI · Closed
54.2%
13
Grok 4.20xAI · Closed
53.3%
14
GPT-5.2OpenAI · Closed
52.9%
15
Grok 4.5xAI · Closed
52.6%
16
Gemini 3 Pro Deep ThinkGoogle · Closed
45.1%
17
Muse SparkMeta · Closed
42.5%
18
Gemini 3 ProGoogle · Closed
31.1%
19
Claude Sonnet 4.5Anthropic · Closed
13.6%

According to BenchLM.ai, GPT-5.6 Sol leads the ARC-AGI-2 benchmark with a score of 92.5%, followed by Claude Opus 5 (90.4%) and GPT-5.5 (85%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

19 models have been evaluated on ARC-AGI-2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, ARC-AGI-2 contributes 31% of the category score, so strong performance here directly affects a model's overall ranking.

About ARC-AGI-2

Year

2025

Tasks

Visual pattern completion and abstract reasoning

Format

Grid transformation puzzles with novel rules

Difficulty

Expert-level — hardest public reasoning benchmark

ARC-AGI-2 extends the original ARC benchmark with harder puzzles designed to test genuine fluid intelligence. Four major AI labs (Anthropic, Google, OpenAI, xAI) now report their model performance on this benchmark. Average individual human performance is 66%, the human panel completion rate is 100%, and the grand prize threshold is greater than 85%. Top frontier models reach 75-85 in BenchLM's tracked data, making it one of the few benchmarks that still separates current reasoning systems.

BenchLM freshness & provenance

Version

ARC-AGI 2

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ARC-AGI-2 measure?

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

Which model scores highest on ARC-AGI-2?

GPT-5.6 Sol by OpenAI currently leads with a score of 92.5% on ARC-AGI-2.

How many models are evaluated on ARC-AGI-2?

19 AI models have been evaluated on ARC-AGI-2 on BenchLM.

Last updated: July 29, 2026 · BenchLM version ARC-AGI 2

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.