Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Abstraction and Reasoning Corpus for AGI v3 (ARC-AGI-3)

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Data verified 24 confirmed releases in the last 30 daysStart free brief

Benchmark score on ARC-AGI-3 — August 22, 2026

We mirror the published score view for ARC-AGI-3. Claude Opus 5 leads the public snapshot at 30.2%, followed by GPT-5.6 Sol (7.8%) and Claude Opus 4.8 (1.5%). We do not use these results to rank models overall.

11 modelsReasoningCurrentDisplay onlyUpdated August 22, 2026

Benchmark score table (11 models)

Score
1
Claude Opus 5Anthropic · Closed
30.2%
2
GPT-5.6 SolOpenAI · Closed
7.8%
3
Claude Opus 4.8Anthropic · Closed
1.5%
4
GPT-5.6 TerraOpenAI · Closed
0.8%
5
GPT-5.5OpenAI · Closed
0.4%
6
Gemini 3.1 ProGoogle · Closed
0.4%
7
Grok 4.5xAI · Closed
0.3%
8
GPT-5.4OpenAI · Closed
0.2%
9
GPT-5.6 LunaOpenAI · Closed
0.2%
10
Claude Opus 4.7 (Adaptive)Anthropic · Closed
0.2%
11
Grok 4.20xAI · Closed
0.1%

The published ARC-AGI-3 snapshot places Claude Opus 5 first at 30.2%. The third row is 28.6 points behind. The broader top-10 range is 30.0 points, so the table still separates the published systems.

11 models have been evaluated on ARC-AGI-3. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. ARC-AGI-3 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ARC-AGI-3

Year

2026

Tasks

Interactive game-like tasks with hidden rules

Format

Agentic task completion under a capped evaluation budget

Difficulty

Frontier agentic reasoning

ARC-AGI-3 is distinct from ARC-AGI-2: it measures interactive, agentic reasoning rather than static grid-puzzle completion. BenchLM tracks published ARC Prize results as display-only until broad, comparable coverage supports a dedicated ranking lane.

BenchLM freshness & provenance

Version

ARC-AGI 3

Refresh cadence

Static

Staleness state

Current

Question availability

Private interactive tasks with public aggregate results

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ARC-AGI-3 measure?

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Which model scores highest on ARC-AGI-3?

Claude Opus 5 by Anthropic currently leads with a score of 30.2% on ARC-AGI-3.

How many models are evaluated on ARC-AGI-3?

11 AI models have been evaluated on ARC-AGI-3 on BenchLM.

Last updated: August 22, 2026 · BenchLM version ARC-AGI 3

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.