Skip to main content
BenchLM
Data

Abstraction and Reasoning Corpus for AGI v3 (ARC-AGI-3)

Data verified 36 confirmed releases in the last 30 daysFollow model changes

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Earlier track

ARC-AGI-2 evaluates static grid-puzzle reasoning. It remains a separate benchmark track, and its scores are not directly comparable with ARC-AGI-3.

View ARC-AGI-2

Top models on ARC-AGI-3 — October 10, 2026

As of October 10, 2026, GPT-6 Astra leads the ARC-AGI-3 leaderboard with 62.7% , followed by GPT-6.1 Sol (52.7%) and Claude Opus 5 (30.2%).

17 modelsReasoning15% of category scoreCurrentUpdated October 10, 2026

Leaderboard (17 models)

ARC-AGI-3 leaderboard
RankModel / configurationScoreParameters (B)Open / closed
1GPT-6 AstraOpenAI
62.7%
Not reportedClosed
2GPT-6.1 SolOpenAI
52.7%
Not reportedClosed
3Claude Opus 5Anthropic
30.2%
Not reportedClosed
4Gemini 3.8 FlashGoogle
10.4%
Not reportedClosed
5GPT-5.6 SolOpenAI
7.8%
Not reportedClosed
6GPT-6 SolOpenAI
4.6%
Not reportedClosed
7Grok 4.6xAI
2.1%
Not reportedClosed
8Claude Opus 4.8Anthropic
1.5%
Not reportedClosed
9GPT-5.6 TerraOpenAI
0.8%
Not reportedClosed
10GPT-5.5OpenAI
0.4%
Not reportedClosed
11Gemini 3.1 ProGoogle
0.4%
Not reportedClosed
12Grok 4.5xAI
0.3%
Not reportedClosed
13GPT-5.4OpenAI
0.2%
Not reportedClosed
14GPT-5.6 LunaOpenAI
0.2%
Not reportedClosed
15Claude Opus 4.7 (Adaptive)Anthropic
0.2%
Not reportedClosed
16GPT-6 LunaOpenAI
0.1%
Not reportedClosed
17Grok 4.20xAI
0.1%
Not reportedClosed

According to BenchLM.ai, GPT-6 Astra leads the ARC-AGI-3 benchmark with a score of 62.7%, followed by GPT-6.1 Sol (52.7%) and Claude Opus 5 (30.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

17 models have been evaluated on ARC-AGI-3. The benchmark falls in the Reasoning category. The Reasoning leaderboard ranks models by a weighted category score, and ARC-AGI-3 contributes 15% of it. BenchAlign v5.8 also uses it in the overall ranking.

About ARC-AGI-3

Year

2026

Tasks

Interactive game-like tasks with hidden rules

Format

Agentic task completion under a capped evaluation budget

Difficulty

Frontier agentic reasoning

ARC-AGI-3 measures interactive reasoning rather than static grid-puzzle completion. ARC Prize reports Standard and Provider Adapter harness runs separately. The single ARC-AGI-3 score field stores the Standard harness result; model evidence notes preserve distinct Provider Adapter results.

Freshness and provenance

Version

ARC-AGI 3

Refresh cadence

Static

Staleness state

Current

Question availability

Private interactive tasks with public aggregate results

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does ARC-AGI-3 measure?

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Which model scores highest on ARC-AGI-3?

GPT-6 Astra by OpenAI currently leads with a score of 62.7% on ARC-AGI-3.

How many models are evaluated on ARC-AGI-3?

17 AI models have published results on ARC-AGI-3 in the BenchLM catalog.

Last updated: October 10, 2026 · BenchLM version ARC-AGI 3

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.