Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Abstraction and Reasoning Corpus for AGI v3 (ARC-AGI-3)

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Earlier track

ARC-AGI-2 evaluates static grid-puzzle reasoning. It remains a separate benchmark track, and its scores are not directly comparable with ARC-AGI-3.

View ARC-AGI-2
Data verified 37 confirmed releases in the last 30 daysSee provider release alerts

Top models on ARC-AGI-3 — September 10, 2026

As of September 10, 2026, GPT-6 Astra leads the ARC-AGI-3 leaderboard with 62.7% , followed by Claude Opus 5 (30.2%) and GPT-5.6 Sol (7.8%).

12 modelsReasoning15% of category scoreCurrentUpdated September 10, 2026

Leaderboard (12 models)

Score
1
GPT-6 AstraOpenAI · Closed
62.7%
2
Claude Opus 5Anthropic · Closed
30.2%
3
GPT-5.6 SolOpenAI · Closed
7.8%
4
Claude Opus 4.8Anthropic · Closed
1.5%
5
GPT-5.6 TerraOpenAI · Closed
0.8%
6
GPT-5.5OpenAI · Closed
0.4%
7
Gemini 3.1 ProGoogle · Closed
0.4%
8
Grok 4.5xAI · Closed
0.3%
9
GPT-5.4OpenAI · Closed
0.2%
10
GPT-5.6 LunaOpenAI · Closed
0.2%
11
Claude Opus 4.7 (Adaptive)Anthropic · Closed
0.2%
12
Grok 4.20xAI · Closed
0.1%

According to BenchLM.ai, GPT-6 Astra leads the ARC-AGI-3 benchmark with a score of 62.7%, followed by Claude Opus 5 (30.2%) and GPT-5.6 Sol (7.8%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

12 models have been evaluated on ARC-AGI-3. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, ARC-AGI-3 contributes 15% of the category score, so strong performance here directly affects a model's overall ranking.

About ARC-AGI-3

Year

2026

Tasks

Interactive game-like tasks with hidden rules

Format

Agentic task completion under a capped evaluation budget

Difficulty

Frontier agentic reasoning

ARC-AGI-3 is distinct from ARC-AGI-2: it measures interactive, agentic reasoning rather than static grid-puzzle completion. BenchLM tracks published ARC Prize results as display-only until broad, comparable coverage supports a dedicated ranking lane.

BenchLM freshness & provenance

Version

ARC-AGI 3

Refresh cadence

Static

Staleness state

Current

Question availability

Private interactive tasks with public aggregate results

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ARC-AGI-3 measure?

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Which model scores highest on ARC-AGI-3?

GPT-6 Astra by OpenAI currently leads with a score of 62.7% on ARC-AGI-3.

How many models are evaluated on ARC-AGI-3?

12 AI models have been evaluated on ARC-AGI-3 on BenchLM.

Last updated: September 10, 2026 · BenchLM version ARC-AGI 3

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.