Abstraction and Reasoning Corpus for AGI v3 (ARC-AGI-3)
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
Earlier track
ARC-AGI-2 evaluates static grid-puzzle reasoning. It remains a separate benchmark track, and its scores are not directly comparable with ARC-AGI-3.
View ARC-AGI-2Top models on ARC-AGI-3 — October 10, 2026
As of October 10, 2026, GPT-6 Astra leads the ARC-AGI-3 leaderboard with 62.7% , followed by GPT-6.1 Sol (52.7%) and Claude Opus 5 (30.2%).
GPT-6 Astra
OpenAI
GPT-6.1 Sol
OpenAI
Claude Opus 5
Anthropic
17 modelsReasoning15% of category scoreCurrentUpdated October 10, 2026
Leaderboard (17 models)
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 62.7% | Not reported | Closed |
| 2 | GPT-6.1 SolOpenAI | 52.7% | Not reported | Closed |
| 3 | Claude Opus 5Anthropic | 30.2% | Not reported | Closed |
| 4 | Gemini 3.8 FlashGoogle | 10.4% | Not reported | Closed |
| 5 | GPT-5.6 SolOpenAI | 7.8% | Not reported | Closed |
| 6 | GPT-6 SolOpenAI | 4.6% | Not reported | Closed |
| 7 | Grok 4.6xAI | 2.1% | Not reported | Closed |
| 8 | Claude Opus 4.8Anthropic | 1.5% | Not reported | Closed |
| 9 | GPT-5.6 TerraOpenAI | 0.8% | Not reported | Closed |
| 10 | GPT-5.5OpenAI | 0.4% | Not reported | Closed |
| 11 | Gemini 3.1 ProGoogle | 0.4% | Not reported | Closed |
| 12 | Grok 4.5xAI | 0.3% | Not reported | Closed |
| 13 | GPT-5.4OpenAI | 0.2% | Not reported | Closed |
| 14 | GPT-5.6 LunaOpenAI | 0.2% | Not reported | Closed |
| 15 | Claude Opus 4.7 (Adaptive)Anthropic | 0.2% | Not reported | Closed |
| 16 | GPT-6 LunaOpenAI | 0.1% | Not reported | Closed |
| 17 | Grok 4.20xAI | 0.1% | Not reported | Closed |
According to BenchLM.ai, GPT-6 Astra leads the ARC-AGI-3 benchmark with a score of 62.7%, followed by GPT-6.1 Sol (52.7%) and Claude Opus 5 (30.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
17 models have been evaluated on ARC-AGI-3. The benchmark falls in the Reasoning category. The Reasoning leaderboard ranks models by a weighted category score, and ARC-AGI-3 contributes 15% of it. BenchAlign v5.8 also uses it in the overall ranking.
About ARC-AGI-3
Year
2026
Tasks
Interactive game-like tasks with hidden rules
Format
Agentic task completion under a capped evaluation budget
Difficulty
Frontier agentic reasoning
ARC-AGI-3 measures interactive reasoning rather than static grid-puzzle completion. ARC Prize reports Standard and Provider Adapter harness runs separately. The single ARC-AGI-3 score field stores the Standard harness result; model evidence notes preserve distinct Provider Adapter results.
Freshness and provenance
Version
ARC-AGI 3
Refresh cadence
Static
Staleness state
Current
Question availability
Private interactive tasks with public aggregate results
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ARC-AGI-3 measure?
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
Which model scores highest on ARC-AGI-3?
GPT-6 Astra by OpenAI currently leads with a score of 62.7% on ARC-AGI-3.
How many models are evaluated on ARC-AGI-3?
17 AI models have published results on ARC-AGI-3 in the BenchLM catalog.
Compare top models on ARC-AGI-3
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.