Abstraction and Reasoning Corpus for AGI v3 (ARC-AGI-3)
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
Earlier track
ARC-AGI-2 evaluates static grid-puzzle reasoning. It remains a separate benchmark track, and its scores are not directly comparable with ARC-AGI-3.
View ARC-AGI-2Top models on ARC-AGI-3 — September 10, 2026
As of September 10, 2026, GPT-6 Astra leads the ARC-AGI-3 leaderboard with 62.7% , followed by Claude Opus 5 (30.2%) and GPT-5.6 Sol (7.8%).
GPT-6 Astra
OpenAI
Claude Opus 5
Anthropic
GPT-5.6 Sol
OpenAI
12 modelsReasoning15% of category scoreCurrentUpdated September 10, 2026
Leaderboard (12 models)
ScoreAccording to BenchLM.ai, GPT-6 Astra leads the ARC-AGI-3 benchmark with a score of 62.7%, followed by Claude Opus 5 (30.2%) and GPT-5.6 Sol (7.8%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
12 models have been evaluated on ARC-AGI-3. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, ARC-AGI-3 contributes 15% of the category score, so strong performance here directly affects a model's overall ranking.
About ARC-AGI-3
Year
2026
Tasks
Interactive game-like tasks with hidden rules
Format
Agentic task completion under a capped evaluation budget
Difficulty
Frontier agentic reasoning
ARC-AGI-3 is distinct from ARC-AGI-2: it measures interactive, agentic reasoning rather than static grid-puzzle completion. BenchLM tracks published ARC Prize results as display-only until broad, comparable coverage supports a dedicated ranking lane.
BenchLM freshness & provenance
Version
ARC-AGI 3
Refresh cadence
Static
Staleness state
Current
Question availability
Private interactive tasks with public aggregate results
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does ARC-AGI-3 measure?
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
Which model scores highest on ARC-AGI-3?
GPT-6 Astra by OpenAI currently leads with a score of 62.7% on ARC-AGI-3.
How many models are evaluated on ARC-AGI-3?
12 AI models have been evaluated on ARC-AGI-3 on BenchLM.
Compare Top Models on ARC-AGI-3
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.