# Abstraction and Reasoning Corpus for AGI v3 (ARC-AGI-3)

> An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Canonical page: https://benchlm.ai/benchmarks/arcagi3

- Category: [Reasoning](/reasoning)
- Last updated: September 10, 2026

## About ARC-AGI-3

- Year: 2026
- Tasks: Interactive game-like tasks with hidden rules
- Format: Agentic task completion under a capped evaluation budget
- Difficulty: Frontier agentic reasoning
- Paper: [ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf)

ARC-AGI-3 is distinct from ARC-AGI-2: it measures interactive, agentic reasoning rather than static grid-puzzle completion. BenchLM tracks published ARC Prize results as display-only until broad, comparable coverage supports a dedicated ranking lane.

ARC-AGI-3 is currently weighted in BenchLM's scoring formula. The Reasoning category carries 17% of the overall score, and ARC-AGI-3 contributes 15% of that category score.

## Leaderboard (12 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 62.7% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 30.2% |
| 3 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 7.8% |
| 4 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 1.5% |
| 5 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 0.8% |
| 6 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 0.4% |
| 7 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 0.4% |
| 8 | [Grok 4.5](/models/grok-4-5) | xAI | 0.3% |
| 9 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 0.2% |
| 10 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 0.2% |
| 11 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 0.2% |
| 12 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 0.1% |

## FAQ

### What does ARC-AGI-3 measure?

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

### Which model scores highest on ARC-AGI-3?

GPT-6 Astra by OpenAI currently leads with a score of 62.7% on ARC-AGI-3.

### How many models are evaluated on ARC-AGI-3?

12 AI models have been evaluated on ARC-AGI-3 on BenchLM.

### Does ARC-AGI-3 affect BenchLM's overall score?

Yes. ARC-AGI-3 is a weighted benchmark inside the Reasoning category, which carries 17% of BenchLM's overall score. ARC-AGI-3 itself contributes 15% of that category score.

## Compare Top Models on ARC-AGI-3

- [GPT-6 Astra vs Claude Opus 5](/compare/claude-opus-5-vs-gpt-6-astra)
- [Claude Opus 5 vs GPT-5.6 Sol](/compare/claude-opus-5-vs-gpt-5-6-sol)
- [GPT-5.6 Sol vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-gpt-5-6-sol)
- [Claude Opus 4.8 vs GPT-5.6 Terra](/compare/claude-opus-4-8-vs-gpt-5-6-terra)
