# ARC-AGI-1 Semi-Private Evaluation (ARC-AGI-1)

> ARC Prize fluid-intelligence benchmark using novel visual grid transformations.

Canonical page: https://benchlm.ai/benchmarks/arcagi1

- Category: [Reasoning](/reasoning)
- Last updated: September 18, 2026

## About ARC-AGI-1

- Year: 2026
- Tasks: Semi-private ARC-AGI-1 evaluation set
- Format: Verified accuracy
- Difficulty: Abstract visual reasoning
- Paper: [ARC Prize leaderboard](https://arcprize.org/)

The Opus 5 system card reports the ARC Prize Foundation's verified semi-private max-effort result.

ARC-AGI-1 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 98.50% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 97.50% |
| 3 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 97.50% |

## FAQ

### What does ARC-AGI-1 measure?

ARC Prize fluid-intelligence benchmark using novel visual grid transformations.

### Which model scores highest on ARC-AGI-1?

GPT-6 Astra by OpenAI currently leads with a score of 98.50% on ARC-AGI-1.

### How many models are evaluated on ARC-AGI-1?

3 AI models have been evaluated on ARC-AGI-1 on BenchLM.

### Does ARC-AGI-1 affect BenchLM's overall score?

Not directly. ARC-AGI-1 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on ARC-AGI-1

- [GPT-6 Astra vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-gpt-6-astra)
- [Claude Fable 5.1 vs Claude Opus 5](/compare/claude-fable-5-1-vs-claude-opus-5)
