# Toolathlon Verified Pass@3

> Fraction of Toolathlon Verified tasks solved in at least one of three trials.

Canonical page: https://benchlm.ai/benchmarks/toolathlonverifiedpass3

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About Toolathlon Verified Pass@3

- Year: 2026
- Tasks: 108 verified real-world tool-use tasks
- Format: Pass@3
- Difficulty: Long-horizon application tool use
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Table 8.13.6.A reports this alongside Pass@1, Pass³, and average turns from Anthropic's pinned internal harness.

Toolathlon Verified Pass@3 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 87.0% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 82.4% |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 81.5% |

## FAQ

### What does Toolathlon Verified Pass@3 measure?

Fraction of Toolathlon Verified tasks solved in at least one of three trials.

### Which model scores highest on Toolathlon Verified Pass@3?

Claude Opus 5 by Anthropic currently leads with a score of 87.0% on Toolathlon Verified Pass@3.

### How many models are evaluated on Toolathlon Verified Pass@3?

3 AI models have been evaluated on Toolathlon Verified Pass@3 on BenchLM.

### Does Toolathlon Verified Pass@3 affect BenchLM's overall score?

Not directly. Toolathlon Verified Pass@3 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on Toolathlon Verified Pass@3

- [Claude Opus 5 vs Claude Opus 5.5](/compare/claude-opus-5-vs-claude-opus-5-5)
- [Claude Opus 5.5 vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-claude-opus-5-5)
