Skip to main content

Benchmark profile

Toolathlon-Verified

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Data verified

Benchmark score on Toolathlon-Verified — July 29, 2026

BenchLM mirrors the published score view for Toolathlon-Verified. Claude Opus 5 leads the public snapshot at 80.6% , followed by Kimi K3 (73.2%) and Laguna S 2.1 (49.7%). BenchLM does not use these results to rank models overall.

3 modelsAgenticCurrentDisplay onlyUpdated July 29, 2026

Benchmark score table (3 models)

Score
1
Claude Opus 5Anthropic · Closed
80.6%
2
Kimi K3Moonshot AI · Closed
73.2%
3
Laguna S 2.1Poolside · Open weight
49.7%

The published Toolathlon-Verified snapshot places Claude Opus 5 first at 80.6%. The third row is 30.9 points behind. The broader top-10 range is 30.9 points, so the table still separates the published systems.

3 models have been evaluated on Toolathlon-Verified. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Toolathlon-Verified is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Toolathlon-Verified

Year

2026

Tasks

Verified multi-tool workflows

Format

Interactive tool-use score

Difficulty

Advanced tool use

BenchLM keeps Toolathlon-Verified separate from the broader Toolathlon key so provider launch values do not collapse distinct benchmark variants.

BenchLM freshness & provenance

Version

Toolathlon-Verified 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Toolathlon-Verified measure?

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Which model scores highest on Toolathlon-Verified?

Claude Opus 5 by Anthropic currently leads with a score of 80.6% on Toolathlon-Verified.

How many models are evaluated on Toolathlon-Verified?

3 AI models have been evaluated on Toolathlon-Verified on BenchLM.

Compare Top Models on Toolathlon-Verified

Last updated: July 29, 2026 · BenchLM version Toolathlon-Verified 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.