Skip to main content
BenchLM

Toolathlon-Verified

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Top models on Toolathlon-Verified — September 27, 2026

As of September 27, 2026, Claude Opus 5 leads the Toolathlon-Verified leaderboard with 80.6% , followed by GLM-5.3-Flash (78.4%) and Claude Fable 5.1 (77.8%).

20 modelsAgentic3% of Agentic reference weightCurrentUpdated September 27, 2026

Leaderboard (20 models)

Score
1
Claude Opus 5Anthropic · Closed
80.6%
2
GLM-5.3-FlashZ.AI · Open weight
78.4%
3
Claude Fable 5.1Anthropic · Closed
77.8%
4
Claude Opus 5.5Anthropic · Closed
77.8%
5
MiMo-V2.6-ProXiaomi · Open weight
76.9%
6
DeepSeek V4 Pro 0813DeepSeek · Open weight
74.1%
7
Step 5 PreviewStepFun · Closed
74.1%
8
Hy4 previewTencent · Open weight
74.1%
9
MiMo-V2.6-FlashXiaomi · Open weight
73.6%
10
Qwen3.8-Flash-NextAlibaba · Open weight
73.5%
11
Kimi K3Moonshot AI · Closed
73.2%
12
GLM-5.3Z.AI · Open weight
73.0%
13
Qwen3.8 MaxAlibaba · Open weight
72.5%
14
Ornith-1.5-397BOrnith AI · Open weight
71.2%
15
DeepSeek V4 Flash 0731DeepSeek · Open weight
70.3%
16
dots3-note PreviewDots Studio · Open weight
55.6%
17
Inkling-SmallThinking Machines Lab · Open weight
54.4%
18
Laguna S 2.1Poolside · Open weight
49.7%
19
Ornith-1.5-35B-A3BOrnith AI · Open weight
48.7%
20
Ornith-1.5-9BOrnith AI · Open weight
41.2%

According to BenchLM.ai, Claude Opus 5 leads the Toolathlon-Verified benchmark with a score of 80.6%, followed by GLM-5.3-Flash (78.4%) and Claude Fable 5.1 (77.8%). The top models are clustered within 2.8 points, suggesting this benchmark is nearing saturation for frontier models.

20 models have been evaluated on Toolathlon-Verified. The benchmark falls in the Agentic category. BenchAlign v5.7 gives Toolathlon-Verified 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About Toolathlon-Verified

Year

2026

Tasks

Verified multi-tool workflows

Format

Interactive tool-use score

Difficulty

Advanced tool use

BenchLM keeps Toolathlon-Verified separate from the broader Toolathlon key so provider launch values do not collapse distinct benchmark variants.

Freshness and provenance

Version

Toolathlon-Verified 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Toolathlon-Verified measure?

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Which model scores highest on Toolathlon-Verified?

Claude Opus 5 by Anthropic currently leads with a score of 80.6% on Toolathlon-Verified.

How many models are evaluated on Toolathlon-Verified?

20 AI models have been evaluated on Toolathlon-Verified on BenchLM.

Last updated: September 27, 2026 · BenchLM version Toolathlon-Verified 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.