# Toolathlon-Verified

> A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Canonical page: https://benchlm.ai/benchmarks/toolathlonverified

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About Toolathlon-Verified

- Year: 2026
- Tasks: Verified multi-tool workflows
- Format: Interactive tool-use score
- Difficulty: Advanced tool use
- Paper: [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3)

BenchLM keeps Toolathlon-Verified separate from the broader Toolathlon key so provider launch values do not collapse distinct benchmark variants.

BenchAlign v5.7 gives Toolathlon-Verified 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Leaderboard (20 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 80.6% |
| 2 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 78.4% |
| 3 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 77.8% |
| 4 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 77.8% |
| 5 | [MiMo-V2.6-Pro](/models/mimo-v2-6-pro) | Xiaomi | 76.9% |
| 6 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 74.1% |
| 7 | [Hy4 preview](/models/hy4-preview) | Tencent | 74.1% |
| 8 | [Step 5 Preview](/models/step-5-preview) | StepFun | 74.1% |
| 9 | [MiMo-V2.6-Flash](/models/mimo-v2-6-flash) | Xiaomi | 73.6% |
| 10 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 73.5% |
| 11 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 73.2% |
| 12 | [GLM-5.3](/models/glm-5-3) | Z.AI | 73.0% |
| 13 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 72.5% |
| 14 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 71.2% |
| 15 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 70.3% |
| 16 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 55.6% |
| 17 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 54.4% |
| 18 | [Laguna S 2.1](/models/laguna-s-2-1) | Poolside | 49.7% |
| 19 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 48.7% |
| 20 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 41.2% |

## FAQ

### What does Toolathlon-Verified measure?

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

### Which model scores highest on Toolathlon-Verified?

Claude Opus 5 by Anthropic currently leads with a score of 80.6% on Toolathlon-Verified.

### How many models are evaluated on Toolathlon-Verified?

20 AI models have been evaluated on Toolathlon-Verified on BenchLM.

### Does Toolathlon-Verified affect BenchLM's overall score?

Yes. BenchAlign v5.7 gives Toolathlon-Verified 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Compare Top Models on Toolathlon-Verified

- [Claude Opus 5 vs GLM-5.3-Flash](/compare/claude-opus-5-vs-glm-5-3-flash)
- [GLM-5.3-Flash vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-glm-5-3-flash)
- [Claude Opus 5.5 vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-claude-opus-5-5)
- [Claude Fable 5.1 vs MiMo-V2.6-Pro](/compare/claude-fable-5-1-vs-mimo-v2-6-pro)
