# Toolathlon

> A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Canonical page: https://benchlm.ai/benchmarks/toolathlon

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About Toolathlon

- Year: 2026
- Tasks: Multi-tool workflows
- Format: Interactive tool-calling evaluation
- Difficulty: Advanced tool use
- Paper: [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/)

Toolathlon is useful for judging whether a model can do more than answer in chat and instead complete multi-step tool workflows.

BenchAlign v5.7 gives Toolathlon 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Leaderboard (22 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 75.6% |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 59.9% |
| 3 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 58% |
| 4 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 56.5% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 55.6% |
| 6 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 54.6% |
| 7 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 53.4% |
| 8 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 53.1% |
| 9 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 51.8% |
| 10 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 50% |
| 11 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 49.5% |
| 12 | [GLM-5.2](/models/glm-5-2) | Z.AI | 48.2% |
| 13 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 47.8% |
| 14 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 46.3% |
| 15 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 43.5% |
| 16 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 42.9% |
| 17 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 39.8% |
| 18 | [GLM-5](/models/glm-5) | Z.AI | 38% |
| 19 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 36.3% |
| 20 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 35.5% |
| 21 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 27.8% |
| 22 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 26.9% |

## FAQ

### What does Toolathlon measure?

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

### Which model scores highest on Toolathlon?

Muse Spark 1.1 by Meta currently leads with a score of 75.6% on Toolathlon.

### How many models are evaluated on Toolathlon?

22 AI models have been evaluated on Toolathlon on BenchLM.

### Does Toolathlon affect BenchLM's overall score?

Yes. BenchAlign v5.7 gives Toolathlon 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Compare Top Models on Toolathlon

- [Muse Spark 1.1 vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-muse-spark-1-1)
- [Claude Opus 4.8 vs GPT-5.6 Sol](/compare/claude-opus-4-8-vs-gpt-5-6-sol)
- [GPT-5.6 Sol vs Gemini 3.5 Flash](/compare/gemini-3-5-flash-vs-gpt-5-6-sol)
- [Gemini 3.5 Flash vs GPT-5.5](/compare/gemini-3-5-flash-vs-gpt-5-5)
