Skip to main content
BenchLM

Toolathlon

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Top models on Toolathlon — September 27, 2026

As of September 27, 2026, Muse Spark 1.1 leads the Toolathlon leaderboard with 75.6% , followed by Claude Opus 4.8 (59.9%) and GPT-5.6 Sol (58%).

26 modelsAgentic3% of Agentic reference weightCurrentUpdated September 27, 2026

Leaderboard (26 models)

Score
1
Muse Spark 1.1Meta · Closed
75.6%
2
Claude Opus 4.8Anthropic · Closed
59.9%
3
GPT-5.6 SolOpenAI · Closed
58%
4
Gemini 3.5 FlashGoogle · Closed
56.5%
5
GPT-5.5OpenAI · Closed
55.6%
6
GPT-5.4OpenAI · Closed
54.6%
7
GPT-5.6 LunaOpenAI · Closed
53.4%
8
GPT-5.6 TerraOpenAI · Closed
53.1%
9
DeepSeek V4 Pro 0813DeepSeek · Open weight
51.8%
10
Kimi K2.6Moonshot AI · Open weight
50%
11
Step 3.7 FlashStepFun · Open weight
49.5%
12
DeepSeek V4 Pro (High)DeepSeek · Open weight
49%
13
GLM-5.2Z.AI · Open weight
48.2%
14
DeepSeek V4 Flash 0731DeepSeek · Open weight
47.8%
15
MiniMax M2.7MiniMax · Open weight
46.3%
16
DeepSeek V4 ProDeepSeek · Open weight
46.3%
17
Claude Opus 4.5Anthropic · Closed
43.5%
18
DeepSeek V4 Flash (High)DeepSeek · Open weight
43.5%
19
GPT-5.4 miniOpenAI · Closed
42.9%
20
DeepSeek V4 FlashDeepSeek · Open weight
40.7%
21
Qwen3.6 PlusAlibaba · Closed
39.8%
22
GLM-5Z.AI · Open weight
38%
23
Qwen3.5 397BAlibaba · Open weight
36.3%
24
GPT-5.4 nanoOpenAI · Closed
35.5%
25
Kimi K2.5Moonshot AI · Open weight
27.8%
26
Qwen3.6-35B-A3BAlibaba · Open weight
26.9%

According to BenchLM.ai, Muse Spark 1.1 leads the Toolathlon benchmark with a score of 75.6%, followed by Claude Opus 4.8 (59.9%) and GPT-5.6 Sol (58%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

26 models have been evaluated on Toolathlon. The benchmark falls in the Agentic category. BenchAlign v5.7 gives Toolathlon 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About Toolathlon

Year

2026

Tasks

Multi-tool workflows

Format

Interactive tool-calling evaluation

Difficulty

Advanced tool use

Toolathlon is useful for judging whether a model can do more than answer in chat and instead complete multi-step tool workflows.

Freshness and provenance

Version

Toolathlon 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Toolathlon measure?

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Which model scores highest on Toolathlon?

Muse Spark 1.1 by Meta currently leads with a score of 75.6% on Toolathlon.

How many models are evaluated on Toolathlon?

26 AI models have been evaluated on Toolathlon on BenchLM.

Last updated: September 27, 2026 · BenchLM version Toolathlon 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.