# Scale Labs Visual Tool Bench (Visual Tool Bench)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-vtb

- Category: [Agentic](/agentic)
- Last updated: September 29, 2026 snapshot

## About Visual Tool Bench

- Year: 2026
- Tasks: 21 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/vtb)

BenchLM mirrors 21 published rows from the Visual Tool Bench public table captured on September 29, 2026 snapshot.

Visual Tool Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (21 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 44.77% |
| 2 | [gpt-5.4-2026-03-05 (reasoning effort = high)](https://labs.scale.com/leaderboard/vtb) | OpenAI | 29.17% |
| 3 | [gemini-3.1-pro-preview](https://labs.scale.com/leaderboard/vtb) | Google | 28.97% |
| 4 | [Claude Opus 4.6 (Adaptive)](/models/claude-opus-4-6-thinking) | Anthropic | 27.52% |
| 5 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/vtb) | Google | 26.85% |
| 6 | [gpt-5-2025-08-07-thinking](https://labs.scale.com/leaderboard/vtb) | OpenAI | 18.68% |
| 7 | [gpt-5-2025-08-07](https://labs.scale.com/leaderboard/vtb) | OpenAI | 16.96% |
| 8 | [o3](/models/o3) | OpenAI | 13.74% |
| 9 | [gemini-2.5-pro-preview-06-05](https://labs.scale.com/leaderboard/vtb) | Google | 11.75% |
| 10 | [o4-mini-2025-04-16](https://labs.scale.com/leaderboard/vtb) | OpenAI | 11.12% |
| 11 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | Anthropic | 6.20% |
| 12 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 5.60% |
| 13 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 5.52% |
| 14 | [claude-opus-4-1-20250805-thinking](https://labs.scale.com/leaderboard/vtb) | Anthropic | 5.16% |
| 15 | [claude-opus-4-1-20250805](https://labs.scale.com/leaderboard/vtb) | Anthropic | 4.71% |
| 16 | [Gemini 2.5 Flash](/models/gemini-2-5-flash) | Google | 4.69% |
| 17 | [claude-sonnet-4](https://labs.scale.com/leaderboard/vtb) | Anthropic | 4.48% |
| 18 | [claude-sonnet-4-thinking](https://labs.scale.com/leaderboard/vtb) | Anthropic | 4.44% |
| 19 | [nova-premier](https://labs.scale.com/leaderboard/vtb) | Amazon | 2.00% |
| 20 | [llama4-scout](https://labs.scale.com/leaderboard/vtb) | Meta | 1.58% |
| 21 | [llama4-maverick](https://labs.scale.com/leaderboard/vtb) | Meta | 1.41% |

## FAQ

### What does Visual Tool Bench measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published Visual Tool Bench snapshot?

Muse Spark 1.1 currently leads the published Visual Tool Bench snapshot with a score of 44.77%.

### How many models are evaluated on Visual Tool Bench?

The September 29, 2026 snapshot contains 21 AI models.

### Does Visual Tool Bench affect BenchLM's overall score?

Not directly. Visual Tool Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
