# Scale Labs TutorBench (TutorBench)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-tutorbench

- Category: [Knowledge](/knowledge)
- Last updated: September 29, 2026 snapshot

## About TutorBench

- Year: 2026
- Tasks: 27 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/tutorbench)

BenchLM mirrors 27 published rows from the TutorBench public table captured on September 29, 2026 snapshot.

TutorBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (27 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark](/models/muse-spark) | Meta | 68.55% |
| 2 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 56.62% |
| 3 | [gemini-2.5-pro-preview-06-05](https://labs.scale.com/leaderboard/tutorbench) | Google | 55.65% |
| 4 | [gpt-5-2025-08-07](https://labs.scale.com/leaderboard/tutorbench) | OpenAI | 55.33% |
| 5 | [o3-pro](/models/o3-pro) | OpenAI | 54.62% |
| 6 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 54.56% |
| 7 | [gpt-5.1-thinking](https://labs.scale.com/leaderboard/tutorbench) | OpenAI | 54.09% |
| 8 | [claude-opus-4-6-thinking-max](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 53.68% |
| 9 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/tutorbench) | Google | 53.67% |
| 10 | [claude-opus-4-6 (Non-Thinking)](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 53.55% |
| 11 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 53.49% |
| 12 | [gemini-3.1-pro-preview](https://labs.scale.com/leaderboard/tutorbench) | Google | 52.99% |
| 13 | [o3-2025-04-16-medium](https://labs.scale.com/leaderboard/tutorbench) | OpenAI | 52.76% |
| 14 | [o3-2025-04-16-high](https://labs.scale.com/leaderboard/tutorbench) | OpenAI | 52.09% |
| 15 | [gemini-3.1-flash-lite-preview](https://labs.scale.com/leaderboard/tutorbench) | Google | 51.50% |
| 16 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 51.20% |
| 17 | [claude-opus-4-1-20250805-thinking](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 50.78% |
| 18 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 49.82% |
| 19 | [claude-4-opus-20250514-thinking](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 49.71% |
| 20 | [gpt-5.1-instant](https://labs.scale.com/leaderboard/tutorbench) | OpenAI | 49.08% |
| 21 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | Anthropic | 49.00% |
| 22 | [claude-opus-4-1-20250805_anthropic](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 47.40% |
| 23 | [claude-37-sonnet-thinking](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 46.45% |
| 24 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 45.70% |
| 25 | [claude-opus-4-20250514](https://labs.scale.com/leaderboard/tutorbench) | Anthropic | 45.46% |
| 26 | [llama4-maverick](https://labs.scale.com/leaderboard/tutorbench) | Meta | 40.20% |
| 27 | [GPT-4o](/models/gpt-4o) | OpenAI | 36.12% |

## FAQ

### What does TutorBench measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published TutorBench snapshot?

Muse Spark currently leads the published TutorBench snapshot with a score of 68.55%.

### How many models are evaluated on TutorBench?

The September 29, 2026 snapshot contains 27 AI models.

### Does TutorBench affect BenchLM's overall score?

Not directly. TutorBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
