# Scale Labs MultiChallenge (MultiChallenge)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-multichallenge

- Category: [Instruction Following](/instruction-following)
- Last updated: September 29, 2026 snapshot

## About MultiChallenge

- Year: 2026
- Tasks: 30 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/multichallenge)

BenchLM mirrors 30 published rows from the MultiChallenge public table captured on September 29, 2026 snapshot.

MultiChallenge is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (30 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark](/models/muse-spark) | Meta | 75.52% |
| 2 | [Muse Spark 1.1\n](https://labs.scale.com/leaderboard/multichallenge) | Meta | 75.30% |
| 3 | [gemini-3.1-pro-preview](https://labs.scale.com/leaderboard/multichallenge) | Google | 71.37% |
| 4 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 69.23% |
| 5 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/multichallenge) | Google | 65.67% |
| 6 | [gpt-5.1-2025-11-13-thinking](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 63.41% |
| 7 | [gpt-5-thinking](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 63.19% |
| 8 | [o3-pro-2025-06-10-reasoning-high](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 62.40% |
| 9 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 61.39% |
| 10 | [gemini-3.1-flash-lite-preview](https://labs.scale.com/leaderboard/multichallenge) | Google | 60.61% |
| 11 | [gpt-5-mini-thinking](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 58.99% |
| 12 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 58.97% |
| 13 | [claude-4-opus-thinking](https://labs.scale.com/leaderboard/multichallenge) | Anthropic | 58.62% |
| 14 | [claude-opus-4-1-20250805-thinking](https://labs.scale.com/leaderboard/multichallenge) | Anthropic | 57.20% |
| 15 | [claude-4-sonnet-thinking](https://labs.scale.com/leaderboard/multichallenge) | Anthropic | 57.11% |
| 16 | [o3-2025-04-16-reasoning-high](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 56.62% |
| 17 | [claude-opus-4-6 (Non-Thinking)](https://labs.scale.com/leaderboard/multichallenge) | Anthropic | 56.02% |
| 18 | [kimi-k2-thinking](https://labs.scale.com/leaderboard/multichallenge) | Moonshot AI | 55.42% |
| 19 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | Anthropic | 55.32% |
| 20 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 53.62% |
| 21 | [claude-3-7-sonnet-thinking](https://labs.scale.com/leaderboard/multichallenge) | Anthropic | 51.58% |
| 22 | [gpt-5.1-2025-11-13-instant](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 51.23% |
| 23 | [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) | Anthropic | 50.49% |
| 24 | [deepseek-v3p1](https://labs.scale.com/leaderboard/multichallenge) | Deepseek | 46.10% |
| 25 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 45.34% |
| 26 | [o4-mini-2025-04-16-reasoning-high](https://labs.scale.com/leaderboard/multichallenge) | OpenAI | 44.90% |
| 27 | [qwen3-235b-a22b](https://labs.scale.com/leaderboard/multichallenge) | Alibaba | 41.22% |
| 28 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 39.43% |
| 29 | [claude-opus-4-6-thinking-max](https://labs.scale.com/leaderboard/multichallenge) | Anthropic | 37.15% |
| 30 | [gemini-2-0-flash](https://labs.scale.com/leaderboard/multichallenge) | Google | 36.35% |

## FAQ

### What does MultiChallenge measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published MultiChallenge snapshot?

Muse Spark currently leads the published MultiChallenge snapshot with a score of 75.52%.

### How many models are evaluated on MultiChallenge?

The September 29, 2026 snapshot contains 30 AI models.

### Does MultiChallenge affect BenchLM's overall score?

Not directly. MultiChallenge is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
