# Scale Labs MultiNRC (MultiNRC)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-multinrc

- Category: [Knowledge](/knowledge)
- Last updated: September 29, 2026 snapshot

## About MultiNRC

- Year: 2026
- Tasks: 44 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/multinrc)

BenchLM mirrors 44 published rows from the MultiNRC public table captured on September 29, 2026 snapshot.

MultiNRC is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (44 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 65.59% |
| 2 | [gpt-5-pro-2025-10-06](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 65.20% |
| 3 | [gemini-3.1-pro-preview](https://labs.scale.com/leaderboard/multinrc) | Google | 64.74% |
| 4 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 62.27% |
| 5 | [Muse Spark](/models/muse-spark) | Meta | 59.05% |
| 6 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/multinrc) | Google | 58.96% |
| 7 | [gpt-5.4-2026-03-05 (xhigh thinking)](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 58.29% |
| 8 | [claude-opus-4-6-thinking-max](https://labs.scale.com/leaderboard/multinrc) | Anthropic | 57.06% |
| 9 | [gpt-5-2025-08-07](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 52.13% |
| 10 | [gpt-5.1-thinking](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 49.00% |
| 11 | [o3-pro-2025-06-10-high](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 49.00% |
| 12 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 48.63% |
| 13 | [claude-opus-4-6 (Non-Thinking)](https://labs.scale.com/leaderboard/multinrc) | Anthropic | 48.34% |
| 14 | [o3-2025-04-16-high](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 45.50% |
| 15 | [Gemini-2.5-Pro-Preview-06-05](https://labs.scale.com/leaderboard/multinrc) | Google | 45.12% |
| 16 | [o3-2025-04-16-medium](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 44.45% |
| 17 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 42.18% |
| 18 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 41.23% |
| 19 | [claude-opus-4-1-20250805-thinking](https://labs.scale.com/leaderboard/multinrc) | Anthropic | 38.39% |
| 20 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | Anthropic | 35.83% |
| 21 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 35.17% |
| 22 | [Claude-4-Opus-20250514-thinking](https://labs.scale.com/leaderboard/multinrc) | Anthropic | 33.93% |
| 23 | [claude-opus-4-1-20250805](https://labs.scale.com/leaderboard/multinrc) | Anthropic | 29.67% |
| 24 | [Claude-4-Opus-20250514](https://labs.scale.com/leaderboard/multinrc) | Anthropic | 29.00% |
| 25 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 28.15% |
| 26 | [Claude-3.7-Sonnet-thinking](https://labs.scale.com/leaderboard/multinrc) | Deepseek | 27.77% |
| 27 | [Deepseek-R1-0528](https://labs.scale.com/leaderboard/multinrc) | Deepseek | 27.58% |
| 28 | [Qwen3-235B-A22B-Thinking-2507](https://labs.scale.com/leaderboard/multinrc) | Alibaba | 27.11% |
| 29 | [gemini-3.1-flash-lite-preview](https://labs.scale.com/leaderboard/multinrc) | Google | 25.02% |
| 30 | [DeepSeek-R1](/models/deepseek-r1) | DeepSeek | 24.27% |
| 31 | [GPT-5 mini](/models/gpt-5-mini) | OpenAI | 23.89% |
| 32 | [DeepSeek V3.1](/models/deepseek-v3-1) | DeepSeek | 23.60% |
| 33 | [o4-mini (high)](/models/o4-mini-high) | OpenAI | 22.18% |
| 34 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 21.23% |
| 35 | [kimi-k2-instruct](https://labs.scale.com/leaderboard/multinrc) | Moonshot AI | 18.48% |
| 36 | [Claude 4 Sonnet](/models/claude-4-sonnet) | Anthropic | 18.39% |
| 37 | [gpt-5.1-instant](https://labs.scale.com/leaderboard/multinrc) | OpenAI | 18.29% |
| 38 | [Qwen3-235B-A22B](https://labs.scale.com/leaderboard/multinrc) | Alibaba | 17.63% |
| 39 | [glm-4p5](https://labs.scale.com/leaderboard/multinrc) | Z.AI | 17.44% |
| 40 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 15.17% |
| 41 | [GPT-4o](/models/gpt-4o) | OpenAI | 12.42% |
| 42 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 10.43% |
| 43 | [glm-4p5-air](https://labs.scale.com/leaderboard/multinrc) | Z.AI | 10.43% |
| 44 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 8.44% |

## FAQ

### What does MultiNRC measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published MultiNRC snapshot?

Muse Spark 1.1 currently leads the published MultiNRC snapshot with a score of 65.59%.

### How many models are evaluated on MultiNRC?

The September 29, 2026 snapshot contains 44 AI models.

### Does MultiNRC affect BenchLM's overall score?

Not directly. MultiNRC is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
