# Scale Labs MASK (MASK)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-mask

- Category: [Instruction Following](/instruction-following)
- Last updated: September 29, 2026 snapshot

## About MASK

- Year: 2026
- Tasks: 67 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/mask)

BenchLM mirrors 67 published rows from the MASK public table captured on September 29, 2026 snapshot.

MASK is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (67 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [claude-opus-4-6 (Non-Thinking)](https://labs.scale.com/leaderboard/mask) | Anthropic | 96.28% |
| 2 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | Anthropic | 96.13% |
| 3 | [Claude Sonnet 4 (Thinking)](https://labs.scale.com/leaderboard/mask) | Anthropic | 95.33% |
| 4 | [claude-opus-4-1-20250805-thinking](https://labs.scale.com/leaderboard/mask) | Anthropic | 94.20% |
| 5 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 92.53% |
| 6 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 92.00% |
| 7 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 91.73% |
| 8 | [gpt-5.4-2026-03-05 (xhigh thinking)](https://labs.scale.com/leaderboard/mask) | OpenAI | 89.67% |
| 9 | [Claude Sonnet 4](https://labs.scale.com/leaderboard/mask) | Anthropic | 89.27% |
| 10 | [Claude Opus 4 (Thinking)](https://labs.scale.com/leaderboard/mask) | Anthropic | 87.87% |
| 11 | [claude-opus-4-1-20250805](https://labs.scale.com/leaderboard/mask) | Anthropic | 87.40% |
| 12 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 87.13% |
| 13 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 86.67% |
| 14 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 86.46% |
| 15 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 86.40% |
| 16 | [gpt-5.1-thinking](https://labs.scale.com/leaderboard/mask) | OpenAI | 86.33% |
| 17 | [gpt-5-pro-2025-10-06](https://labs.scale.com/leaderboard/mask) | OpenAI | 85.99% |
| 18 | [claude-opus-4-6-thinking-max](https://labs.scale.com/leaderboard/mask) | Anthropic | 85.40% |
| 19 | [o3 (high) (April 2025)](https://labs.scale.com/leaderboard/mask) | OpenAI | 84.47% |
| 20 | [o3 (medium) (April 2025)](https://labs.scale.com/leaderboard/mask) | OpenAI | 82.60% |
| 21 | [GPT-5 mini](/models/gpt-5-mini) | OpenAI | 82.60% |
| 22 | [o3 Pro (high) (June 2025)](https://labs.scale.com/leaderboard/mask) | OpenAI | 82.50% |
| 23 | [Claude 3.7 Sonnet (Thinking) (February 2025)](https://labs.scale.com/leaderboard/mask) | Anthropic | 82.13% |
| 24 | [Claude Opus 4](https://labs.scale.com/leaderboard/mask) | Anthropic | 80.28% |
| 25 | [gpt-5-2025-08-07](https://labs.scale.com/leaderboard/mask) | OpenAI | 79.33% |
| 26 | [Claude 3 Opus](/models/claude-3-opus) | Anthropic | 79.00% |
| 27 | [o4-mini (high) (April 2025)](https://labs.scale.com/leaderboard/mask) | OpenAI | 78.60% |
| 28 | [o4-mini (medium) (April 2025)](https://labs.scale.com/leaderboard/mask) | OpenAI | 72.93% |
| 29 | [Claude 3.5 Sonnet (October 2024)](https://labs.scale.com/leaderboard/mask) | Anthropic | 72.33% |
| 30 | [Claude 3.7 Sonnet (February 2025)](https://labs.scale.com/leaderboard/mask) | Anthropic | 72.27% |
| 31 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 70.47% |
| 32 | [gpt-5.1-instant](https://labs.scale.com/leaderboard/mask) | OpenAI | 63.33% |
| 33 | [o1-pro](/models/o1-pro) | OpenAI | 61.60% |
| 34 | [glm-4p5](https://labs.scale.com/leaderboard/mask) | Z.AI | 61.46% |
| 35 | [GPT-4.1 nano\n](https://labs.scale.com/leaderboard/mask) | OpenAI | 61.40% |
| 36 | [Llama 3.1 405B Instruct](https://labs.scale.com/leaderboard/mask) | Meta | 61.40% |
| 37 | [glm-4p5-air](https://labs.scale.com/leaderboard/mask) | Z.AI | 60.80% |
| 38 | [gpt 4o (November 2024)](https://labs.scale.com/leaderboard/mask) | OpenAI | 60.07% |
| 39 | [o1 (December 2024)](https://labs.scale.com/leaderboard/mask) | OpenAI | 59.27% |
| 40 | [Deepseek R1 (Jan 2025)](https://labs.scale.com/leaderboard/mask) | Deepseek | 57.32% |
| 41 | [GPT 4.5 Preview](https://labs.scale.com/leaderboard/mask) | OpenAI | 56.93% |
| 42 | [Mistral Magistral](https://labs.scale.com/leaderboard/mask) | Mistral | 56.50% |
| 43 | [Qwen3-235B-A22B](https://labs.scale.com/leaderboard/mask) | Alibaba | 56.40% |
| 44 | [Gemini 2.5 Pro Experimental (March 2025)](https://labs.scale.com/leaderboard/mask) | Google | 55.93% |
| 45 | [gemini-2.5-pro-preview-06-05](https://labs.scale.com/leaderboard/mask) | Google | 55.67% |
| 46 | [Llama 3.2 90B Vision Instruct](https://labs.scale.com/leaderboard/mask) | Meta | 54.07% |
| 47 | [Gemini 2.5 Pro Preview (May 06 2025)](https://labs.scale.com/leaderboard/mask) | Google | 53.07% |
| 48 | [DeepSeek-R1-0528](https://labs.scale.com/leaderboard/mask) | Deepseek | 53.00% |
| 49 | [Llama 3.3 70B Instruct](https://labs.scale.com/leaderboard/mask) | Meta | 51.93% |
| 50 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 51.13% |
| 51 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 50.00% |
| 52 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 49.73% |
| 53 | [o3 mini (Low)](https://labs.scale.com/leaderboard/mask) | OpenAI | 49.73% |
| 54 | [Gemini 2.0 Flash Thinking (January 2025)](https://labs.scale.com/leaderboard/mask) | Google | 49.53% |
| 55 | [Gemini 2.5 Flash Preview (May 2025)](https://labs.scale.com/leaderboard/mask) | Google | 49.13% |
| 56 | [Gemini 2.0 Flash](https://labs.scale.com/leaderboard/mask) | Google | 49.07% |
| 57 | [o3 mini (Medium)](https://labs.scale.com/leaderboard/mask) | OpenAI | 48.93% |
| 58 | [Gemini 2.0 Pro Experimental (February 2025)](https://labs.scale.com/leaderboard/mask) | Google | 48.67% |
| 59 | [gemini-3.1-flash-lite-preview](https://labs.scale.com/leaderboard/mask) | Google | 48.40% |
| 60 | [Mistral Large 2411](https://labs.scale.com/leaderboard/mask) | Mistral | 47.53% |
| 61 | [o3 mini (High)](https://labs.scale.com/leaderboard/mask) | OpenAI | 46.80% |
| 62 | [kimi-k2-instruct](https://labs.scale.com/leaderboard/mask) | Moonshot AI | 46.67% |
| 63 | [DeepSeek V3.1](/models/deepseek-v3-1) | DeepSeek | 46.27% |
| 64 | [Deepseek V3 (March 2025)](https://labs.scale.com/leaderboard/mask) | Deepseek | 44.53% |
| 65 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/mask) | Google | 42.60% |
| 66 | [Mistral Medium 3](/models/mistral-medium-3) | Mistral | 42.60% |
| 67 | [gemini-3.1-pro-preview](https://labs.scale.com/leaderboard/mask) | Google | 42.40% |

## FAQ

### What does MASK measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published MASK snapshot?

claude-opus-4-6 (Non-Thinking) currently leads the published MASK snapshot with a score of 96.28%.

### How many models are evaluated on MASK?

The September 29, 2026 snapshot contains 67 AI models.

### Does MASK affect BenchLM's overall score?

Not directly. MASK is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
