# Scale Labs FORTRESS (FORTRESS)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings. The catalog still links this board, but its public route was unavailable during the latest refresh; these are the last successfully captured rows.

Canonical page: https://benchlm.ai/benchmarks/scale-fortress

- Category: [Knowledge](/knowledge)
- Last updated: 2026-09-28 last available snapshot

## About FORTRESS

- Year: 2026
- Tasks: 68 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/fortress)

BenchLM mirrors 68 published rows from the FORTRESS public table captured on 2026-09-28 last available snapshot.

FORTRESS is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (68 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [DeepSeek-R1](/models/deepseek-r1) | DeepSeek | 74.39% |
| 2 | [glm-4p5-air](https://labs.scale.com/leaderboard/fortress) | Z.AI | 63.18% |
| 3 | [Gemini 2.5 Pro (06-05)](https://labs.scale.com/leaderboard/fortress) | Google | 61.68% |
| 4 | [Qwen3-235B-A22B](https://labs.scale.com/leaderboard/fortress) | Alibaba | 61.39% |
| 5 | [DeepSeek V3.1](/models/deepseek-v3-1) | DeepSeek | 60.55% |
| 6 | [glm-4p5](https://labs.scale.com/leaderboard/fortress) | Z.AI | 59.58% |
| 7 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 59.19% |
| 8 | [Qwen 2.5 72B](https://labs.scale.com/leaderboard/fortress) | Alibaba | 56.43% |
| 9 | [Mixtral 8x22B](https://labs.scale.com/leaderboard/fortress) | Mistral | 56.06% |
| 10 | [kimi-k2-instruct](https://labs.scale.com/leaderboard/fortress) | Moonshot AI | 55.47% |
| 11 | [Gemini 2.5 Pro (03-25)](https://labs.scale.com/leaderboard/fortress) | Google | 54.89% |
| 12 | [Gemini 1.5 Pro](/models/gemini-1-5-pro) | Google | 53.86% |
| 13 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 53.02% |
| 14 | [gemini-3.1-flash-lite-preview](https://labs.scale.com/leaderboard/fortress) | Google | 51.05% |
| 15 | [Gemini 1.5 Flash](https://labs.scale.com/leaderboard/fortress) | Google | 50.61% |
| 16 | [GPT-4o mini](/models/gpt-4o-mini) | OpenAI | 48.07% |
| 17 | [GPT-4o](/models/gpt-4o) | OpenAI | 47.18% |
| 18 | [Llama 3.3 70B](https://labs.scale.com/leaderboard/fortress) | Meta | 44.79% |
| 19 | [Llama 3.1 70B](https://labs.scale.com/leaderboard/fortress) | Meta | 44.18% |
| 20 | [gemini-3-pro-preview](https://labs.scale.com/leaderboard/fortress) | Google | 41.69% |
| 21 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 41.09% |
| 22 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 40.09% |
| 23 | [Claude 3.7 Sonnet](https://labs.scale.com/leaderboard/fortress) | Anthropic | 38.01% |
| 24 | [Claude 3.5 Haiku](https://labs.scale.com/leaderboard/fortress) | Anthropic | 30.41% |
| 25 | [gpt-5.1-instant](https://labs.scale.com/leaderboard/fortress) | OpenAI | 30.36% |
| 26 | [o3-mini](/models/o3-mini) | OpenAI | 30.05% |
| 27 | [gemini-3.1-pro-preview](https://labs.scale.com/leaderboard/fortress) | Google | 29.76% |
| 28 | [glm-5p3 max](https://labs.scale.com/leaderboard/fortress) | Z.AI | 28.23% |
| 29 | [Claude Opus 4](https://labs.scale.com/leaderboard/fortress) | Anthropic | 27.61% |
| 30 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 27.25% |
| 31 | [kimi-k3 max](https://labs.scale.com/leaderboard/fortress) | Moonshot AI | 26.65% |
| 32 | [nemotron-3-ultra-nvfp4 high](https://labs.scale.com/leaderboard/fortress) | Nvidia | 26.35% |
| 33 | [gpt-5.1-thinking](https://labs.scale.com/leaderboard/fortress) | OpenAI | 25.72% |
| 34 | [inkling xhigh](https://labs.scale.com/leaderboard/fortress) | Thinking Machines | 24.99% |
| 35 | [Claude 4 Opus (thinking)](https://labs.scale.com/leaderboard/fortress) | Anthropic | 24.76% |
| 36 | [Claude Sonnet 4](https://labs.scale.com/leaderboard/fortress) | Anthropic | 24.37% |
| 37 | [qwen3p8-2p4t-a95b xhigh](https://labs.scale.com/leaderboard/fortress) | Alibaba | 21.60% |
| 38 | [o4-mini](https://labs.scale.com/leaderboard/fortress) | OpenAI | 21.49% |
| 39 | [Llama 3.1 405B](/models/llama-3-1-405b) | Meta | 20.61% |
| 40 | [claude-opus-4-6 (Non-Thinking)](https://labs.scale.com/leaderboard/fortress) | Anthropic | 20.52% |
| 41 | [Muse Spark](/models/muse-spark) | Meta | 20.22% |
| 42 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 20.14% |
| 43 | [o1](/models/o1) | OpenAI | 19.38% |
| 44 | [Claude Opus 4.8 max](https://labs.scale.com/leaderboard/fortress) | Anthropic | 18.19% |
| 45 | [Claude 4 Sonnet (thinking)](https://labs.scale.com/leaderboard/fortress) | Anthropic | 18.06% |
| 46 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 17.62% |
| 47 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 17.50% |
| 48 | [gpt-5-2025-08-07](https://labs.scale.com/leaderboard/fortress) | OpenAI | 17.04% |
| 49 | [GPT-5 mini](/models/gpt-5-mini) | OpenAI | 17.00% |
| 50 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 16.96% |
| 51 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 16.84% |
| 52 | [gpt-5.5-xhigh](https://labs.scale.com/leaderboard/fortress) | OpenAI | 16.28% |
| 53 | [claude-opus-4-1-20250805](https://labs.scale.com/leaderboard/fortress) | Anthropic | 16.15% |
| 54 | [o3](/models/o3) | OpenAI | 16.01% |
| 55 | [gpt-5-pro-2025-10-06](https://labs.scale.com/leaderboard/fortress) | OpenAI | 15.20% |
| 56 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 14.84% |
| 57 | [claude-opus-4-1-20250805-thinking](https://labs.scale.com/leaderboard/fortress) | Anthropic | 14.79% |
| 58 | [Fable-5.1](https://labs.scale.com/leaderboard/fortress) | Anthropic | 13.72% |
| 59 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 13.56% |
| 60 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 13.46% |
| 61 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 13.18% |
| 62 | [claude-opus-4-6-thinking-max](https://labs.scale.com/leaderboard/fortress) | Anthropic | 13.00% |
| 63 | [Claude 3.5 Sonnet](/models/claude-3-5-sonnet) | Anthropic | 12.96% |
| 64 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | Anthropic | 12.80% |
| 65 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 12.40% |
| 66 | [GPT-6 Luna](/models/gpt-6-luna) | OpenAI | 10.78% |
| 67 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 9.63% |
| 68 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 8.24% |

## FAQ

### What does FORTRESS measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings. The catalog still links this board, but its public route was unavailable during the latest refresh; these are the last successfully captured rows.

### Which model leads the published FORTRESS snapshot?

DeepSeek-R1 currently leads the published FORTRESS snapshot with a score of 74.39%.

### How many models are evaluated on FORTRESS?

The 2026-09-28 last available snapshot contains 68 AI models.

### Does FORTRESS affect BenchLM's overall score?

Not directly. FORTRESS is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
