# Scale Labs SWE-Bench Pro V2 Full (SWE-Bench Pro V2 Full)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-swe-bench-pro-public-v2-full

- Category: [Coding](/coding)
- Last updated: September 23, 2026 snapshot

## About SWE-Bench Pro V2 Full

- Year: 2026
- Tasks: 10 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2)

BenchLM mirrors 10 published rows from the SWE-Bench Pro V2 Full public table captured on September 23, 2026 snapshot.

SWE-Bench Pro V2 Full is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (10 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5 (Claude Code) xhigh](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Anthropic | 99.40% |
| 2 | [Fable 5.1 (Claude Code) high](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Anthr | 99.10% |
| 3 | [Kimi-K3 (mini-swe-agent) max](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Moonshot AI | 97.70% |
| 4 | [GPT - 6- Astra (Codex) high](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Unknown | 96.90% |
| 5 | [GLM-5.3 (mini-swe-agent) max](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Z.AI | 95.60% |
| 6 | [GPT-5.6-sol (codex) xhigh](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | OpenAI | 95.50% |
| 7 | [Gemini 3.8 Flash (mini-swe-agent) high](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Google | 94.86% |
| 8 | [Claude Sonnet 5 (Claude Code) xhigh](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Anthropic | 93.15% |
| 9 | [GPT-5.6-Terra (codex) xhigh](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | OpenAI | 92.37% |
| 10 | [Inkling (mini-swe-agent) xhigh](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2) | SWE-Bench Pro V2 Full | Thinkingmachines | 89.88% |

## FAQ

### What does SWE-Bench Pro V2 Full measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published SWE-Bench Pro V2 Full snapshot?

Claude Opus 5 (Claude Code) xhigh currently leads the published SWE-Bench Pro V2 Full snapshot with a score of 99.40%.

### How many models are evaluated on SWE-Bench Pro V2 Full?

The September 23, 2026 snapshot contains 10 AI models.

### Does SWE-Bench Pro V2 Full affect BenchLM's overall score?

Not directly. SWE-Bench Pro V2 Full is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
