Skip to main content
BenchLM

Scale Labs SWE-Bench Pro V2 Full (SWE-Bench Pro V2 Full)

We show this table for reference; we do not rank on it.

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Scale score on SWE-Bench Pro V2 Full — September 23, 2026 snapshot

We mirror the published scale score view for SWE-Bench Pro V2 Full. Claude Opus 5 (Claude Code) xhigh leads the public snapshot at 99.4%, followed by Fable 5.1 (Claude Code) high (99.1%) and Kimi-K3 (mini-swe-agent) max (97.7%). We do not use these results to rank models overall.

10 modelsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot

Scale score table (10 models)

Score
1
Claude Opus 5 (Claude Code) xhighAnthropicSWE-Bench Pro V2 Full
99.4%
2
Fable 5.1 (Claude Code) highAnthrSWE-Bench Pro V2 Full
99.1%
3
Kimi-K3 (mini-swe-agent) maxMoonshot AISWE-Bench Pro V2 Full
97.7%
4
GPT - 6- Astra (Codex) highUnknownSWE-Bench Pro V2 Full
96.9%
5
GLM-5.3 (mini-swe-agent) maxZ.AISWE-Bench Pro V2 Full
95.6%
6
GPT-5.6-sol (codex) xhighOpenAISWE-Bench Pro V2 Full
95.5%
7
Gemini 3.8 Flash (mini-swe-agent) highGoogleSWE-Bench Pro V2 Full
94.9%
8
Claude Sonnet 5 (Claude Code) xhighAnthropicSWE-Bench Pro V2 Full
93.2%
9
GPT-5.6-Terra (codex) xhighOpenAISWE-Bench Pro V2 Full
92.4%
10
Inkling (mini-swe-agent) xhighThinkingmachinesSWE-Bench Pro V2 Full
89.9%

How to read this leaderboard

Compare the published configurations as complete evaluation systems. The source can combine a base model, agent scaffold, tools, budget, and inference setting in each result.

Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 (Claude Code) xhigh at 99.4%.

Honest limit: This Scale table is display-only context, not benchmark provenance or a weighted model-only comparison.

How BenchLM shows SWE-Bench Pro V2 Full

BenchLM mirrors 10 published rows from Scale Labs’ public SWE-Bench Pro V2 Full leaderboard, captured on September 23, 2026 snapshot.

The table is display only. It is useful context for a published agent or model configuration, but it does not enter BenchLM’s overall or category rankings.

Snapshot

10 published rowsScale Labs sourceSWE-Bench Pro V2 FullDisplay only

The published SWE-Bench Pro V2 Full snapshot places Claude Opus 5 (Claude Code) xhigh first at 99.4%. The third row is 1.7 points behind. The broader top-10 range is 9.5 points, so many of the published results sit in a relatively narrow band.

10 models have been evaluated on SWE-Bench Pro V2 Full. The benchmark falls in the Coding category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. SWE-Bench Pro V2 Full is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE-Bench Pro V2 Full

Year

2026

Tasks

10 published rows

Format

Published Scale leaderboard score

Difficulty

External agent and model evaluation

BenchLM mirrors 10 published rows from the SWE-Bench Pro V2 Full public table captured on September 23, 2026 snapshot.

Freshness and provenance

Version

SWE-Bench Pro V2 Full 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SWE-Bench Pro V2 Full measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Which model leads the published SWE-Bench Pro V2 Full snapshot?

Claude Opus 5 (Claude Code) xhigh currently leads the published SWE-Bench Pro V2 Full snapshot with 99.4% scale score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SWE-Bench Pro V2 Full?

The September 23, 2026 snapshot snapshot contains 10 AI models.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.