Skip to main content
BenchLM

Scale Labs SWE-Bench Pro V2 HARD (SWE-Bench Pro V2 HARD)

We show this table for reference; we do not rank on it.

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Scale score on SWE-Bench Pro V2 HARD — September 23, 2026 snapshot

We mirror the published scale score view for SWE-Bench Pro V2 HARD. Opus 5 (Claude Code) xhigh leads the public snapshot at 98%, followed by Fable 5.1 (Claude Code) high (92.2%) and GPT-6 Astra (Codex) high (90.2%). We do not use these results to rank models overall.

11 modelsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot

Scale score table (11 models)

Score
1
Opus 5 (Claude Code) xhighUnknownSWE-Bench Pro V2 HARD
98%
2
Fable 5.1 (Claude Code) highUnknownSWE-Bench Pro V2 HARD
92.2%
3
GPT-6 Astra (Codex) highUnknownSWE-Bench Pro V2 HARD
90.2%
4
Sonnet 5 (Claude Code) xhighAnthropicSWE-Bench Pro V2 HARD
88.2%
5
Kimi-K3 (mini-swe-agent) maxMoonshot AISWE-Bench Pro V2 HARD
88.2%
6
GPT-5.6 Terra (Codex) xhighOpenAISWE-Bench Pro V2 HARD
86.3%
7
GLM-5.3 (mini-swe-agent) maxZ.AISWE-Bench Pro V2 HARD
84.3%
8
GPT-5.6 Sol (Codex) xhighOpenAISWE-Bench Pro V2 HARD
82.4%
9
Gemini 3.8 Flash (mini-swe-agent) highGoogleSWE-Bench Pro V2 HARD
58.8%
10
Inkling (mini-swe-agent) xhighThinkingmachinesSWE-Bench Pro V2 HARD
56.9%
11
Haiku 4.5 (Claude Code) xhighAnthropicSWE-Bench Pro V2 HARD
25.5%

How to read this leaderboard

Compare the published configurations as complete evaluation systems. The source can combine a base model, agent scaffold, tools, budget, and inference setting in each result.

Operator receipt: 11 sourced rows are currently displayable on this page; the leading published row is Opus 5 (Claude Code) xhigh at 98%.

Honest limit: This Scale table is display-only context, not benchmark provenance or a weighted model-only comparison.

How BenchLM shows SWE-Bench Pro V2 HARD

BenchLM mirrors 11 published rows from Scale Labs’ public SWE-Bench Pro V2 HARD leaderboard, captured on September 23, 2026 snapshot.

The table is display only. It is useful context for a published agent or model configuration, but it does not enter BenchLM’s overall or category rankings.

Snapshot

11 published rowsScale Labs sourceSWE-Bench Pro V2 HARDDisplay only

The published SWE-Bench Pro V2 HARD snapshot places Opus 5 (Claude Code) xhigh first at 98%. The third row is 7.8 points behind. The broader top-10 range is 41.1 points, so the table still separates the published systems.

11 models have been evaluated on SWE-Bench Pro V2 HARD. The benchmark falls in the Coding category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. SWE-Bench Pro V2 HARD is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE-Bench Pro V2 HARD

Year

2026

Tasks

11 published rows

Format

Published Scale leaderboard score

Difficulty

External agent and model evaluation

BenchLM mirrors 11 published rows from the SWE-Bench Pro V2 HARD public table captured on September 23, 2026 snapshot.

Freshness and provenance

Version

SWE-Bench Pro V2 HARD 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SWE-Bench Pro V2 HARD measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Which model leads the published SWE-Bench Pro V2 HARD snapshot?

Opus 5 (Claude Code) xhigh currently leads the published SWE-Bench Pro V2 HARD snapshot with 98% scale score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SWE-Bench Pro V2 HARD?

The September 23, 2026 snapshot snapshot contains 11 AI models.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.