Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

Data verified 37 confirmed releases in the last 30 daysSee provider release alerts

Claude Fable 5.1 leads the SWE-bench Pro leaderboard on BenchLM's September 2026 update with 81.2%, ahead of Claude Mythos 5 (80.3%) and Claude Fable 5 (80%), across 72 models.

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use SWE-bench Pro as evidence for long-horizon repository work only after matching the split, scaffold, tool budget, retry policy, and run count. The table preserves exact published rows from different providers, so a small score gap is directional unless the setups match.

Operator receipt: 72 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5.1 at 81.2%.

Honest limit: OpenAI's July 2026 audit estimated that about 30% of the 731-task public split is broken and retracted its earlier recommendation to adopt the benchmark. The current rows remain useful published receipts, but this page should not decide a coding-agent purchase by itself.

Top models on SWE-bench Pro — September 10, 2026

As of September 10, 2026, Claude Fable 5.1 leads the SWE-bench Pro leaderboard with 81.2% , followed by Claude Mythos 5 (80.3%) and Claude Fable 5 (80%).

72 modelsCoding25% of category scoreCurrentUpdated September 10, 2026

Leaderboard (72 models)

Score
1
Claude Fable 5.1Anthropic · Closed
81.2%
2
Claude Mythos 5Anthropic · Closed
80.3%
3
Claude Fable 5Anthropic · Closed
80%
4
Claude Opus 5Anthropic · Closed
79.2%
5
Sakana Fugu-UltraSakana AI · Closed
73.7%
6
Claude Opus 4.8Anthropic · Closed
69.2%
7
Qwen3.8 MaxAlibaba · Open weight
67.7%
8
Hy4 previewTencent · Open weight
65.7%
9
Ornith-1.5-397BOrnith AI · Open weight
65.1%
10
Grok 4.5xAI · Closed
64.7%
11
GPT-5.6 SolOpenAI · Closed
64.6%
12
Claude Opus 4.7 (Adaptive)Anthropic · Closed
64.3%
13
GPT-5.6 TerraOpenAI · Closed
63.4%
14
Claude Sonnet 5Anthropic · Closed
63.2%
15
GPT-5.6 LunaOpenAI · Closed
62.7%
16
Qwen3.8-Flash-NextAlibaba · Open weight
62.5%
17
Ornith-1.0-397BDeepReinforce AI · Open weight
62.2%
18
GLM-5.2Z.AI · Open weight
62.1%
19
Qwen3.8-27BAlibaba · Open weight
61.7%
20
Muse Spark 1.1Meta · Closed
61.5%
21
dots3-note PreviewDots Studio · Open weight
61%
22
Qwen3.7 MaxAlibaba · Closed
60.6%
23
Ornith-1.5-35B-A3BOrnith AI · Open weight
59.6%
24
Laguna S 2.1Poolside · Open weight
59.4%
25
MiniMax M3MiniMax · Open weight
59%
26
Sakana FuguSakana AI · Closed
59%
27
GPT-5.5OpenAI · Closed
58.6%
28
Kimi K2.6Moonshot AI · Open weight
58.6%
29
GLM-5.1Z.AI · Open weight
58.4%
30
GPT-5.4OpenAI · Closed
57.7%
31
Qwen3.7 PlusAlibaba · Closed
57.6%
32
Qwen 3.6 Max (preview)Alibaba · Closed
57.3%
33
MiMo-V2.5-ProXiaomi · Closed
57.2%
34
Claude Opus 4.5Anthropic · Closed
57.1%
35
GPT-5.3 CodexOpenAI · Closed
56.8%
36
Qwen3.6 PlusAlibaba · Closed
56.6%
37
Ling 3.0 FlashInclusionAI · Open weight
56.6%
38
Step 3.7 FlashStepFun · Open weight
56.3%
39
MiniMax M2.7MiniMax · Open weight
56.2%
40
MiMo-V2.5Xiaomi · Closed
56.1%
41
Inkling-SmallThinking Machines Lab · Open weight
55.9%
42
GPT-5.2OpenAI · Closed
55.6%
43
DeepSeek V4 Pro 0813DeepSeek · Closed
55.4%
44
Gemini 3.5 FlashGoogle · Closed
55.1%
45
GLM-5Z.AI · Open weight
55.1%
46
DeepSeek V4 Pro (High)DeepSeek · Open weight
54.4%
47
InklingThinking Machines Lab · Open weight
54.3%
48
Gemini 3.5 Flash-LiteGoogle · Closed
54.2%
49
Qwen3.6-27BAlibaba · Open weight
53.5%
50
Claude Opus 4.6Anthropic · Closed
53.4%
51
MAI-Thinking-1Microsoft · Closed
52.8%
52
DeepSeek V4 Flash 0731DeepSeek · Closed
52.6%
53
Muse SparkMeta · Closed
52.4%
54
DeepSeek V4 Flash (High)DeepSeek · Closed
52.3%
55
DeepSeek V4 ProDeepSeek · Open weight
52.1%
56
Grok 4.20xAI · Closed
51.8%
57
Muse Glimmer 30BMeta · Open weight
51.2%
58
Qwen3.5 397BAlibaba · Open weight
50.9%
59
Kimi K2.5Moonshot AI · Open weight
50.7%
60
Ornith-1.0-35BDeepReinforce AI · Open weight
50.4%
61
Qwen3.6-35B-A3BAlibaba · Open weight
49.5%
62
Laguna M.1Poolside · Closed
49.2%
63
DeepSeek V4 FlashDeepSeek · Closed
49.1%
64
Laguna XS 2.1Poolside · Open weight
47.6%
65
Ornith-1.5-9BOrnith AI · Open weight
47.5%
66
Laguna XS.2Poolside · Open weight
46.3%
67
Ornith-1.0-9BDeepReinforce AI · Open weight
42.9%
68
LongCat-Flash-Lite-SparseMeituan · Open weight
40.6%
69
Granite 4.2 30BIBM · Open weight
33.3%
70
LLaDA2.2-flashInclusionAI · Open weight
30.1%
71
Granite 4.2 8BIBM · Open weight
19.1%
72
MiniCPM5-2BOpenBMB · Open weight
14.4%

According to BenchLM.ai, Claude Fable 5.1 leads the SWE-bench Pro benchmark with a score of 81.2%, followed by Claude Mythos 5 (80.3%) and Claude Fable 5 (80%). The top models are clustered within 1.2 points, suggesting this benchmark is nearing saturation for frontier models.

72 models have been evaluated on SWE-bench Pro. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Pro contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE-bench Pro

Year

2025

Tasks

1,865 repository problems

Format

Repository task completion

Difficulty

Long-horizon professional engineering

The authors assembled 1,865 problems from 41 repositories across public, held-out, and commercial splits. Agents receive a repository and issue, then produce a patch that must pass the evaluation tests without breaking existing behavior.

BenchLM freshness & provenance

Version

SWE-bench Pro 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-bench Pro measure?

SWE-bench Pro gives an agent a repository and issue description, then checks whether its patch passes new tests without breaking existing behavior. The full benchmark contains 1,865 problems from 41 repositories across public, held-out, and commercial splits, with work that can span multiple files and long execution horizons.

Are SWE-bench Pro scores directly comparable?

Only when the split and evaluation setup match. Public, held-out, and commercial tasks are different pools, while scaffold, tool budget, retry policy, and token budget can change pass rates. This page preserves exact published rows, but it does not pretend every provider ran the same harness.

Should SWE-bench Pro decide which coding agent to use?

No. The benchmark covers realistic repository work, but OpenAI's July 2026 audit estimated that about 30% of the public tasks are broken and retracted its earlier adoption recommendation. Use SWE-bench Pro alongside LiveCodeBench, other repository evaluations, and a workload-specific trial instead of treating one score as a procurement decision.

Last updated: September 10, 2026 · BenchLM version SWE-bench Pro 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.