Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Software Engineering Benchmark Verified (SWE-bench Verified)

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Data verified 37 confirmed releases in the last 30 daysSee provider release alerts

Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's September 2026 update with 96%, ahead of Claude Mythos 5 (95.5%) and Claude Fable 5 (95%), across 78 models.

Top models on SWE-bench Verified — September 10, 2026

As of September 10, 2026, Claude Opus 5 leads the SWE-bench Verified leaderboard with 96% , followed by Claude Mythos 5 (95.5%) and Claude Fable 5 (95%).

78 modelsCoding10% of category scoreRefreshingUpdated September 10, 2026

Leaderboard (78 models)

Score
1
Claude Opus 5Anthropic · Closed
96%
2
Claude Mythos 5Anthropic · Closed
95.5%
3
Claude Fable 5Anthropic · Closed
95%
4
Claude Opus 4.8Anthropic · Closed
88.6%
5
Claude Opus 4.7 (Adaptive)Anthropic · Closed
87.6%
6
Ornith-1.5-397BOrnith AI · Open weight
86%
7
Claude Sonnet 5Anthropic · Closed
85.2%
8
GPT-5.3 CodexOpenAI · Closed
85%
9
Ornith-1.0-397BDeepReinforce AI · Open weight
82.4%
10
Claude Opus 4.5Anthropic · Closed
80.9%
11
Claude Opus 4.6Anthropic · Closed
80.8%
12
DeepSeek V4 Pro 0813DeepSeek · Closed
80.6%
13
MiniMax M3MiniMax · Open weight
80.5%
14
Qwen3.7 MaxAlibaba · Closed
80.4%
15
Kimi K2.6Moonshot AI · Open weight
80.2%
16
Inkling-SmallThinking Machines Lab · Open weight
80.2%
17
GPT-5.2OpenAI · Closed
80%
18
Claude Sonnet 4.6Anthropic · Closed
79.6%
19
DeepSeek V4 Pro (High)DeepSeek · Open weight
79.4%
20
DeepSeek V4 Flash 0731DeepSeek · Closed
79%
21
Ornith-1.5-35B-A3BOrnith AI · Open weight
79%
22
Qwen3.6 PlusAlibaba · Closed
78.8%
23
DeepSeek V4 Flash (High)DeepSeek · Closed
78.6%
24
dots3-note PreviewDots Studio · Open weight
78.4%
25
BTL-4Bad Theory Labs · Open weight
78.4%
26
MiMo-V2-ProXiaomi · Closed
78%
27
GLM-5Z.AI · Open weight
77.8%
28
Qwen3.7 PlusAlibaba · Closed
77.7%
29
Apodex 1.1Apodex · Closed
77.7%
30
InklingThinking Machines Lab · Open weight
77.6%
31
Mistral Medium 3.5 128BMistral · Open weight
77.6%
32
Muse SparkMeta · Closed
77.4%
33
Qwen3.6-27BAlibaba · Open weight
77.2%
34
Claude Sonnet 4.5Anthropic · Closed
77.2%
35
Kimi K2.5Moonshot AI · Open weight
76.8%
36
Kimi K2.5 (Reasoning)Moonshot AI · Closed
76.8%
37
Grok 4.20xAI · Closed
76.7%
38
Qwen3.5 397BAlibaba · Open weight
76.2%
39
Muse Glimmer 30BMeta · Open weight
76%
40
Ornith-1.0-35BDeepReinforce AI · Open weight
75.6%
41
MiMo-V2-OmniXiaomi · Closed
74.8%
42
Laguna M.1Poolside · Closed
74.6%
43
Claude 4.1 OpusAnthropic · Closed
74.5%
44
Hy3 PreviewTencent · Open weight
74.4%
45
GLM-4.7Z.AI · Open weight
73.8%
46
DeepSeek V4 FlashDeepSeek · Closed
73.7%
47
DeepSeek V4 ProDeepSeek · Open weight
73.6%
48
MAI-Thinking-1Microsoft · Closed
73.5%
49
Qwen3.6-35B-A3BAlibaba · Open weight
73.4%
50
MiMo-V2-FlashXiaomi · Open weight
73.4%
51
Claude Haiku 4.5Anthropic · Closed
73.3%
52
Claude 4 SonnetAnthropic · Closed
72.7%
53
MAI-Code-1.1-FlashMicrosoft · Closed
72.6%
54
Qwen3.5-27BAlibaba · Open weight
72.4%
55
Qwen3.5-122B-A10BAlibaba · Open weight
72%
56
Nemotron 3 UltraNVIDIA · Open weight
71.9%
57
Laguna XS 2.1Poolside · Open weight
70.9%
58
Grok Code Fast 1xAI · Closed
70.8%
59
Ornith-1.5-9BOrnith AI · Open weight
70.6%
60
Solar Pro 4Upstage · Closed
70.6%
61
Solar Open 2Upstage · Open weight
70.4%
62
Laguna XS.2Poolside · Open weight
69.9%
63
Ornith-1.0-9BDeepReinforce AI · Open weight
69.4%
64
Qwen3.5-35B-A3BAlibaba · Open weight
69.2%
65
LongCat-Flash-Lite-SparseMeituan · Open weight
68.2%
66
K-EXAONE 2.0LG AI Research · Open weight
68.2%
67
Gemini 2.5 ProGoogle · Closed
63.8%
68
Granite 4.2 30BIBM · Open weight
57%
69
GPT-4.1OpenAI · Closed
54.6%
70
ZAYA1-74B-PreviewZyphra · Open weight
53.2%
71
52.8%
72
o3-miniOpenAI · Closed
49.3%
73
LLaDA2.2-flashInclusionAI · Open weight
49.3%
74
Claude 3.5 SonnetAnthropic · Closed
49%
75
Granite 4.2 8BIBM · Open weight
47.7%
76
MiniCPM5-2BOpenBMB · Open weight
46.4%
77
DeepSeek V3DeepSeek · Open weight
42%
78
GPT-4.1 miniOpenAI · Closed
23.6%

According to BenchLM.ai, Claude Opus 5 leads the SWE-bench Verified benchmark with a score of 96%, followed by Claude Mythos 5 (95.5%) and Claude Fable 5 (95%). The top models are clustered within 1.0 points, suggesting this benchmark is nearing saturation for frontier models.

78 models have been evaluated on SWE-bench Verified. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Verified contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE-bench Verified

Year

2024

Tasks

500 verified issues

Format

Code patch generation

Difficulty

Professional software engineering

SWE-bench Verified is the most-cited benchmark for AI coding agents on real repository tasks. Each task requires understanding codebases, writing patches, and passing test suites.

BenchLM freshness & provenance

Version

SWE-bench Verified 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-bench Verified measure?

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Which model scores highest on SWE-bench Verified?

Claude Opus 5 by Anthropic currently leads with a score of 96% on SWE-bench Verified.

How many models are evaluated on SWE-bench Verified?

78 AI models have been evaluated on SWE-bench Verified on BenchLM.

Last updated: September 10, 2026 · BenchLM version SWE-bench Verified 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.