Skip to main content

Benchmark profile

Software Engineering Benchmark Verified (SWE-bench Verified)

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Data verified

Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's July 2026 update with 96%, ahead of Claude Mythos 5 (95.5%) and Claude Fable 5 (95%), across 59 tracked models.

Top models on SWE-bench Verified — July 29, 2026

As of July 29, 2026, Claude Opus 5 leads the SWE-bench Verified leaderboard with 96% , followed by Claude Mythos 5 (95.5%) and Claude Fable 5 (95%).

59 modelsCoding16% of category scoreRefreshingUpdated July 29, 2026

Leaderboard (59 models)

Score
1
Claude Opus 5Anthropic · Closed
96%
2
Claude Mythos 5Anthropic · Closed
95.5%
3
Claude Fable 5Anthropic · Closed
95%
4
Claude Opus 4.8Anthropic · Closed
88.6%
5
Claude Opus 4.7 (Adaptive)Anthropic · Closed
87.6%
6
Claude Sonnet 5Anthropic · Closed
85.2%
7
GPT-5.3 CodexOpenAI · Closed
85%
8
Ornith-1.0-397BDeepReinforce AI · Open weight
82.4%
9
Claude Opus 4.5Anthropic · Closed
80.9%
10
Claude Opus 4.6Anthropic · Closed
80.8%
11
DeepSeek V4 Pro (Max)DeepSeek · Open weight
80.6%
12
MiniMax M3MiniMax · Open weight
80.5%
13
Qwen3.7 MaxAlibaba · Closed
80.4%
14
Kimi K2.6Moonshot AI · Open weight
80.2%
15
GPT-5.2OpenAI · Closed
80%
16
Claude Sonnet 4.6Anthropic · Closed
79.6%
17
DeepSeek V4 Pro (High)DeepSeek · Open weight
79.4%
18
DeepSeek V4 Flash (Max)DeepSeek · Open weight
79%
19
Qwen3.6 PlusAlibaba · Closed
78.8%
20
DeepSeek V4 Flash (High)DeepSeek · Open weight
78.6%
21
MiMo-V2-ProXiaomi · Closed
78%
22
GLM-5Z.AI · Open weight
77.8%
23
Qwen3.7 PlusAlibaba · Closed
77.7%
24
InklingThinking Machines Lab · Open weight
77.6%
25
Mistral Medium 3.5 128BMistral · Open weight
77.6%
26
Muse SparkMeta · Closed
77.4%
27
Qwen3.6-27BAlibaba · Open weight
77.2%
28
Claude Sonnet 4.5Anthropic · Closed
77.2%
29
Kimi K2.5Moonshot AI · Open weight
76.8%
30
Kimi K2.5 (Reasoning)Moonshot AI · Closed
76.8%
31
Grok 4.20xAI · Closed
76.7%
32
Qwen3.5 397BAlibaba · Open weight
76.2%
33
Ornith-1.0-35BDeepReinforce AI · Open weight
75.6%
34
MiMo-V2-OmniXiaomi · Closed
74.8%
35
Laguna M.1Poolside · Closed
74.6%
36
Claude 4.1 OpusAnthropic · Closed
74.5%
37
Hy3 PreviewTencent · Open weight
74.4%
38
GLM-4.7Z.AI · Open weight
73.8%
39
DeepSeek V4 FlashDeepSeek · Open weight
73.7%
40
DeepSeek V4 ProDeepSeek · Open weight
73.6%
41
MAI-Thinking-1Microsoft · Closed
73.5%
42
Qwen3.6-35B-A3BAlibaba · Open weight
73.4%
43
MiMo-V2-FlashXiaomi · Open weight
73.4%
44
Claude Haiku 4.5Anthropic · Closed
73.3%
45
Claude 4 SonnetAnthropic · Closed
72.7%
46
Qwen3.5-27BAlibaba · Open weight
72.4%
47
Qwen3.5-122B-A10BAlibaba · Open weight
72%
48
Nemotron 3 UltraNVIDIA · Open weight
71.9%
49
Grok Code Fast 1xAI · Closed
70.8%
50
Laguna XS.2Poolside · Open weight
69.9%
51
Ornith-1.0-9BDeepReinforce AI · Open weight
69.4%
52
Qwen3.5-35B-A3BAlibaba · Open weight
69.2%
53
Gemini 2.5 ProGoogle · Closed
63.8%
54
GPT-4.1OpenAI · Closed
54.6%
55
ZAYA1-74B-PreviewZyphra · Open weight
53.2%
56
o3-miniOpenAI · Closed
49.3%
57
Claude 3.5 SonnetAnthropic · Closed
49%
58
DeepSeek V3DeepSeek · Open weight
42%
59
GPT-4.1 miniOpenAI · Closed
23.6%

According to BenchLM.ai, Claude Opus 5 leads the SWE-bench Verified benchmark with a score of 96%, followed by Claude Mythos 5 (95.5%) and Claude Fable 5 (95%). The top models are clustered within 1.0 points, suggesting this benchmark is nearing saturation for frontier models.

59 models have been evaluated on SWE-bench Verified. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Verified contributes 16% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE-bench Verified

Year

2024

Tasks

500 verified issues

Format

Code patch generation

Difficulty

Professional software engineering

SWE-bench Verified is the most-cited benchmark for AI coding agents on real repository tasks. Each task requires understanding codebases, writing patches, and passing test suites.

BenchLM freshness & provenance

Version

SWE-bench Verified 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-bench Verified measure?

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Which model scores highest on SWE-bench Verified?

Claude Opus 5 by Anthropic currently leads with a score of 96% on SWE-bench Verified.

How many models are evaluated on SWE-bench Verified?

59 AI models have been evaluated on SWE-bench Verified on BenchLM.

Last updated: July 29, 2026 · BenchLM version SWE-bench Verified 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.