Skip to main content

Benchmark profile

Software Engineering Benchmark Verified (SWE-bench Verified)

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Data verified

Claude Mythos 5 leads the SWE-bench Verified leaderboard on BenchLM's July 2026 update with 95.5%, ahead of Claude Fable 5 (95%) and Claude Opus 4.8 (88.6%), across 58 tracked models.

Top models on SWE-bench Verified — July 20, 2026

As of July 20, 2026, Claude Mythos 5 leads the SWE-bench Verified leaderboard with 95.5% , followed by Claude Fable 5 (95%) and Claude Opus 4.8 (88.6%).

58 modelsCoding16% of category scoreRefreshingUpdated July 20, 2026

Leaderboard (58 models)

Score
1
Claude Mythos 5Anthropic · Closed
95.5%
2
Claude Fable 5Anthropic · Closed
95%
3
Claude Opus 4.8Anthropic · Closed
88.6%
4
Claude Opus 4.7 (Adaptive)Anthropic · Closed
87.6%
5
Claude Sonnet 5Anthropic · Closed
85.2%
6
GPT-5.3 CodexOpenAI · Closed
85%
7
Ornith-1.0-397BDeepReinforce AI · Open weight
82.4%
8
Claude Opus 4.5Anthropic · Closed
80.9%
9
Claude Opus 4.6Anthropic · Closed
80.8%
10
DeepSeek V4 Pro (Max)DeepSeek · Open weight
80.6%
11
MiniMax M3MiniMax · Open weight
80.5%
12
Qwen3.7 MaxAlibaba · Closed
80.4%
13
Kimi K2.6Moonshot AI · Open weight
80.2%
14
GPT-5.2OpenAI · Closed
80%
15
Claude Sonnet 4.6Anthropic · Closed
79.6%
16
DeepSeek V4 Pro (High)DeepSeek · Open weight
79.4%
17
DeepSeek V4 Flash (Max)DeepSeek · Open weight
79%
18
Qwen3.6 PlusAlibaba · Closed
78.8%
19
DeepSeek V4 Flash (High)DeepSeek · Open weight
78.6%
20
MiMo-V2-ProXiaomi · Closed
78%
21
GLM-5Z.AI · Open weight
77.8%
22
Qwen3.7 PlusAlibaba · Closed
77.7%
23
InklingThinking Machines Lab · Open weight
77.6%
24
Mistral Medium 3.5 128BMistral · Open weight
77.6%
25
Muse SparkMeta · Closed
77.4%
26
Qwen3.6-27BAlibaba · Open weight
77.2%
27
Claude Sonnet 4.5Anthropic · Closed
77.2%
28
Kimi K2.5Moonshot AI · Open weight
76.8%
29
Kimi K2.5 (Reasoning)Moonshot AI · Closed
76.8%
30
Grok 4.20xAI · Closed
76.7%
31
Qwen3.5 397BAlibaba · Open weight
76.2%
32
Ornith-1.0-35BDeepReinforce AI · Open weight
75.6%
33
MiMo-V2-OmniXiaomi · Closed
74.8%
34
Laguna M.1Poolside · Closed
74.6%
35
Claude 4.1 OpusAnthropic · Closed
74.5%
36
Hy3 PreviewTencent · Open weight
74.4%
37
GLM-4.7Z.AI · Open weight
73.8%
38
DeepSeek V4 FlashDeepSeek · Open weight
73.7%
39
DeepSeek V4 ProDeepSeek · Open weight
73.6%
40
MAI-Thinking-1Microsoft · Closed
73.5%
41
Qwen3.6-35B-A3BAlibaba · Open weight
73.4%
42
MiMo-V2-FlashXiaomi · Open weight
73.4%
43
Claude Haiku 4.5Anthropic · Closed
73.3%
44
Claude 4 SonnetAnthropic · Closed
72.7%
45
Qwen3.5-27BAlibaba · Open weight
72.4%
46
Qwen3.5-122B-A10BAlibaba · Open weight
72%
47
Nemotron 3 UltraNVIDIA · Open weight
71.9%
48
Grok Code Fast 1xAI · Closed
70.8%
49
Laguna XS.2Poolside · Open weight
69.9%
50
Ornith-1.0-9BDeepReinforce AI · Open weight
69.4%
51
Qwen3.5-35B-A3BAlibaba · Open weight
69.2%
52
Gemini 2.5 ProGoogle · Closed
63.8%
53
GPT-4.1OpenAI · Closed
54.6%
54
ZAYA1-74B-PreviewZyphra · Open weight
53.2%
55
o3-miniOpenAI · Closed
49.3%
56
Claude 3.5 SonnetAnthropic · Closed
49%
57
DeepSeek V3DeepSeek · Open weight
42%
58
GPT-4.1 miniOpenAI · Closed
23.6%

According to BenchLM.ai, Claude Mythos 5 leads the SWE-bench Verified benchmark with a score of 95.5%, followed by Claude Fable 5 (95%) and Claude Opus 4.8 (88.6%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

58 models have been evaluated on SWE-bench Verified. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Verified contributes 16% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE-bench Verified

Year

2024

Tasks

500 verified issues

Format

Code patch generation

Difficulty

Professional software engineering

SWE-bench Verified is the gold standard for evaluating AI coding agents on real-world software engineering tasks. Each task requires understanding codebases, writing patches, and passing test suites.

BenchLM freshness & provenance

Version

SWE-bench Verified 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-bench Verified measure?

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Which model scores highest on SWE-bench Verified?

Claude Mythos 5 by Anthropic currently leads with a score of 95.5% on SWE-bench Verified.

How many models are evaluated on SWE-bench Verified?

58 AI models have been evaluated on SWE-bench Verified on BenchLM.

Last updated: July 20, 2026 · BenchLM version SWE-bench Verified 2024

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.