Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Software Engineering Benchmark Verified (SWE-bench Verified)

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Data verified 25 confirmed releases in the last 30 daysStart free brief

Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's August 2026 update with 96%, ahead of Claude Mythos 5 (95.5%) and Claude Fable 5 (95%), across 67 tracked models.

Top models on SWE-bench Verified — August 21, 2026

As of August 21, 2026, Claude Opus 5 leads the SWE-bench Verified leaderboard with 96% , followed by Claude Mythos 5 (95.5%) and Claude Fable 5 (95%).

67 modelsCoding16% of category scoreRefreshingUpdated August 21, 2026

Leaderboard (67 models)

Score
1
Claude Opus 5Anthropic · Closed
96%
2
Claude Mythos 5Anthropic · Closed
95.5%
3
Claude Fable 5Anthropic · Closed
95%
4
Claude Opus 4.8Anthropic · Closed
88.6%
5
Claude Opus 4.7 (Adaptive)Anthropic · Closed
87.6%
6
Ornith-1.5-397BOrnith AI · Open weight
86%
7
Claude Sonnet 5Anthropic · Closed
85.2%
8
GPT-5.3 CodexOpenAI · Closed
85%
9
Ornith-1.0-397BDeepReinforce AI · Open weight
82.4%
10
Claude Opus 4.5Anthropic · Closed
80.9%
11
Claude Opus 4.6Anthropic · Closed
80.8%
12
DeepSeek V4 Pro 0813DeepSeek · Closed
80.6%
13
MiniMax M3MiniMax · Open weight
80.5%
14
Qwen3.7 MaxAlibaba · Closed
80.4%
15
Kimi K2.6Moonshot AI · Open weight
80.2%
16
Inkling-SmallThinking Machines Lab · Open weight
80.2%
17
GPT-5.2OpenAI · Closed
80%
18
Claude Sonnet 4.6Anthropic · Closed
79.6%
19
DeepSeek V4 Pro (High)DeepSeek · Open weight
79.4%
20
DeepSeek V4 Flash 0731DeepSeek · Closed
79%
21
Ornith-1.5-35B-A3BOrnith AI · Open weight
79%
22
Qwen3.6 PlusAlibaba · Closed
78.8%
23
DeepSeek V4 Flash (High)DeepSeek · Closed
78.6%
24
dots3-note PreviewDots Studio · Open weight
78.4%
25
BTL-4Bad Theory Labs · Open weight
78.4%
26
MiMo-V2-ProXiaomi · Closed
78%
27
GLM-5Z.AI · Open weight
77.8%
28
Qwen3.7 PlusAlibaba · Closed
77.7%
29
InklingThinking Machines Lab · Open weight
77.6%
30
Mistral Medium 3.5 128BMistral · Open weight
77.6%
31
Muse SparkMeta · Closed
77.4%
32
Qwen3.6-27BAlibaba · Open weight
77.2%
33
Claude Sonnet 4.5Anthropic · Closed
77.2%
34
Kimi K2.5Moonshot AI · Open weight
76.8%
35
Kimi K2.5 (Reasoning)Moonshot AI · Closed
76.8%
36
Grok 4.20xAI · Closed
76.7%
37
Qwen3.5 397BAlibaba · Open weight
76.2%
38
Muse Glimmer 30BMeta · Open weight
76%
39
Ornith-1.0-35BDeepReinforce AI · Open weight
75.6%
40
MiMo-V2-OmniXiaomi · Closed
74.8%
41
Laguna M.1Poolside · Closed
74.6%
42
Claude 4.1 OpusAnthropic · Closed
74.5%
43
Hy3 PreviewTencent · Open weight
74.4%
44
GLM-4.7Z.AI · Open weight
73.8%
45
DeepSeek V4 FlashDeepSeek · Closed
73.7%
46
DeepSeek V4 ProDeepSeek · Open weight
73.6%
47
MAI-Thinking-1Microsoft · Closed
73.5%
48
Qwen3.6-35B-A3BAlibaba · Open weight
73.4%
49
MiMo-V2-FlashXiaomi · Open weight
73.4%
50
Claude Haiku 4.5Anthropic · Closed
73.3%
51
Claude 4 SonnetAnthropic · Closed
72.7%
52
Qwen3.5-27BAlibaba · Open weight
72.4%
53
Qwen3.5-122B-A10BAlibaba · Open weight
72%
54
Nemotron 3 UltraNVIDIA · Open weight
71.9%
55
Grok Code Fast 1xAI · Closed
70.8%
56
Ornith-1.5-9BOrnith AI · Open weight
70.6%
57
Laguna XS.2Poolside · Open weight
69.9%
58
Ornith-1.0-9BDeepReinforce AI · Open weight
69.4%
59
Qwen3.5-35B-A3BAlibaba · Open weight
69.2%
60
Gemini 2.5 ProGoogle · Closed
63.8%
61
GPT-4.1OpenAI · Closed
54.6%
62
ZAYA1-74B-PreviewZyphra · Open weight
53.2%
63
52.8%
64
o3-miniOpenAI · Closed
49.3%
65
Claude 3.5 SonnetAnthropic · Closed
49%
66
DeepSeek V3DeepSeek · Open weight
42%
67
GPT-4.1 miniOpenAI · Closed
23.6%

According to BenchLM.ai, Claude Opus 5 leads the SWE-bench Verified benchmark with a score of 96%, followed by Claude Mythos 5 (95.5%) and Claude Fable 5 (95%). The top models are clustered within 1.0 points, suggesting this benchmark is nearing saturation for frontier models.

67 models have been evaluated on SWE-bench Verified. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Verified contributes 16% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE-bench Verified

Year

2024

Tasks

500 verified issues

Format

Code patch generation

Difficulty

Professional software engineering

SWE-bench Verified is the most-cited benchmark for AI coding agents on real repository tasks. Each task requires understanding codebases, writing patches, and passing test suites.

BenchLM freshness & provenance

Version

SWE-bench Verified 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-bench Verified measure?

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Which model scores highest on SWE-bench Verified?

Claude Opus 5 by Anthropic currently leads with a score of 96% on SWE-bench Verified.

How many models are evaluated on SWE-bench Verified?

67 AI models have been evaluated on SWE-bench Verified on BenchLM.

Last updated: August 21, 2026 · BenchLM version SWE-bench Verified 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.