Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

SWE-bench, Vals AI run (SWE-bench (Vals))

Vals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals removed SWE-bench Verified from its index as saturated on 2026-05-04; this board is the standalone SWE-bench run.

Data verified 23 confirmed releases in the last 30 daysSee the free Radar Brief

How BenchLM shows Vals SWE-bench mirror

BenchLM mirrors the public Vals AI Vals SWE-bench mirror leaderboard captured from https://www.vals.ai/benchmarks/swebench and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals SWE-bench mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

88 Vals rows5 task viewspublic datasetTasks: Overall, 1-4 hours, 15 min - 1 hour, <15 min fix, >4 hoursDisplay only

Vals SWE-bench score on SWE-bench (Vals) — September 1, 2026

We mirror the published vals swe-bench score view for SWE-bench (Vals). Claude Opus 5 leads the public snapshot at 97.0%, followed by DeepSeek V4 Pro 0813 (96.4%) and GPT-5.6 Sol (96.2%). We do not use these results to rank models overall.

88 modelsCodingCurrentDisplay onlyUpdated September 1, 2026

Vals SWE-bench score table (88 models)

Score
1
Claude Opus 5Anthropic · Closed
97.0%
2
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
96.4%
3
GPT-5.6 SolOpenAI · Closedmax reasoning
96.2%
4
Grok 4.6xAI · Closedhigh reasoning
95.6%
5
GPT-5.6 TerraOpenAI · Closedmax reasoning
95.4%
6
GLM-5.3Z.AI · Open weightmax reasoning
95.4%
7
Claude Fable 5Anthropic · Closed
95.0%
8
Kimi K3Moonshot AI · Closed
93.4%
9
GPT-5.6 LunaOpenAI · Closedmax reasoning
93.0%
10
GLM-5.3-FlashZ.AI · Open weightmax reasoning
92.0%
11
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
88.8%
12
Claude Opus 4.8Anthropic · Closed
88.6%
13
Grok 4.5xAI · Closedhigh reasoning
86.6%
14
Muse Spark 1.2Meta · Closedxhigh reasoning
86.6%
15
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
86.0%
16
Claude Opus 4.8 Claude CodeAnthropicClaude CodeClaude Code
85.8%
17
Qwen3.8 MaxAlibaba · Open weight
85.6%
18
GLM-5.2Z.AI · Open weightmax reasoning
82.8%
19
GPT-5.5OpenAI · Closedxhigh reasoning
82.6%
20
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
82.2%
21
Muse Spark 1.1Meta · Closedxhigh reasoning
82.0%
22
Claude Opus 4.7Anthropic · Closed
82.0%
23
Gemini 3.7 FlashGoogle · Closedhigh reasoning
80.8%
24
Gemini 3.8 FlashGoogle · Closedhigh reasoning
80.0%
25
Composer 2.5Cursor · ClosedCursor CLICursor CLI
79.6%
26
Gemini 3.6 FlashGoogle · Closedhigh reasoning
79.6%
27
Claude Sonnet 5Anthropic · Closed
79.6%
28
Gemini 3.5 FlashGoogle · Closedhigh reasoning
78.8%
29
Gemini 3.1 Pro PreviewGooglehigh reasoning
78.8%
30
GPT-5.4OpenAI · Closedxhigh reasoning
78.2%
31
Claude Opus 4.6 (Adaptive)Anthropic · Closed
78.2%
32
Kimi K2.7 CodeMoonshot AI · Open weight
78.2%
33
GPT-5.3 CodexOpenAI · Closedxhigh reasoning
78.0%
34
InklingThinking Machines Lab · Open weight0.99 reasoning
77.6%
35
Claude Sonnet 4.6Anthropic · Closed
77.4%
36
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
77.4%
37
GPT-5.5 CodexOpenAICodexCodex
76.4%
38
Claude Opus 4.5 ThinkingAnthropic · Closed
76.4%
39
Gemini 3 Pro PreviewGooglehigh reasoning
76.4%
40
GLM-5.1Z.AI · Open weight
76.4%
41
GPT-5.5 FactoryOpenAIFactoryFactory
76.2%
42
Kimi K2.6Moonshot AI · Open weight
76.2%
43
GPT-5.2OpenAI · Closedxhigh reasoning
75.8%
44
Gemini 3 Flash PreviewGooglehigh reasoning
75.0%
45
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
75.0%
46
MiniMax M3MiniMax · Open weight
75.0%
47
74.8%
48
Muse SparkMeta · Closed
74.4%
49
MiniMax M2.5MiniMax · Closed
74.2%
50
MiMo-V2.5-ProXiaomi · Closed
74.0%
51
MiniMax M2.7MiniMax · Open weight
73.8%
52
Qwen3.6 PlusAlibaba · Closed
73.4%
53
GPT-5.4 miniOpenAI · Closedxhigh reasoning
73.0%
54
Qwen 3.6 Max (preview)Alibaba · Closed
72.8%
55
GPT-5.2-CodexOpenAI · Closedhigh reasoning
72.4%
57
Grok 4.3xAI · Closedhigh reasoning
71.4%
58
71.4%
59
71.2%
60
MiMo-V2.5Xiaomi · Closed
71.0%
61
70.0%
62
Claude Sonnet 4.5 ThinkingAnthropic · Closed
70.0%
63
Qwen3.6-27BAlibaba · Open weight
70.0%
64
GPT-5.4 nanoOpenAI · Closedhigh reasoning
69.8%
65
GPT-5.1OpenAI · Closedhigh reasoning
69.8%
66
GLM-4.7Z.AI · Open weight
69.4%
67
GPT-5OpenAIhigh reasoning
69.0%
69
Qwen3.7 MaxAlibaba · Closed
68.8%
70
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
67.6%
71
Claude Haiku 4.5 ThinkingAnthropic · Closed
66.6%
72
Mistral Medium 3.5Mistral AIhigh reasoning
66.4%
74
Qwen3.5 FlashAlibaba · Closed
64.4%
75
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
62.8%
76
Devstral 2512Mistral AI
62.8%
77
GPT-5 miniOpenAI · Closedhigh reasoning
60.8%
78
Kimi K2 ThinkingMoonshot AI
60.2%
79
57.8%
80
Laguna M.1Poolside · Closed
57.6%
81
Laguna XS.2Poolside · Open weight
55.2%
82
Gemini 2.5 ProGoogle · Closed
54.4%
84
45.4%
85
41.4%
86
41.4%
87
GPT-OSS 120BOpenAI · Open weight
33.6%
88
7.8%

The published SWE-bench (Vals) snapshot places Claude Opus 5 first at 97.0%. The third row is 0.8 points behind. The broader top-10 range is 5.0 points, so many of the published results sit in a relatively narrow band.

88 models have been evaluated on SWE-bench (Vals). The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. SWE-bench (Vals) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE-bench (Vals)

Year

2026

Tasks

Real repository issues by human time bucket

Format

Resolved rate

Difficulty

Frontier coding agents

BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).

BenchLM freshness & provenance

Version

SWE-bench (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-bench (Vals) measure?

Vals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals removed SWE-bench Verified from its index as saturated on 2026-05-04; this board is the standalone SWE-bench run.

Which model leads the published SWE-bench (Vals) snapshot?

Claude Opus 5 currently leads the published SWE-bench (Vals) snapshot with 97.0% vals swe-bench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SWE-bench (Vals)?

The September 1, 2026 snapshot contains 88 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.