Skip to main content

Benchmark profile

Vals-hosted SWE-bench mirror (Vals SWE-bench mirror)

Vals AI hosted SWE-bench view for solving production software engineering tasks.

Data verified

How BenchLM shows Vals SWE-bench mirror

BenchLM mirrors the public Vals AI Vals SWE-bench mirror leaderboard captured from https://www.vals.ai/benchmarks/swebench and updated by Vals on July 22, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals SWE-bench mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

75 Vals rows5 task viewspublic datasetTasks: Overall, 1-4 hours, 15 min - 1 hour, <15 min fix, >4 hoursDisplay only

Vals SWE-bench score on Vals SWE-bench mirror — July 22, 2026

BenchLM mirrors the published vals swe-bench score view for Vals SWE-bench mirror. Claude Opus 5 leads the public snapshot at 97.00% , followed by GPT-5.6 Sol (96.20%) and Claude Fable 5 (95.00%). BenchLM does not use these results to rank models overall.

75 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 22, 2026

Vals SWE-bench score table (75 models)

Score
1
Claude Opus 5Anthropic
97.00%
2
96.20%
3
95.00%
4
Kimi K3Moonshot AI
93.40%
5
93.00%
6
88.60%
7
86.60%
8
Claude Opus 4.8 Claude CodeAnthropicClaude Code
85.80%
9
GLM 5.2Zhipu AI
82.80%
10
GPT-5.5OpenAI
82.60%
11
82.00%
12
82.00%
13
Composer 2.5CursorCursor CLI
79.60%
14
79.60%
15
79.60%
16
78.80%
18
GPT-5.4OpenAI
78.20%
19
78.20%
20
Kimi K2.7 CodeMoonshot AI
78.20%
21
78.00%
22
InklingThinkingmachines
77.60%
23
77.40%
24
77.40%
25
GPT-5.5 CodexOpenAICodex
76.40%
27
76.40%
28
GLM 5.1Zhipu AI
76.40%
29
GPT-5.5 FactoryOpenAIFactory
76.20%
30
Kimi K2.6Moonshot AI
76.20%
31
GPT-5.2OpenAI
75.80%
32
75.20%
34
75.00%
35
MiniMax M3MiniMax
75.00%
36
74.80%
37
74.40%
38
74.20%
39
74.00%
40
73.80%
41
73.40%
42
73.00%
43
72.80%
44
72.40%
46
71.40%
47
71.40%
48
71.20%
49
Mimo V2.5Xiaomi
71.00%
50
70.00%
52
70.00%
53
69.80%
54
GPT-5.1OpenAI
69.80%
55
GLM 4.7Zhipu AI
69.40%
56
GPT-5OpenAI
69.00%
58
68.80%
59
67.60%
61
66.40%
62
64.40%
64
Devstral 2512Mistral AI
62.80%
65
60.80%
66
Kimi K2 ThinkingMoonshot AI
60.20%
67
57.80%
68
Laguna M.1Poolside
57.60%
69
Laguna Xs.2Poolside
55.20%
70
54.40%
73
41.40%
74
GPT Oss 120bFireworks AI
33.60%
75
7.80%

The published Vals SWE-bench mirror snapshot places Claude Opus 5 first at 97.00%. The third row is 2.00 points behind. The broader top-10 range is 14.40 points, so the table still separates the published systems.

75 models have been evaluated on Vals SWE-bench mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals SWE-bench mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals SWE-bench mirror

Year

2026

Tasks

Software engineering issue-resolution tasks

Format

Accuracy score

Difficulty

Production software engineering

BenchLM keeps this separate from its canonical SWE-bench Verified page so Vals-hosted results remain secondary context rather than source-of-record data.

BenchLM freshness & provenance

Version

Vals SWE-bench mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vals SWE-bench mirror measure?

Vals AI hosted SWE-bench view for solving production software engineering tasks.

Which model leads the published Vals SWE-bench mirror snapshot?

Claude Opus 5 currently leads the published Vals SWE-bench mirror snapshot with 97.00% vals swe-bench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals SWE-bench mirror?

75 AI models are included in BenchLM's mirrored Vals SWE-bench mirror snapshot, based on the public leaderboard captured on July 22, 2026.

Last updated: July 22, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.