Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

SWE Multilingual

Data verified 35 confirmed releases in the last 30 daysFollow model changes

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

Top models on SWE Multilingual — September 18, 2026

As of September 18, 2026, Claude Opus 5 leads the SWE Multilingual leaderboard with 89.5% , followed by Claude Fable 5.1 (89.1%) and Claude Opus 4.8 (84.4%).

45 modelsCoding5% of category scoreCurrentUpdated September 18, 2026

Leaderboard (45 models)

Score
1
Claude Opus 5Anthropic · Closed
89.5%
2
Claude Fable 5.1Anthropic · Closed
89.1%
3
Claude Opus 4.8Anthropic · Closed
84.4%
4
Hy4 previewTencent · Open weight
82.9%
5
Qwen3.8-Flash-NextAlibaba · Open weight
81%
6
Qwen3.8-Omni-FlashAlibaba · Closed
80.5%
7
Composer 2.5Cursor · Closed
79.8%
8
Ornith-1.5-397BOrnith AI · Open weight
79.6%
9
Ornith-1.0-397BDeepReinforce AI · Open weight
78.9%
10
Laguna S 2.1Poolside · Open weight
78.5%
11
Claude Sonnet 5Anthropic · Closed
78.3%
12
Qwen3.7 MaxAlibaba · Closed
78.3%
13
Grok 4.5xAI · Closed
78%
14
SWE-1.7Cognition · Closed
77.8%
15
Claude Opus 4.5Anthropic · Closed
77.5%
16
Kimi K2.6Moonshot AI · Open weight
76.7%
17
MiniMax M2.7MiniMax · Open weight
76.5%
18
DeepSeek V4 Pro 0813DeepSeek · Closed
76.2%
19
Qwen3.7 PlusAlibaba · Closed
75.8%
20
dots3-note PreviewDots Studio · Open weight
75.7%
21
DeepSeek V4 Pro (High)DeepSeek · Open weight
74.1%
22
Qwen3.6 PlusAlibaba · Closed
73.8%
23
Composer 2Cursor · Closed
73.7%
24
GLM-5Z.AI · Open weight
73.3%
25
DeepSeek V4 Flash 0731DeepSeek · Closed
73.3%
26
Kimi K2.5Moonshot AI · Open weight
73%
27
Ling 3.0 FlashInclusionAI · Open weight
72.4%
28
Ornith-1.5-35B-A3BOrnith AI · Open weight
71.4%
29
Qwen3.6-27BAlibaba · Open weight
71.3%
30
DeepSeek V4 Flash (High)DeepSeek · Closed
70.2%
31
DeepSeek V4 ProDeepSeek · Open weight
69.8%
32
DeepSeek V4 FlashDeepSeek · Closed
69.7%
33
Ornith-1.0-35BDeepReinforce AI · Open weight
69.3%
34
Nemotron 3 UltraNVIDIA · Open weight
67.7%
35
Qwen3.6-35B-A3BAlibaba · Open weight
67.2%
36
Laguna M.1Poolside · Closed
63.1%
37
Laguna XS 2.1Poolside · Open weight
63.1%
38
LongCat-Flash-Lite-SparseMeituan · Open weight
59.3%
39
Laguna XS.2Poolside · Open weight
57.7%
40
Ornith-1.5-9BOrnith AI · Open weight
54.4%
41
Ornith-1.0-9BDeepReinforce AI · Open weight
52%
42
Granite 4.2 30BIBM · Open weight
41.9%
43
36.5%
44
Granite 4.2 8BIBM · Open weight
30.8%
45
LLaDA2.2-flashInclusionAI · Open weight
25%

According to BenchLM.ai, Claude Opus 5 leads the SWE Multilingual benchmark with a score of 89.5%, followed by Claude Fable 5.1 (89.1%) and Claude Opus 4.8 (84.4%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

45 models have been evaluated on SWE Multilingual. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE Multilingual contributes 5% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE Multilingual

Year

2026

Tasks

Multilingual software-engineering tasks

Format

Repository task completion

Difficulty

Professional software engineering

MiniMax reports SWE Multilingual as a coding benchmark focused on multilingual software-engineering tasks beyond single-language Python issue fixing.

Freshness and provenance

Version

SWE Multilingual 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SWE Multilingual measure?

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

Which model scores highest on SWE Multilingual?

Claude Opus 5 by Anthropic currently leads with a score of 89.5% on SWE Multilingual.

How many models are evaluated on SWE Multilingual?

45 AI models have been evaluated on SWE Multilingual on BenchLM.

Last updated: September 18, 2026 · BenchLM version SWE Multilingual 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.