Skip to main content
BenchLM

SWE-bench Multilingual — Mercor evaluation (SWE-bench Multilingual (Mercor))

We show this table for reference; we do not rank on it.

Resolving repository issues across programming languages. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Evaluation results published by Mercor. Published evaluation aggregates only; task contents and dataset license grants are not included.

Pass@1 on SWE-bench Multilingual (Mercor) — October 1, 2026 capture

We mirror the published pass@1 view for SWE-bench Multilingual (Mercor). DeepSeek-V4.1-Flash leads the public snapshot at 98.20%, followed by DeepSeek-V4-Pro-0813 (96.40%) and Opus 5 (96.20%). We do not use these results to rank models overall.

38 configurationsCodingDated Mercor configurationsDisplay onlyUpdated October 1, 2026 capture

Pass@1 table (38 configurations)

Score
1
DeepSeek-V4.1-FlashDeepSeekmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±1.1 · 894 reported samples
98.20%
2
DeepSeek-V4-Pro-0813DeepSeekmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±1.8 · 900 reported samples
96.40%
3
Opus 5Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±1.7 · 900 reported samples
96.20%
4
Opus 5.5Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.2 · 894 reported samples
95.50%
5
DeepSeek-V4-FlashDeepSeekmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±1.9 · 894 reported samples
95.40%
6
Fable 5Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.1 · 900 reported samples
94.80%
7
GPT-5.6 TerraOpenAImini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.3 · 900 reported samples
94.40%
8
GPT-5.6 Sol (Pro)OpenAImini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.3 · 900 reported samples
94.20%
9
Grok 4.6xAImini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoningPublished margin ±2.6 · 900 reported samples
92.70%
10
GPT-5.6 LunaOpenAImini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.7 · 900 reported samples
92.40%
11
Kimi K3Kimimini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.6 · 900 reported samples
91.40%
12
Fable 5.1Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±3 · 894 reported samples
91.10%
13
Sonnet 5.5Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±3 · 893 reported samples
90.90%
14
Sonnet 5Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.9 · 900 reported samples
89.20%
15
GLM-5.3-FlashZhipumini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±2.8 · 900 reported samples
88.90%
16
Qwen3.8-MaxAlibabamini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoningPublished margin ±2.9 · 900 reported samples
88.30%
17
Grok 4.5xAImini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±3.1 · 900 reported samples
87.70%
18
GLM-5.3Zhipumini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±3 · 900 reported samples
87.60%
19
Opus 4.8Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±3.4 · 900 reported samples
86.60%
20
Gemini 3.8 FlashGooglemini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±3.7 · 894 reported samples
85.80%
21
Opus 4.7Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±3.4 · 900 reported samples
85.30%
22
GPT-6.1 SolOpenAImini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±3.9 · 894 reported samples
84.90%
23
Gemini 3.7 FlashGooglemini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4 · 900 reported samples
82.60%
24
GPT-6 SolOpenAImini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±4.2 · 894 reported samples
81.70%
25
GLM-5.2Zhipumini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±4 · 900 reported samples
78.60%
26
GPT-6 LunaOpenAImini-swe-agent (1000 max steps, 3 hour time limit) · max reasoningPublished margin ±4.4 · 894 reported samples
78.20%
27
Sonnet 4.6Anthropicmini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4.2 · 900 reported samples
78.10%
28
GPT-5.5OpenAImini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoningPublished margin ±4.2 · 900 reported samples
77.70%
29
InklingThinking Machinesmini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±3.7 · 900 reported samples
77.70%
30
GPT-5.4OpenAImini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoningPublished margin ±4.2 · 900 reported samples
77.00%
31
MiniMax-M3MiniMaxmini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±3.9 · 900 reported samples
76.00%
32
Kimi K2.7 CodeKimimini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4.1 · 900 reported samples
75.60%
33
Gemini 3.6 FlashGooglemini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4.5 · 900 reported samples
75.10%
34
Gemini 3.5 FlashGooglemini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4.5 · 900 reported samples
68.20%
35
Gemini 3.1 ProGooglemini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4.3 · 900 reported samples
65.10%
36
Qwen3.5Alibabamini-swe-agent (1000 max steps, 3 hour time limit)Published margin ±4.7 · 900 reported samples
61.70%
37
MiniMax-M2.7MiniMaxmini-swe-agent (1000 max steps, 3 hour time limit) · high reasoningPublished margin ±4.8 · 900 reported samples
60.80%
38
DeepSeek-V3.2DeepSeekmini-swe-agent (1000 max steps, 3 hour time limit)Published margin ±4.8 · 900 reported samples
60.70%

How to read this Mercor evaluation

We captured Mercor's SWE-bench Multilingual results on October 1, 2026. The table preserves 38 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

This dated source table is display only and does not enter overall or category rankings. Source configurations retain their external identity, including Pro descriptors and different effort settings.

The headline is the source page's selected Pass@1 metric. The task set, harness, tools, and grader belong to this Mercor run and may differ from another evaluator's results.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists programmatic fail-to-pass and pass-to-pass unit tests (all must pass) and a judge of None. Its run-count label is k = 3. The linked result page takes precedence where its setup differs, as noted above.

The published SWE-bench Multilingual (Mercor) snapshot places DeepSeek-V4.1-Flash first at 98.20%. The third row is 2.00 points behind. The broader top-10 range is 5.80 points, so many of the published results sit in a relatively narrow band.

38 configurations are shown for SWE-bench Multilingual (Mercor). The benchmark falls in the Coding category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. SWE-bench Multilingual (Mercor) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE-bench Multilingual (Mercor)

Tasks

298 source-reported tasks

Format

Pass@1

Difficulty

Source-specific evaluation

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

Freshness and provenance

Version

SWE-bench Multilingual; Mercor Pass@1

Refresh cadence

Reviewed dated source capture

Staleness state

Dated Mercor configurations

Question availability

Upstream task access and terms; only aggregate results mirrored

Dated Mercor configurationsDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SWE-bench Multilingual (Mercor) measure?

Resolving repository issues across programming languages. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Which model leads the published SWE-bench Multilingual (Mercor) snapshot?

DeepSeek-V4.1-Flash currently leads the published SWE-bench Multilingual (Mercor) snapshot with 98.20% pass@1. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SWE-bench Multilingual (Mercor)?

The October 1, 2026 capture snapshot contains 38 source configurations.

Last updated: October 1, 2026 capture · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.