# SWE-bench Multilingual — Mercor evaluation (SWE-bench Multilingual (Mercor))

> Resolving repository issues across programming languages. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Canonical page: https://benchlm.ai/benchmarks/mercorswemultilingual

- Category: [Coding](/coding)
- Last updated: October 1, 2026 capture

Evaluation results published by [Mercor](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/). Published evaluation aggregates only; task contents and dataset license grants are not included.

We captured Mercor's SWE-bench Multilingual results on October 1, 2026. The table preserves 38 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

This dated source table is display only and does not enter overall or category rankings. Source configurations retain their external identity, including Pro descriptors and different effort settings.

The headline is the source page's selected Pass@1 metric. The task set, harness, tools, and grader belong to this Mercor run and may differ from another evaluator's results.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists programmatic fail-to-pass and pass-to-pass unit tests (all must pass) and a judge of None. Its run-count label is k = 3. The linked result page takes precedence where its setup differs, as noted above.

- [Mercor SWE-bench Multilingual results](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/)
- [Mercor APEX benchmark index](https://www.mercor.com/apex/)
- [Mercor methodology](https://www.mercor.com/apex/methodology/)
- [Source paper](https://arxiv.org/abs/2310.06770)
- [Upstream evaluation code](https://github.com/SWE-bench/SWE-bench)
- [Upstream dataset](https://huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual)

## About SWE-bench Multilingual (Mercor)

- Tasks: 298 source-reported tasks
- Format: Pass@1
- Difficulty: Source-specific evaluation
- Paper: [SWE-bench Multilingual evaluation results](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/)

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

SWE-bench Multilingual (Mercor) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (38 configurations)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [DeepSeek-V4.1-Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±1.1 · 894 reported samples | DeepSeek | 98.20% |
| 2 | [DeepSeek-V4-Pro-0813](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±1.8 · 900 reported samples | DeepSeek | 96.40% |
| 3 | [Opus 5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±1.7 · 900 reported samples | Anthropic | 96.20% |
| 4 | [Opus 5.5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.2 · 894 reported samples | Anthropic | 95.50% |
| 5 | [DeepSeek-V4-Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±1.9 · 894 reported samples | DeepSeek | 95.40% |
| 6 | [Fable 5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.1 · 900 reported samples | Anthropic | 94.80% |
| 7 | [GPT-5.6 Terra](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.3 · 900 reported samples | OpenAI | 94.40% |
| 8 | [GPT-5.6 Sol (Pro)](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.3 · 900 reported samples | OpenAI | 94.20% |
| 9 | [Grok 4.6](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoning · Published margin ±2.6 · 900 reported samples | xAI | 92.70% |
| 10 | [GPT-5.6 Luna](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.7 · 900 reported samples | OpenAI | 92.40% |
| 11 | [Kimi K3](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.6 · 900 reported samples | Kimi | 91.40% |
| 12 | [Fable 5.1](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±3 · 894 reported samples | Anthropic | 91.10% |
| 13 | [Sonnet 5.5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±3 · 893 reported samples | Anthropic | 90.90% |
| 14 | [Sonnet 5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.9 · 900 reported samples | Anthropic | 89.20% |
| 15 | [GLM-5.3-Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±2.8 · 900 reported samples | Zhipu | 88.90% |
| 16 | [Qwen3.8-Max](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoning · Published margin ±2.9 · 900 reported samples | Alibaba | 88.30% |
| 17 | [Grok 4.5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±3.1 · 900 reported samples | xAI | 87.70% |
| 18 | [GLM-5.3](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±3 · 900 reported samples | Zhipu | 87.60% |
| 19 | [Opus 4.8](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±3.4 · 900 reported samples | Anthropic | 86.60% |
| 20 | [Gemini 3.8 Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±3.7 · 894 reported samples | Google | 85.80% |
| 21 | [Opus 4.7](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±3.4 · 900 reported samples | Anthropic | 85.30% |
| 22 | [GPT-6.1 Sol](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±3.9 · 894 reported samples | OpenAI | 84.90% |
| 23 | [Gemini 3.7 Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4 · 900 reported samples | Google | 82.60% |
| 24 | [GPT-6 Sol](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±4.2 · 894 reported samples | OpenAI | 81.70% |
| 25 | [GLM-5.2](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±4 · 900 reported samples | Zhipu | 78.60% |
| 26 | [GPT-6 Luna](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · max reasoning · Published margin ±4.4 · 894 reported samples | OpenAI | 78.20% |
| 27 | [Sonnet 4.6](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4.2 · 900 reported samples | Anthropic | 78.10% |
| 28 | [GPT-5.5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoning · Published margin ±4.2 · 900 reported samples | OpenAI | 77.70% |
| 29 | [Inkling](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±3.7 · 900 reported samples | Thinking Machines | 77.70% |
| 30 | [GPT-5.4](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · xhigh reasoning · Published margin ±4.2 · 900 reported samples | OpenAI | 77.00% |
| 31 | [MiniMax-M3](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±3.9 · 900 reported samples | MiniMax | 76.00% |
| 32 | [Kimi K2.7 Code](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4.1 · 900 reported samples | Kimi | 75.60% |
| 33 | [Gemini 3.6 Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4.5 · 900 reported samples | Google | 75.10% |
| 34 | [Gemini 3.5 Flash](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4.5 · 900 reported samples | Google | 68.20% |
| 35 | [Gemini 3.1 Pro](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4.3 · 900 reported samples | Google | 65.10% |
| 36 | [Qwen3.5](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · Published margin ±4.7 · 900 reported samples | Alibaba | 61.70% |
| 37 | [MiniMax-M2.7](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · high reasoning · Published margin ±4.8 · 900 reported samples | MiniMax | 60.80% |
| 38 | [DeepSeek-V3.2](https://www.mercor.com/apex/oss-benchmarks/oss-swe-bench-multilingual-leaderboard/) | mini-swe-agent (1000 max steps, 3 hour time limit) · Published margin ±4.8 · 900 reported samples | DeepSeek | 60.70% |

## FAQ

### What does SWE-bench Multilingual (Mercor) measure?

Resolving repository issues across programming languages. This table shows Mercor-run configurations for reference and is excluded from model rankings.

### Which model leads the published SWE-bench Multilingual (Mercor) snapshot?

DeepSeek-V4.1-Flash currently leads the published SWE-bench Multilingual (Mercor) snapshot with a score of 98.20%.

### How many models are evaluated on SWE-bench Multilingual (Mercor)?

The October 1, 2026 capture contains 38 source configurations.

### Does SWE-bench Multilingual (Mercor) affect BenchLM's overall score?

Not directly. SWE-bench Multilingual (Mercor) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
