# SWE-bench Multilingual (SWE Multilingual)

> A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.

Canonical page: https://benchlm.ai/benchmarks/swe-bench-multilingual

- Category: [Multilingual](/multilingual)
- Last updated: September 18, 2026

## About SWE Multilingual

- Year: 2025
- Tasks: 300 problems across 9 languages
- Format: Multi-language code patch generation
- Difficulty: Professional multilingual software engineering
- Paper: [SWE-bench Multilingual](https://www.swebench.com/multilingual)

SWE-bench Multilingual extends evaluation to Java, JavaScript, TypeScript, C++, Go, Rust, Ruby, PHP, and Swift. Mythos Preview achieves 87.3% averaged over 5 trials.

SWE Multilingual is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 92.2% |

## FAQ

### What does SWE Multilingual measure?

A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.

### Which model scores highest on SWE Multilingual?

Claude Mythos 5 by Anthropic currently leads with a score of 92.2% on SWE Multilingual.

### How many models are evaluated on SWE Multilingual?

1 AI models have been evaluated on SWE Multilingual on BenchLM.

### Does SWE Multilingual affect BenchLM's overall score?

Not directly. SWE Multilingual is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
