# Multi-SWE Bench

> A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.

Canonical page: https://benchlm.ai/benchmarks/multiswebench

- Category: [Coding](/coding)
- Last updated: September 15, 2026

## About Multi-SWE Bench

- Year: 2026
- Tasks: Multi-language repo tasks
- Format: Repository task completion
- Difficulty: Professional software engineering
- Paper: [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en)

MiniMax positions Multi-SWE Bench as a benchmark closer to real engineering work than isolated code generation, emphasizing multi-language repository workflows.

Multi-SWE Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 52.7% |

## FAQ

### What does Multi-SWE Bench measure?

A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.

### Which model scores highest on Multi-SWE Bench?

MiniMax M2.7 by MiniMax currently leads with a score of 52.7% on Multi-SWE Bench.

### How many models are evaluated on Multi-SWE Bench?

1 AI models have been evaluated on Multi-SWE Bench on BenchLM.

### Does Multi-SWE Bench affect BenchLM's overall score?

Not directly. Multi-SWE Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
