# SWE-Rebench

> A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

Canonical page: https://benchlm.ai/benchmarks/swe-rebench

- Category: [Coding](/coding)
- Last updated: September 15, 2026

## About SWE-Rebench

- Year: 2026
- Tasks: Fresh GitHub issues (rolling window)
- Format: Code patch generation
- Difficulty: Professional software engineering
- Paper: [SWE-Rebench: Contamination-Free Evaluation of Software Engineering Agents](https://swe-rebench.com)

SWE-Rebench uses a rolling window of fresh problems from real GitHub repositories, sourced after each model's release date to prevent contamination. Each model runs 5 times with a standardized 128K-context ReAct scaffold. Unlike SWE-bench Verified (2023 problems), scores reflect consistent, up-to-date difficulty.

SWE-Rebench is currently weighted in BenchLM's scoring formula. The Coding category carries 20% of the overall score, and SWE-Rebench contributes 10% of that category score.

## Leaderboard (13 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 65.3% |
| 2 | [GLM-5](/models/glm-5) | Z.AI | 62.8% |
| 3 | [GLM-5.1](/models/glm-5-1) | Z.AI | 62.7% |
| 4 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 60.9% |
| 5 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 60.7% |
| 6 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 58.9% |
| 7 | [GLM-4.7](/models/glm-4-7) | Z.AI | 58.7% |
| 8 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 58.5% |
| 9 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 58.2% |
| 10 | [Composer 2](/models/composer-2) | Cursor | 58% |
| 11 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 53.7% |
| 12 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 51.9% |
| 13 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 41.6% |

## FAQ

### What does SWE-Rebench measure?

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

### Which model scores highest on SWE-Rebench?

Claude Opus 4.6 by Anthropic currently leads with a score of 65.3% on SWE-Rebench.

### How many models are evaluated on SWE-Rebench?

13 AI models have been evaluated on SWE-Rebench on BenchLM.

### Does SWE-Rebench affect BenchLM's overall score?

Yes. SWE-Rebench is a weighted benchmark inside the Coding category, which carries 20% of BenchLM's overall score. SWE-Rebench itself contributes 10% of that category score.

## Compare Top Models on SWE-Rebench

- [Claude Opus 4.6 vs GLM-5](/compare/claude-opus-4-6-vs-glm-5)
- [GLM-5 vs GLM-5.1](/compare/glm-5-vs-glm-5-1)
- [GLM-5.1 vs DeepSeek V3.2](/compare/deepseek-v3-2-vs-glm-5-1)
- [DeepSeek V3.2 vs Claude Sonnet 4.6](/compare/claude-sonnet-4-6-vs-deepseek-v3-2)
