Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

SWE-Rebench

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

Claude Opus 4.6 leads the SWE-Rebench leaderboard on BenchLM's September 2026 update with 65.3%, ahead of GLM-5 (62.8%) and GLM-5.1 (62.7%), across 13 models.

Top models on SWE-Rebench — September 15, 2026

As of September 15, 2026, Claude Opus 4.6 leads the SWE-Rebench leaderboard with 65.3% , followed by GLM-5 (62.8%) and GLM-5.1 (62.7%).

13 modelsCoding10% of category scoreCurrentUpdated September 15, 2026

Leaderboard (13 models)

Score
1
Claude Opus 4.6Anthropic · Closed
65.3%
2
GLM-5Z.AI · Open weight
62.8%
3
GLM-5.1Z.AI · Open weight
62.7%
4
DeepSeek V3.2DeepSeek · Open weight
60.9%
5
Claude Sonnet 4.6Anthropic · Closed
60.7%
6
Qwen3.5-27BAlibaba · Open weight
58.9%
7
GLM-4.7Z.AI · Open weight
58.7%
8
Kimi K2.5Moonshot AI · Open weight
58.5%
9
GPT-5.3 CodexOpenAI · Closed
58.2%
10
Composer 2Cursor · Closed
58%
11
Qwen3.5-35B-A3BAlibaba · Open weight
53.7%
12
MiniMax M2.7MiniMax · Open weight
51.9%
13
Gemma 4 31BGoogle · Open weight
41.6%

According to BenchLM.ai, Claude Opus 4.6 leads the SWE-Rebench benchmark with a score of 65.3%, followed by GLM-5 (62.8%) and GLM-5.1 (62.7%). The top models are clustered within 2.6 points, suggesting this benchmark is nearing saturation for frontier models.

13 models have been evaluated on SWE-Rebench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-Rebench contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.

About SWE-Rebench

Year

2026

Tasks

Fresh GitHub issues (rolling window)

Format

Code patch generation

Difficulty

Professional software engineering

SWE-Rebench uses a rolling window of fresh problems from real GitHub repositories, sourced after each model's release date to prevent contamination. Each model runs 5 times with a standardized 128K-context ReAct scaffold. Unlike SWE-bench Verified (2023 problems), scores reflect consistent, up-to-date difficulty.

BenchLM freshness & provenance

Version

Rolling 2026 window

Refresh cadence

Rolling

Staleness state

Current

Question availability

Rolling public issues

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-Rebench measure?

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

Which model scores highest on SWE-Rebench?

Claude Opus 4.6 by Anthropic currently leads with a score of 65.3% on SWE-Rebench.

How many models are evaluated on SWE-Rebench?

13 AI models have been evaluated on SWE-Rebench on BenchLM.

Last updated: September 15, 2026 · BenchLM version Rolling 2026 window

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.