Skip to main content
BenchLM
Data

SWE-sweep: Can Agents Autonomously Find and Fix Bugs? (SWE-sweep)

We show this table for reference; we do not rank on it.

Tests whether coding agents can find and fix multiple bugs across a repository without issue descriptions, while preserving existing behavior.

Bugs resolved on SWE-sweep — September 24, 2026 snapshot

We mirror the published bugs resolved view for SWE-sweep. GPT-5.6 Sol leads the public snapshot at 4.7%, followed by GPT-5.6 Luna (2.5%) and GPT-5.6 Terra (1.5%). We do not use these results to rank models overall.

10 configurationsAgenticCurrentDisplay onlyUpdated September 24, 2026 snapshot

Bugs resolved table (10 configurations)

Score
1
GPT-5.6 SolOpenAI · Closedmini-SWE-agent · xhigh reasoning$72.30 per repository · 232.54 turns · 104,813.27 output tokens
4.7%
2
GPT-5.6 LunaOpenAI · Closedmini-SWE-agent · xhigh reasoning$2.24 per repository · 204.13 turns · 61,210.82 output tokens
2.5%
3
GPT-5.6 TerraOpenAI · Closedmini-SWE-agent · xhigh reasoning$3.57 per repository · 76.32 turns · 46,647.27 output tokens
1.5%
4
GPT-5.6 LunaOpenAI · Closedmini-SWE-agent · high reasoning$0.28 per repository · 74.58 turns · 20,288.3 output tokens
1.4%
5
Claude Opus 5Anthropic · Closedmini-SWE-agent · xhigh reasoning$53.63 per repository · 322.78 turns · 192,749.15 output tokens
1.3%
6
Kimi K3Moonshot AI · Closedmini-SWE-agent · max reasoning$24.51 per repository · 337.27 turns · 152,612.73 output tokens
0.6%
7
GPT-5.6 LunaOpenAI · Closedmini-SWE-agent · default reasoning$0.04 per repository · 22.49 turns · 4,465.53 output tokens
0.5%
8
GPT-5.4 miniOpenAI · Closedmini-SWE-agent · high reasoning$1.22 per repository · 75.04 turns · 40,513.27 output tokens
0.5%
9
GPT-5.4 miniOpenAI · Closedmini-SWE-agent · default reasoning$0.05 per repository · 12.68 turns · 2,147.18 output tokens
0.2%
10
Gemini 3.5 Flash-LiteGoogle · Closedmini-SWE-agent · default reasoning$0.06 per repository · 33.70 turns · 5,228.24 output tokens
0.1%

How to read this leaderboard

The score divides resolved reference bugs by all reference bugs across the repository set. Each bug has equal weight, so repositories with more reference bugs contribute more to the micro-average. A repository where the patch introduces a new failure in the restored original test suite contributes no resolved bugs.

Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is GPT-5.6 Sol at 4.7%.

Honest limit: Each row measures a model, reasoning effort, and mini-SWE-agent setup. The reference-bug set and regression tests define which repairs receive credit; the result does not measure every possible defect in a repository.

How we show SWE-sweep

We mirror the official SWE-sweep table from the September 24, 2026 snapshot, retrieved on October 2, 2026. The source reports 10 mini-SWE-agent configurations on 100 repositories, scored by the percentage of reference bugs resolved across the full set.

Agents must find and fix bugs without issue descriptions. Hidden tests check each fix. A new failure in the restored original test suite sets that repository's score to zero.

Each reasoning-effort setting stays on its own row. The table is display only and excluded from overall and category rankings because its scores belong to the published model, effort, and agent setup.

The table averages cost, turns, and output tokens per evaluated repository, including non-deprecated retries. The paper supplies Kimi K3’s max reasoning effort, which the website row label omits.

Snapshot

7 models10 configurations100 repositories4,068 reference bugsmini-SWE-agentDisplay only

The published SWE-sweep snapshot places GPT-5.6 Sol first at 4.7%. The third row is 3.1 points behind. The broader top-10 range is 4.6 points, so many of the published results sit in a relatively narrow band.

10 configurations have been evaluated on SWE-sweep. The benchmark falls in the Agentic category. SWE-sweep is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE-sweep

Year

2026

Tasks

4,068 reference bugs across 100 repositories in 22 languages

Format

Micro-average percentage of bugs resolved with a repository regression gate

Difficulty

Repository-wide bug discovery and repair

SWE-sweep evaluates bug discovery and repair on 4,068 reference bugs across 100 repositories in 22 languages. Agents receive the codebase without issue descriptions or bug-location hints and have up to 1,000 steps and 12 hours per repository. Hidden regression tests check the fixes, and a repository receives zero credit if the patch introduces a new failure in the restored test suite. We mirror the published mini-SWE-agent and reasoning-effort configurations as display-only evidence.

Freshness and provenance

Version

SWE-sweep 2026

Refresh cadence

Manual snapshot

Staleness state

Current

Question availability

Public repositories and evaluation framework; regression tests hidden from agents

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SWE-sweep measure?

Tests whether coding agents can find and fix multiple bugs across a repository without issue descriptions, while preserving existing behavior.

Which model leads the published SWE-sweep snapshot?

GPT-5.6 Sol currently leads the published SWE-sweep snapshot with 4.7% bugs resolved. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SWE-sweep?

The September 24, 2026 snapshot snapshot contains 10 configurations across 7 AI models.

Last updated: September 24, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.