SWE-sweep: Can Agents Autonomously Find and Fix Bugs? (SWE-sweep)
We show this table for reference; we do not rank on it.
Tests whether coding agents can find and fix multiple bugs across a repository without issue descriptions, while preserving existing behavior.
Bugs resolved on SWE-sweep — September 24, 2026 snapshot
We mirror the published bugs resolved view for SWE-sweep. GPT-5.6 Sol leads the public snapshot at 4.7%, followed by GPT-5.6 Luna (2.5%) and GPT-5.6 Terra (1.5%). We do not use these results to rank models overall.
GPT-5.6 Sol
OpenAI
mini-SWE-agent · xhigh reasoning
$72.30 per repository · 232.54 turns · 104,813.27 output tokens
GPT-5.6 Luna
OpenAI
mini-SWE-agent · xhigh reasoning
$2.24 per repository · 204.13 turns · 61,210.82 output tokens
GPT-5.6 Terra
OpenAI
mini-SWE-agent · xhigh reasoning
$3.57 per repository · 76.32 turns · 46,647.27 output tokens
10 configurationsAgenticCurrentDisplay onlyUpdated September 24, 2026 snapshot
Bugs resolved table (10 configurations)
ScoreHow to read this leaderboard
The score divides resolved reference bugs by all reference bugs across the repository set. Each bug has equal weight, so repositories with more reference bugs contribute more to the micro-average. A repository where the patch introduces a new failure in the restored original test suite contributes no resolved bugs.
Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is GPT-5.6 Sol at 4.7%.
Honest limit: Each row measures a model, reasoning effort, and mini-SWE-agent setup. The reference-bug set and regression tests define which repairs receive credit; the result does not measure every possible defect in a repository.
How we show SWE-sweep
We mirror the official SWE-sweep table from the September 24, 2026 snapshot, retrieved on October 2, 2026. The source reports 10 mini-SWE-agent configurations on 100 repositories, scored by the percentage of reference bugs resolved across the full set.
Agents must find and fix bugs without issue descriptions. Hidden tests check each fix. A new failure in the restored original test suite sets that repository's score to zero.
Each reasoning-effort setting stays on its own row. The table is display only and excluded from overall and category rankings because its scores belong to the published model, effort, and agent setup.
The table averages cost, turns, and output tokens per evaluated repository, including non-deprecated retries. The paper supplies Kimi K3’s max reasoning effort, which the website row label omits.
Snapshot
The published SWE-sweep snapshot places GPT-5.6 Sol first at 4.7%. The third row is 3.1 points behind. The broader top-10 range is 4.6 points, so many of the published results sit in a relatively narrow band.
10 configurations have been evaluated on SWE-sweep. The benchmark falls in the Agentic category. SWE-sweep is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About SWE-sweep
Year
2026
Tasks
4,068 reference bugs across 100 repositories in 22 languages
Format
Micro-average percentage of bugs resolved with a repository regression gate
Difficulty
Repository-wide bug discovery and repair
SWE-sweep evaluates bug discovery and repair on 4,068 reference bugs across 100 repositories in 22 languages. Agents receive the codebase without issue descriptions or bug-location hints and have up to 1,000 steps and 12 hours per repository. Hidden regression tests check the fixes, and a repository receives zero credit if the patch introduces a new failure in the restored test suite. We mirror the published mini-SWE-agent and reasoning-effort configurations as display-only evidence.
Freshness and provenance
Version
SWE-sweep 2026
Refresh cadence
Manual snapshot
Staleness state
Current
Question availability
Public repositories and evaluation framework; regression tests hidden from agents
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does SWE-sweep measure?
Tests whether coding agents can find and fix multiple bugs across a repository without issue descriptions, while preserving existing behavior.
Which model leads the published SWE-sweep snapshot?
GPT-5.6 Sol currently leads the published SWE-sweep snapshot with 4.7% bugs resolved. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on SWE-sweep?
The September 24, 2026 snapshot snapshot contains 10 configurations across 7 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.