# SWE-sweep: Can Agents Autonomously Find and Fix Bugs? (SWE-sweep)

> Tests whether coding agents can find and fix multiple bugs across a repository without issue descriptions, while preserving existing behavior.

Canonical page: https://benchlm.ai/benchmarks/swe-sweep

- Category: [Agentic](/agentic)
- Last updated: September 24, 2026 snapshot

## About SWE-sweep

- Year: 2026
- Tasks: 4,068 reference bugs across 100 repositories in 22 languages
- Format: Micro-average percentage of bugs resolved with a repository regression gate
- Difficulty: Repository-wide bug discovery and repair
- Paper: [SWE-sweep: Can Agents Autonomously Find and Fix Bugs?](https://swesweep.com/paper)

SWE-sweep evaluates bug discovery and repair on 4,068 reference bugs across 100 repositories in 22 languages. Agents receive the codebase without issue descriptions or bug-location hints and have up to 1,000 steps and 12 hours per repository. Hidden regression tests check the fixes, and a repository receives zero credit if the patch introduces a new failure in the restored test suite. We mirror the published mini-SWE-agent and reasoning-effort configurations as display-only evidence.

SWE-sweep is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## How we show SWE-sweep

We mirror the official SWE-sweep table from the September 24, 2026 snapshot, retrieved on October 2, 2026. The source reports 10 mini-SWE-agent configurations on 100 repositories, scored by the percentage of reference bugs resolved across the full set.

Agents must find and fix bugs without issue descriptions. Hidden tests check each fix. A new failure in the restored original test suite sets that repository's score to zero.

Each reasoning-effort setting stays on its own row. The table is display only and excluded from overall and category rankings because its scores belong to the published model, effort, and agent setup.

The table averages cost, turns, and output tokens per evaluated repository, including non-deprecated retries. The paper supplies Kimi K3’s max reasoning effort, which the website row label omits.

Sources: [SWE-sweep leaderboard](https://swesweep.com/), [SWE-sweep paper](https://swesweep.com/paper), [GitHub repository](https://github.com/facebookresearch/swe-sweep).

## Leaderboard (10 configurations)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | mini-SWE-agent · xhigh reasoning · $72.30 per repository · 232.54 turns · 104,813.27 output tokens | OpenAI | 4.7% |
| 2 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | mini-SWE-agent · xhigh reasoning · $2.24 per repository · 204.13 turns · 61,210.82 output tokens | OpenAI | 2.5% |
| 3 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | mini-SWE-agent · xhigh reasoning · $3.57 per repository · 76.32 turns · 46,647.27 output tokens | OpenAI | 1.5% |
| 4 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | mini-SWE-agent · high reasoning · $0.28 per repository · 74.58 turns · 20,288.3 output tokens | OpenAI | 1.4% |
| 5 | [Claude Opus 5](/models/claude-opus-5) | mini-SWE-agent · xhigh reasoning · $53.63 per repository · 322.78 turns · 192,749.15 output tokens | Anthropic | 1.3% |
| 6 | [Kimi K3](/models/kimi-k3) | mini-SWE-agent · max reasoning · $24.51 per repository · 337.27 turns · 152,612.73 output tokens | Moonshot AI | 0.6% |
| 7 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | mini-SWE-agent · default reasoning · $0.04 per repository · 22.49 turns · 4,465.53 output tokens | OpenAI | 0.5% |
| 8 | [GPT-5.4 mini](/models/gpt-5-4-mini) | mini-SWE-agent · high reasoning · $1.22 per repository · 75.04 turns · 40,513.27 output tokens | OpenAI | 0.5% |
| 9 | [GPT-5.4 mini](/models/gpt-5-4-mini) | mini-SWE-agent · default reasoning · $0.05 per repository · 12.68 turns · 2,147.18 output tokens | OpenAI | 0.2% |
| 10 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | mini-SWE-agent · default reasoning · $0.06 per repository · 33.70 turns · 5,228.24 output tokens | Google | 0.1% |

## FAQ

### What does SWE-sweep measure?

Tests whether coding agents can find and fix multiple bugs across a repository without issue descriptions, while preserving existing behavior.

### Which model leads the published SWE-sweep snapshot?

GPT-5.6 Sol currently leads the published SWE-sweep snapshot with a score of 4.7%.

### How many models are evaluated on SWE-sweep?

The September 24, 2026 snapshot contains 10 configurations across 7 AI models.

### Does SWE-sweep affect BenchLM's overall score?

Not directly. SWE-sweep is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
