SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? (SWE Refactor Bench)
Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.
How to read this leaderboard
A nonzero composite means the run completed the requested migration and passed every frozen behavioral check. The score then rises with the number of six independent coding-agent verifiers that failed to find a behavioral difference.
Operator receipt: 26 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 at 47.0.
Honest limit: Each result belongs to a model, reasoning-effort level, agent client, and task budget. Treat the rows as complete-system configurations, not controlled measurements of model weights alone.
How we show SWE Refactor Bench
We mirror the benchmark-owned SWE Refactor Bench configuration table from August 25, 2026 snapshot. Every one of the 26 configurations ran all 20 whole-repository migrations, for 520 graded runs in total.
The composite awards credit only after the submitted repository passes a migration audit and every frozen behavioral check. Six coding agents then search for executable counterexamples. We keep each reasoning-effort and client configuration separate, and the table stays display only because those choices are part of the result.
Snapshot
Composite score on SWE Refactor Bench — August 25, 2026 snapshot
We mirror the published composite score view for SWE Refactor Bench. Claude Opus 5 leads the public snapshot at 47.0, followed by Claude Opus 5 (34.5) and Claude Opus 5 (31.0). We do not use these results to rank models overall.
Claude Opus 5
Anthropic
Claude Code · xhigh reasoning
$74.9 per task · 5 accepted · 6 broken · 1 blind
Claude Opus 5
Anthropic
Claude Code · high reasoning
$55.7 per task · 4 accepted · 4 broken · 2 blind
Claude Opus 5
Anthropic
Claude Code · max reasoning
$72.4 per task · 3 accepted · 4 broken · 2 blind
Composite score table (26 configurations)
ScoreThe published SWE Refactor Bench snapshot places Claude Opus 5 first at 47.0. The third row is 16.0 score units behind. The broader top-10 range is 36.5 score units, so the table still separates the published systems.
26 configurations have been evaluated on SWE Refactor Bench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. SWE Refactor Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About SWE Refactor Bench
Year
2026
Tasks
20 whole-repository stack migrations
Format
Migration audit, frozen behavioral checks, and agentic verification
Difficulty
6- to 30-hour autonomous repository migrations
SWE Refactor Bench contains 20 repository-wide migrations across languages, frameworks, platforms, and build systems. A run receives credit only after the migration audit confirms that the old stack is gone, every frozen behavioral check passes, and six coding agents fail to find an executable counterexample. We mirror all 26 model, client, and reasoning-effort configurations as display-only evidence.
BenchLM freshness & provenance
Version
SWE Refactor Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
20 public tasks with released score table and trajectories
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does SWE Refactor Bench measure?
Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.
Which model leads the published SWE Refactor Bench snapshot?
Claude Opus 5 currently leads the published SWE Refactor Bench snapshot with 47.0 composite score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on SWE Refactor Bench?
The August 25, 2026 snapshot contains 26 configurations across 8 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.