SWE-Atlas Refactoring
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
How BenchLM shows SWE-Atlas Refactoring
BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard from September 11, 2026 snapshot. The source reports 17 agent/model rows with confidence intervals and harness labels such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.
SWE-Atlas Refactoring is display only on BenchLM. It is useful evidence about software-engineering agents, but the rows mix base model quality with agent harness choices, so BenchLM keeps it out of weighted model-only rankings.
Snapshot
Refactoring score on SWE-Atlas Refactoring — September 11, 2026 snapshot
We mirror the published refactoring score view for SWE-Atlas Refactoring. GPT-6 Astra leads the public snapshot at 59.0%, followed by Fable-5.1 (Claude Code) xHigh (56.7%) and Fable-5 (Claude Code) xHigh (54.8%). We do not use these results to rank models overall.
GPT-6 Astra
OpenAI
Fable-5.1 (Claude Code) xHigh
Anthropic
Fable-5 (Claude Code) xHigh
Anthropic
17 modelsAgenticCurrentDisplay onlyUpdated September 11, 2026 snapshot
Refactoring score table (17 models)
ScoreThe published SWE-Atlas Refactoring snapshot places GPT-6 Astra first at 59.0%. The third row is 4.3 points behind. The broader top-10 range is 16.7 points, so the table still separates the published systems.
17 models have been evaluated on SWE-Atlas Refactoring. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. SWE-Atlas Refactoring is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About SWE-Atlas Refactoring
Year
2026
Tasks
SWE-Atlas refactoring tasks
Format
Refactoring score with confidence intervals
Difficulty
Real-world software-engineering agent tasks
BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard as a display-only agentic software-engineering benchmark. The source compares model-agent combinations such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.
BenchLM freshness & provenance
Version
SWE-Atlas Refactoring 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does SWE-Atlas Refactoring measure?
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
Which model leads the published SWE-Atlas Refactoring snapshot?
GPT-6 Astra currently leads the published SWE-Atlas Refactoring snapshot with 59.0% refactoring score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on SWE-Atlas Refactoring?
The September 11, 2026 snapshot snapshot contains 17 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.