Benchmark profile
SWE-Atlas Refactoring
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
How BenchLM shows SWE-Atlas Refactoring
BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard from July 27, 2026 snapshot. The source reports 14 agent/model rows with confidence intervals and harness labels such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.
SWE-Atlas Refactoring is display only on BenchLM. It is useful evidence about software-engineering agents, but the rows mix base model quality with agent harness choices, so BenchLM keeps it out of weighted model-only rankings.
Refactoring score on SWE-Atlas Refactoring — July 27, 2026 snapshot
BenchLM mirrors the published refactoring score view for SWE-Atlas Refactoring. Fable-5 (Claude Code) xHigh leads the public snapshot at 54.8% , followed by Claude Opus 4.7 (Adaptive) (48.6%) and Opus 4.8 (Claude Code)\n (46.7%). BenchLM does not use these results to rank models overall.
Fable-5 (Claude Code) xHigh
Anthropic
Claude Opus 4.7 (Adaptive)
Anthropic
Opus-4.7 (Claude Code)
Opus 4.8 (Claude Code)\n
Anthropic
Refactoring score table (14 models)
ScoreThe published SWE-Atlas Refactoring snapshot places Fable-5 (Claude Code) xHigh first at 54.8%. The third row is 8.1 points behind. The broader top-10 range is 22.5 points, so the table still separates the published systems.
14 models have been evaluated on SWE-Atlas Refactoring. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. SWE-Atlas Refactoring is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About SWE-Atlas Refactoring
Year
2026
Tasks
SWE-Atlas refactoring tasks
Format
Refactoring score with confidence intervals
Difficulty
Real-world software-engineering agent tasks
BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard as a display-only agentic software-engineering benchmark. The source compares model-agent combinations such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.
BenchLM freshness & provenance
Version
SWE-Atlas Refactoring 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does SWE-Atlas Refactoring measure?
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
Which model leads the published SWE-Atlas Refactoring snapshot?
Fable-5 (Claude Code) xHigh currently leads the published SWE-Atlas Refactoring snapshot with 54.8% refactoring score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on SWE-Atlas Refactoring?
14 AI models are included in BenchLM's mirrored SWE-Atlas Refactoring snapshot, based on the public leaderboard captured on July 27, 2026 snapshot.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.