# SWE-Atlas Refactoring

> A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

Canonical page: https://benchlm.ai/benchmarks/sweatlasrefactoring

- Category: [Agentic](/agentic)
- Last updated: September 11, 2026 snapshot

## About SWE-Atlas Refactoring

- Year: 2026
- Tasks: SWE-Atlas refactoring tasks
- Format: Refactoring score with confidence intervals
- Difficulty: Real-world software-engineering agent tasks
- Paper: [SWE-Atlas](https://labs.scale.com/papers/sweatlas)

BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard as a display-only agentic software-engineering benchmark. The source compares model-agent combinations such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.

SWE-Atlas Refactoring is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (17 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 59.0% |
| 2 | [Fable-5.1 (Claude Code) xHigh](https://labs.scale.com/leaderboard/sweatlas-refactoring) | Anthropic | 56.7% |
| 3 | [Fable-5 (Claude Code) xHigh](https://labs.scale.com/leaderboard/sweatlas-refactoring) | Anthropic | 54.8% |
| 4 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 48.6% |
| 5 | [Opus 4.8 (Claude Code)\n](https://labs.scale.com/leaderboard/sweatlas-refactoring) | Anthropic | 46.7% |
| 6 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 44.8% |
| 7 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 44.8% |
| 8 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 44.3% |
| 9 | [GLM-5.2](/models/glm-5-2) | Z.AI | 42.4% |
| 10 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 42.4% |
| 11 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 35.6% |
| 12 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 33.8% |
| 13 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 32.2% |
| 14 | [GLM-5](/models/glm-5) | Z.AI | 24.2% |
| 15 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 20.9% |
| 16 | [MiniMax M2.5](/models/minimax-m2-5) | MiniMax | 19.5% |
| 17 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 10.0% |

## FAQ

### What does SWE-Atlas Refactoring measure?

A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

### Which model leads the published SWE-Atlas Refactoring snapshot?

GPT-6 Astra currently leads the published SWE-Atlas Refactoring snapshot with a score of 59.0%.

### How many models are evaluated on SWE-Atlas Refactoring?

The September 11, 2026 snapshot contains 17 AI models.

### Does SWE-Atlas Refactoring affect BenchLM's overall score?

Not directly. SWE-Atlas Refactoring is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
