Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

SWE-Atlas Refactoring

A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

How BenchLM shows SWE-Atlas Refactoring

BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard from September 11, 2026 snapshot. The source reports 17 agent/model rows with confidence intervals and harness labels such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.

SWE-Atlas Refactoring is display only on BenchLM. It is useful evidence about software-engineering agents, but the rows mix base model quality with agent harness choices, so BenchLM keeps it out of weighted model-only rankings.

Snapshot

17 model variantsSWE-Atlas task familyRefactoring scoreScale Labs sourceDisplay only

Refactoring score on SWE-Atlas Refactoring — September 11, 2026 snapshot

We mirror the published refactoring score view for SWE-Atlas Refactoring. GPT-6 Astra leads the public snapshot at 59.0%, followed by Fable-5.1 (Claude Code) xHigh (56.7%) and Fable-5 (Claude Code) xHigh (54.8%). We do not use these results to rank models overall.

17 modelsAgenticCurrentDisplay onlyUpdated September 11, 2026 snapshot

Refactoring score table (17 models)

Score
1
GPT-6 AstraOpenAI · Closed
59.0%
4
Claude Opus 4.7 (Adaptive)Anthropic · Closed
48.6%
6
GPT-5.5OpenAI · Closed
44.8%
7
Gemini 3.8 FlashGoogle · Closed
44.8%
8
GPT-5.4OpenAI · Closed
44.3%
9
GLM-5.2Z.AI · Open weight
42.4%
10
GPT-5.3 CodexOpenAI · Closed
42.4%
11
Claude Opus 4.6Anthropic · Closed
35.6%
12
Gemini 3.1 ProGoogle · Closed
33.8%
13
Claude Sonnet 4.6Anthropic · Closed
32.2%
14
GLM-5Z.AI · Open weight
24.2%
15
Kimi K2.5Moonshot AI · Open weight
20.9%
16
MiniMax M2.5MiniMax · Closed
19.5%
17
Gemini 3 FlashGoogle · Closed
10.0%

The published SWE-Atlas Refactoring snapshot places GPT-6 Astra first at 59.0%. The third row is 4.3 points behind. The broader top-10 range is 16.7 points, so the table still separates the published systems.

17 models have been evaluated on SWE-Atlas Refactoring. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. SWE-Atlas Refactoring is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE-Atlas Refactoring

Year

2026

Tasks

SWE-Atlas refactoring tasks

Format

Refactoring score with confidence intervals

Difficulty

Real-world software-engineering agent tasks

BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard as a display-only agentic software-engineering benchmark. The source compares model-agent combinations such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.

BenchLM freshness & provenance

Version

SWE-Atlas Refactoring 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE-Atlas Refactoring measure?

A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

Which model leads the published SWE-Atlas Refactoring snapshot?

GPT-6 Astra currently leads the published SWE-Atlas Refactoring snapshot with 59.0% refactoring score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SWE-Atlas Refactoring?

The September 11, 2026 snapshot snapshot contains 17 AI models.

Last updated: September 11, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.