# Scale Labs SWE Atlas – Test Writing (SWE Atlas – Test Writing)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-sweatlas-test-writing

- Category: [Coding](/coding)
- Last updated: September 29, 2026 snapshot

## About SWE Atlas – Test Writing

- Year: 2026
- Tasks: 24 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/sweatlas-tw)

BenchLM mirrors 24 published rows from the SWE Atlas – Test Writing public table captured on September 29, 2026 snapshot.

SWE Atlas – Test Writing is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (24 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Fable-5.1 (Claude Code) xHigh*](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 67.04% |
| 2 | [Opus 5 (Claude Code) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 62.22% |
| 3 | [Fable-5 (Claude Code) xHigh*](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 55.60% |
| 4 | [Gemini 3.8 Flash (Mini-SWE-Agent)](https://labs.scale.com/leaderboard/sweatlas-tw) | Google | 53.70% |
| 5 | [GPT 6 Astra (Codex) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | OpenAI | 50.74% |
| 6 | [Opus 4.8 (Claude Code) xhigh](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 49.63% |
| 7 | [GPT-5.6-Sol (Codex) xHigh*\n](https://labs.scale.com/leaderboard/sweatlas-tw) | OpenAI | 45.90% |
| 8 | [GPT-5.4 (Codex) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | OpenAI | 44.36% |
| 9 | [GPT-5.5 (Codex) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | OpenAI | 42.59% |
| 10 | [GLM 5.2 (Mini-SWE-Agent)](https://labs.scale.com/leaderboard/sweatlas-tw) | Z.AI | 41.48% |
| 11 | [Muse Spark 1.1 (Mini-SWE-Agent) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | Meta | 41.48% |
| 12 | [GPT-5.4 (Mini-SWE) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | OpenAI | 40.00% |
| 13 | [GPT 5.3 (Codex) xHigh](https://labs.scale.com/leaderboard/sweatlas-tw) | OpenAI | 38.98% |
| 14 | [Opus 4.7 (Claude Code)](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 38.52% |
| 15 | [Opus-4.6 (Claude Code)](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 36.67% |
| 16 | [Opus-4.6 (Mini-SWE)](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 36.08% |
| 17 | [Sonnet-4.6 (Claude Code)](https://labs.scale.com/leaderboard/sweatlas-tw) | Anthropic | 31.76% |
| 18 | [Muse Spark](/models/muse-spark) | Meta | 31.11% |
| 19 | [Gemini-3-Flash (Mini-SWE)](https://labs.scale.com/leaderboard/sweatlas-tw) | Google | 30.30% |
| 20 | [Gemini-3.1-Pro (Mini-SWE)](https://labs.scale.com/leaderboard/sweatlas-tw) | Google | 29.84% |
| 21 | [Glm-5 (Mini-SWE)](https://labs.scale.com/leaderboard/sweatlas-tw) | Z.AI | 28.74% |
| 22 | [Deepseek V4 Pro (Mini-SWE-Agent)](https://labs.scale.com/leaderboard/sweatlas-tw) | Deepseek | 27.05% |
| 23 | [Kimi-K2.5 (Mini-SWE)](https://labs.scale.com/leaderboard/sweatlas-tw) | Moonshot AI | 25.77% |
| 24 | [Minimax-M2.5 (Mini-SWE)](https://labs.scale.com/leaderboard/sweatlas-tw) | MiniMax | 18.60% |

## FAQ

### What does SWE Atlas – Test Writing measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published SWE Atlas – Test Writing snapshot?

Fable-5.1 (Claude Code) xHigh* currently leads the published SWE Atlas – Test Writing snapshot with a score of 67.04%.

### How many models are evaluated on SWE Atlas – Test Writing?

The September 29, 2026 snapshot contains 24 AI models.

### Does SWE Atlas – Test Writing affect BenchLM's overall score?

Not directly. SWE Atlas – Test Writing is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
