# FrontierSWE

> An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.

Canonical page: https://benchlm.ai/benchmarks/frontierswe

- Category: [Coding](/coding)
- Last updated: September 15, 2026

## About FrontierSWE

- Year: 2026
- Tasks: 17 ultra-long-horizon engineering and research tasks
- Format: Mean@5, best@5, average rank, and dominance
- Difficulty: Ultra-long-horizon frontier software engineering
- Paper: [FrontierSWE: Benchmarking coding agents at the limits of human abilities](https://www.frontierswe.com/blog)

FrontierSWE's first release contains 17 tasks across implementation, performance, and research. Agents receive up to 20 hours per task, run five trials, and are compared using mean@5, best@5, average rank, and dominance. Proximal superseded this release with the 34-task FrontierSWE v2 in September 2026 and archived the original table at frontierswe.com/v1; BenchLM keeps the two releases as separate keys. Moonshot reports Kimi K3's dominance score after recomputing it from raw scores with the official evaluation script.

FrontierSWE is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (17 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | Claude Code | Anthropic | 88.2% |
| 2 | [GLM-5.3](/models/glm-5-3) | Claude Code | Z.AI | 78.1% |
| 3 | [Grok 4.6](/models/grok-4-6) | Grok CLI | xAI | 77.9% |
| 4 | [Grok 4.5](/models/grok-4-5) | Grok CLI | xAI | 72.1% |
| 5 | [GLM-5.2](/models/glm-5-2) | Claude Code | Z.AI | 67.5% |
| 6 | [Claude Opus 4.8](/models/claude-opus-4-8) | Claude Code | Anthropic | 66.5% |
| 7 | [GPT-5.5](/models/gpt-5-5) | Codex | OpenAI | 64.5% |
| 8 | [Claude Opus 4.7](/models/claude-opus-4-7) | Claude Code | Anthropic | 56.3% |
| 9 | [Claude Opus 4.6](/models/claude-opus-4-6) | Claude Code | Anthropic | 48.9% |
| 10 | [GPT-5.4](/models/gpt-5-4) | Codex | OpenAI | 46.0% |
| 11 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Gemini CLI | Google | 34.4% |
| 12 | [Composer 2.5](/models/composer-2-5) | Cursor CLI | Cursor | 33.8% |
| 13 | [GLM-5.1](/models/glm-5-1) | Claude Code | Z.AI | 25.7% |
| 14 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | Claude Code | DeepSeek | 24.6% |
| 15 | [Kimi K2.5](/models/kimi-k2-5) | Kimi CLI | Moonshot AI | 23.3% |
| 16 | [Kimi K2.6](/models/kimi-2-6) | Kimi CLI | Moonshot AI | 22.2% |
| 17 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Qwen Code | Alibaba | 19.9% |

## FAQ

### What does FrontierSWE measure?

An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.

### Which model leads the published FrontierSWE snapshot?

Claude Fable 5 currently leads the published FrontierSWE snapshot with a score of 88.2%.

### How many models are evaluated on FrontierSWE?

The September 15, 2026 contains 17 AI models.

### Does FrontierSWE affect BenchLM's overall score?

Not directly. FrontierSWE is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
