Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

FrontierSWE

An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.

Current release

FrontierSWE v2 is the current release. It expands the suite to 34 tasks, runs every model in Proximal's own Proximus harness, and reports a mean task score instead of dominance, so its scores are not directly comparable with this original 17-task table.

View FrontierSWE v2
Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

How we show the original FrontierSWE

We mirror Proximal's archived Mean@5 dominance table for the original 17-task release. Each row preserves the agent harness and the task-category rank breakdown published by the benchmark owner.

Proximal superseded this release with the 34-task FrontierSWE v2 in September 2026 and moved the original table to frontierswe.com/v1. The two releases use different task sets, harnesses, and score scales, so BenchLM keeps them on separate pages.

Snapshot

17 model-agent rows17 public tasks5 trials per taskMean@5 dominanceArchived v1Display only

Mean@5 dominance on FrontierSWE — September 15, 2026

We mirror the published mean@5 dominance view for FrontierSWE. Claude Fable 5 leads the public snapshot at 88.2%, followed by GLM-5.3 (78.1%) and Grok 4.6 (77.9%). We do not use these results to rank models overall.

17 modelsCodingCurrentDisplay onlyUpdated September 15, 2026

Mean@5 dominance table (17 models)

Score
1
Claude Fable 5Anthropic · ClosedClaude Code
88.2%
2
GLM-5.3Z.AI · Open weightClaude Code
78.1%
3
Grok 4.6xAI · ClosedGrok CLI
77.9%
4
Grok 4.5xAI · ClosedGrok CLI
72.1%
5
GLM-5.2Z.AI · Open weightClaude Code
67.5%
6
Claude Opus 4.8Anthropic · ClosedClaude Code
66.5%
7
GPT-5.5OpenAI · ClosedCodex
64.5%
8
Claude Opus 4.7Anthropic · ClosedClaude Code
56.3%
9
Claude Opus 4.6Anthropic · ClosedClaude Code
48.9%
10
GPT-5.4OpenAI · ClosedCodex
46.0%
11
Gemini 3.1 ProGoogle · ClosedGemini CLI
34.4%
12
Composer 2.5Cursor · ClosedCursor CLI
33.8%
13
GLM-5.1Z.AI · Open weightClaude Code
25.7%
14
DeepSeek V4 Pro 0813DeepSeek · ClosedClaude Code
24.6%
15
Kimi K2.5Moonshot AI · Open weightKimi CLI
23.3%
16
Kimi K2.6Moonshot AI · Open weightKimi CLI
22.2%
17
Qwen3.6 PlusAlibaba · ClosedQwen Code
19.9%

The published FrontierSWE snapshot places Claude Fable 5 first at 88.2%. The third row is 10.3 points behind. The broader top-10 range is 42.2 points, so the table still separates the published systems.

17 models have been evaluated on FrontierSWE. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. FrontierSWE is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About FrontierSWE

Year

2026

Tasks

17 ultra-long-horizon engineering and research tasks

Format

Mean@5, best@5, average rank, and dominance

Difficulty

Ultra-long-horizon frontier software engineering

FrontierSWE's first release contains 17 tasks across implementation, performance, and research. Agents receive up to 20 hours per task, run five trials, and are compared using mean@5, best@5, average rank, and dominance. Proximal superseded this release with the 34-task FrontierSWE v2 in September 2026 and archived the original table at frontierswe.com/v1; BenchLM keeps the two releases as separate keys. Moonshot reports Kimi K3's dominance score after recomputing it from raw scores with the official evaluation script.

BenchLM freshness & provenance

Version

FrontierSWE 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does FrontierSWE measure?

An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.

Which model leads the published FrontierSWE snapshot?

Claude Fable 5 currently leads the published FrontierSWE snapshot with 88.2% mean@5 dominance. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on FrontierSWE?

The September 15, 2026 snapshot contains 17 AI models.

Is the original FrontierSWE still updated?

No. Proximal superseded this release with FrontierSWE v2 in September 2026 and archived the original table at frontierswe.com/v1. New models are evaluated on v2, so this page is a historical snapshot of the 17-task release. BenchLM keeps mirroring the archived table so earlier results stay traceable.

How is this table different from FrontierSWE v2?

This table ranks model-and-harness pairs by dominance, the share of other rows a pair beats, across the original 17 tasks. FrontierSWE v2 keeps 13 of those tasks, retires four, adds 21 new ones for a total of 34, runs every model in Proximal's own Proximus harness, and reports a mean task score on a 0-100 scale.

Can I compare these scores with FrontierSWE v2?

Not directly. The task set, the harness, and the metric all changed between releases, and v2 also measures performance tasks by a weighted instruction count instead of wall-clock time. A model can move between the two tables without its underlying capability changing, so treat them as separate benchmarks.

Last updated: September 15, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.