FrontierSWE
An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.
Current release
FrontierSWE v2 is the current release. It expands the suite to 34 tasks, runs every model in Proximal's own Proximus harness, and reports a mean task score instead of dominance, so its scores are not directly comparable with this original 17-task table.
View FrontierSWE v2How we show the original FrontierSWE
We mirror Proximal's archived Mean@5 dominance table for the original 17-task release. Each row preserves the agent harness and the task-category rank breakdown published by the benchmark owner.
Proximal superseded this release with the 34-task FrontierSWE v2 in September 2026 and moved the original table to frontierswe.com/v1. The two releases use different task sets, harnesses, and score scales, so BenchLM keeps them on separate pages.
Snapshot
Mean@5 dominance on FrontierSWE — September 15, 2026
We mirror the published mean@5 dominance view for FrontierSWE. Claude Fable 5 leads the public snapshot at 88.2%, followed by GLM-5.3 (78.1%) and Grok 4.6 (77.9%). We do not use these results to rank models overall.
Claude Fable 5
Anthropic
Claude Code
GLM-5.3
Z.AI
Claude Code
Grok 4.6
xAI
Grok CLI
17 modelsCodingCurrentDisplay onlyUpdated September 15, 2026
Mean@5 dominance table (17 models)
ScoreThe published FrontierSWE snapshot places Claude Fable 5 first at 88.2%. The third row is 10.3 points behind. The broader top-10 range is 42.2 points, so the table still separates the published systems.
17 models have been evaluated on FrontierSWE. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. FrontierSWE is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About FrontierSWE
Year
2026
Tasks
17 ultra-long-horizon engineering and research tasks
Format
Mean@5, best@5, average rank, and dominance
Difficulty
Ultra-long-horizon frontier software engineering
FrontierSWE's first release contains 17 tasks across implementation, performance, and research. Agents receive up to 20 hours per task, run five trials, and are compared using mean@5, best@5, average rank, and dominance. Proximal superseded this release with the 34-task FrontierSWE v2 in September 2026 and archived the original table at frontierswe.com/v1; BenchLM keeps the two releases as separate keys. Moonshot reports Kimi K3's dominance score after recomputing it from raw scores with the official evaluation script.
BenchLM freshness & provenance
Version
FrontierSWE 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does FrontierSWE measure?
An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.
Which model leads the published FrontierSWE snapshot?
Claude Fable 5 currently leads the published FrontierSWE snapshot with 88.2% mean@5 dominance. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on FrontierSWE?
The September 15, 2026 snapshot contains 17 AI models.
Is the original FrontierSWE still updated?
No. Proximal superseded this release with FrontierSWE v2 in September 2026 and archived the original table at frontierswe.com/v1. New models are evaluated on v2, so this page is a historical snapshot of the 17-task release. BenchLM keeps mirroring the archived table so earlier results stay traceable.
How is this table different from FrontierSWE v2?
This table ranks model-and-harness pairs by dominance, the share of other rows a pair beats, across the original 17 tasks. FrontierSWE v2 keeps 13 of those tasks, retires four, adds 21 new ones for a total of 34, runs every model in Proximal's own Proximus harness, and reports a mean task score on a 0-100 scale.
Can I compare these scores with FrontierSWE v2?
Not directly. The task set, the harness, and the metric all changed between releases, and v2 also measures performance tasks by a weighted instruction count instead of wall-clock time. A model can move between the two tables without its underlying capability changing, so treat them as separate benchmarks.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.