Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

FrontierSWE v2

A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.

Previous release

The original 17-task FrontierSWE remains available as an archived table. It ranks model-and-harness pairs by dominance rather than mean task score, so movement between the two tables is not a like-for-like model comparison.

View the original FrontierSWE
Data verified 27 confirmed releases in the last 30 daysSee the free Radar Brief

How we show FrontierSWE v2

We mirror Proximal's live Mean@5 table across all 34 v2 tasks. Every model runs in Proximal's own Proximus harness at maximum reasoning effort, with 5 trials per task and a 20-hour budget. Scores are the mean task score on a 0-100 scale, not the dominance metric used by the original release.

Rows on this page are benchmark-owned. Models that only appear in provider material, such as the Claude Opus 5 and Claude Fable 5 results in Anthropic's Fable 5.1 system card, keep provider-exact rows on their model pages and do not appear here.

Snapshot

10 model rows34 public tasks5 trials per task20-hour budgetMean@5 scoreDisplay only

Mean@5 score on FrontierSWE v2 — September 3, 2026

We mirror the published mean@5 score view for FrontierSWE v2. Claude Fable 5.1 leads the public snapshot at 56.3%, followed by GPT-5.6 Sol (32.2%) and GLM-5.3 (30.2%). We do not use these results to rank models overall.

10 modelsCodingCurrentDisplay onlyUpdated September 3, 2026

Mean@5 score table (10 models)

Score
1
Claude Fable 5.1Anthropic · Closedproximus
56.3%
2
GPT-5.6 SolOpenAI · Closedproximus
32.2%
3
GLM-5.3Z.AI · Open weightproximus
30.2%
4
Kimi K3Moonshot AI · Closedproximus
25.9%
5
Grok 4.6xAI · Closedproximus
25.3%
6
Gemini 3.7 FlashGoogle · Closedproximus
20.3%
7
Qwen3.8 MaxAlibaba · Open weightproximus
15.8%
8
DeepSeek V4 Flash ExpDeepSeekproximus
14.8%
9
Muse Spark 1.2Meta · Closedproximus
12.0%
10
InklingThinking Machines Lab · Open weightproximus
4.1%

The published FrontierSWE v2 snapshot places Claude Fable 5.1 first at 56.3%. The third row is 26.1 points behind. The broader top-10 range is 52.2 points, so the table still separates the published systems.

10 models have been evaluated on FrontierSWE v2. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. FrontierSWE v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About FrontierSWE v2

Year

2026

Tasks

34 ultra-long-horizon engineering and research tasks

Format

Five-trial mean task score (Mean@5), 0-100

Difficulty

Ultra-long-horizon frontier software engineering

FrontierSWE v2 keeps 13 of the original 17 tasks, retires four as saturated or unreliably scorable, and adds 21 new tasks across visual reasoning, scientific computing, and AI research. Proximal runs every model at maximum reasoning effort in its own Proximus harness with a 20-hour budget and five trials per task, and publishes mean, best, and worst scores on a 0-100 scale. BenchLM mirrors the public Mean@5 table as benchmark-owned rows; models that only appear in provider material stay provider-exact on the same scale.

BenchLM freshness & provenance

Version

FrontierSWE v2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does FrontierSWE v2 measure?

A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.

Which model leads the published FrontierSWE v2 snapshot?

Claude Fable 5.1 currently leads the published FrontierSWE v2 snapshot with 56.3% mean@5 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on FrontierSWE v2?

The September 3, 2026 snapshot contains 10 AI models.

How is FrontierSWE v2 different from the original FrontierSWE?

The original release had 17 tasks. Version 2 keeps 13 of them, retires four that Proximal found saturated or impossible to score deterministically (PCQM4Mv2 gap prediction, Pyright type-checking optimization, Revideo rendering optimization, and the dependent type checker), and adds 21 new tasks across visual reasoning, scientific computing, and AI research, for 34 in total. Every model now runs in Proximal's Proximus harness, and each task reports a normalized 0-100 score instead of a dominance rank.

Can FrontierSWE v2 scores be compared with the original FrontierSWE?

No. The two releases are not comparable: the task set changed, the harness changed from each model's native agent to Proximus, the headline metric changed from dominance to mean task score, and performance tasks are now measured by a weighted instruction count on a pinned CPU microarchitecture rather than wall-clock time in shared sandboxes. BenchLM keeps the original 17-task table on its own page for that reason.

What is the Proximus harness?

Proximus is the minimal coding-agent harness Proximal built for ultra-long-horizon tasks. It compacts context by summarizing the trajectory, supports images and plots, and lets a model record a clean workspace as a candidate submission while reporting how much of the 20-hour budget remains, so models keep working instead of submitting early. Proximal reports that Claude Opus 5 and especially GPT-5.6 Sol score higher under Proximus than under their native harnesses.

How does FrontierSWE v2 guard against gaming the verifier?

Each workspace ships a documented self-check tool whose output tracks the verifier score, so models can measure progress honestly. Verification then applies one-to-one structural test mutation to defeat hardcoded answers, and it runs in a fresh verifier container booted from a pinned image after the agent's container is stopped. The agent also runs as a non-root user with root-owned protected directories.

Where do the Claude Opus 5 and Claude Fable 5 results come from?

Proximal's public v2 table does not list those two models. Their scores come from Anthropic's Claude Fable 5.1 and Claude Mythos 5.1 system card, which reports Proximal's own five-trial runs. BenchLM shows them on the model pages as provider-exact rows on the same 0-100 scale but keeps them off this benchmark-owned leaderboard.

Last updated: September 3, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.