FrontierSWE v2
A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.
Previous release
The original 17-task FrontierSWE remains available as an archived table. It ranks model-and-harness pairs by dominance rather than mean task score, so movement between the two tables is not a like-for-like model comparison.
View the original FrontierSWEHow we show FrontierSWE v2
We mirror Proximal's live Mean@5 table across all 34 v2 tasks. Every model runs in Proximal's own Proximus harness at maximum reasoning effort, with 5 trials per task and a 20-hour budget. Scores are the mean task score on a 0-100 scale, not the dominance metric used by the original release.
Rows on this page are benchmark-owned. Models that only appear in provider material, such as the Claude Opus 5 and Claude Fable 5 results in Anthropic's Fable 5.1 system card, keep provider-exact rows on their model pages and do not appear here.
Snapshot
Mean@5 score on FrontierSWE v2 — September 3, 2026
We mirror the published mean@5 score view for FrontierSWE v2. Claude Fable 5.1 leads the public snapshot at 56.3%, followed by GPT-5.6 Sol (32.2%) and GLM-5.3 (30.2%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
proximus
GPT-5.6 Sol
OpenAI
proximus
GLM-5.3
Z.AI
proximus
10 modelsCodingCurrentDisplay onlyUpdated September 3, 2026
Mean@5 score table (10 models)
ScoreThe published FrontierSWE v2 snapshot places Claude Fable 5.1 first at 56.3%. The third row is 26.1 points behind. The broader top-10 range is 52.2 points, so the table still separates the published systems.
10 models have been evaluated on FrontierSWE v2. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. FrontierSWE v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About FrontierSWE v2
Year
2026
Tasks
34 ultra-long-horizon engineering and research tasks
Format
Five-trial mean task score (Mean@5), 0-100
Difficulty
Ultra-long-horizon frontier software engineering
FrontierSWE v2 keeps 13 of the original 17 tasks, retires four as saturated or unreliably scorable, and adds 21 new tasks across visual reasoning, scientific computing, and AI research. Proximal runs every model at maximum reasoning effort in its own Proximus harness with a 20-hour budget and five trials per task, and publishes mean, best, and worst scores on a 0-100 scale. BenchLM mirrors the public Mean@5 table as benchmark-owned rows; models that only appear in provider material stay provider-exact on the same scale.
BenchLM freshness & provenance
Version
FrontierSWE v2 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does FrontierSWE v2 measure?
A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.
Which model leads the published FrontierSWE v2 snapshot?
Claude Fable 5.1 currently leads the published FrontierSWE v2 snapshot with 56.3% mean@5 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on FrontierSWE v2?
The September 3, 2026 snapshot contains 10 AI models.
How is FrontierSWE v2 different from the original FrontierSWE?
The original release had 17 tasks. Version 2 keeps 13 of them, retires four that Proximal found saturated or impossible to score deterministically (PCQM4Mv2 gap prediction, Pyright type-checking optimization, Revideo rendering optimization, and the dependent type checker), and adds 21 new tasks across visual reasoning, scientific computing, and AI research, for 34 in total. Every model now runs in Proximal's Proximus harness, and each task reports a normalized 0-100 score instead of a dominance rank.
Can FrontierSWE v2 scores be compared with the original FrontierSWE?
No. The two releases are not comparable: the task set changed, the harness changed from each model's native agent to Proximus, the headline metric changed from dominance to mean task score, and performance tasks are now measured by a weighted instruction count on a pinned CPU microarchitecture rather than wall-clock time in shared sandboxes. BenchLM keeps the original 17-task table on its own page for that reason.
What is the Proximus harness?
Proximus is the minimal coding-agent harness Proximal built for ultra-long-horizon tasks. It compacts context by summarizing the trajectory, supports images and plots, and lets a model record a clean workspace as a candidate submission while reporting how much of the 20-hour budget remains, so models keep working instead of submitting early. Proximal reports that Claude Opus 5 and especially GPT-5.6 Sol score higher under Proximus than under their native harnesses.
How does FrontierSWE v2 guard against gaming the verifier?
Each workspace ships a documented self-check tool whose output tracks the verifier score, so models can measure progress honestly. Verification then applies one-to-one structural test mutation to defeat hardcoded answers, and it runs in a fresh verifier container booted from a pinned image after the agent's container is stopped. The agent also runs as a non-root user with root-owned protected directories.
Where do the Claude Opus 5 and Claude Fable 5 results come from?
Proximal's public v2 table does not list those two models. Their scores come from Anthropic's Claude Fable 5.1 and Claude Mythos 5.1 system card, which reports Proximal's own five-trial runs. BenchLM shows them on the model pages as provider-exact rows on the same 0-100 scale but keeps them off this benchmark-owned leaderboard.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.