Benchmark profile
SWE-bench Pro
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
Data verifiedClaude Mythos 5 leads the SWE-bench Pro leaderboard on BenchLM's July 2026 update with 80.3%, ahead of Claude Fable 5 (80%) and Claude Opus 5 (79.2%), across 54 tracked models.
How to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use SWE-bench Pro as evidence for long-horizon repository work only after matching the split, scaffold, tool budget, retry policy, and run count. The table preserves exact published rows from different providers, so a small score gap is directional unless the setups match.
Operator receipt: 54 sourced rows are currently displayable on this page; the leading published row is Claude Mythos 5 at 80.3%.
Honest limit: OpenAI's July 2026 audit estimated that about 30% of the 731-task public split is broken and retracted its earlier recommendation to adopt the benchmark. The current rows remain useful published receipts, but this page should not decide a coding-agent purchase by itself.
Top models on SWE-bench Pro — July 28, 2026
As of July 28, 2026, Claude Mythos 5 leads the SWE-bench Pro leaderboard with 80.3% , followed by Claude Fable 5 (80%) and Claude Opus 5 (79.2%).
Claude Mythos 5
Anthropic
claude-mythos-5
Claude Fable 5
Anthropic
claude-fable-5
Claude Opus 5
Anthropic
claude-opus-5
Leaderboard (54 models)
ScoreAccording to BenchLM.ai, Claude Mythos 5 leads the SWE-bench Pro benchmark with a score of 80.3%, followed by Claude Fable 5 (80%) and Claude Opus 5 (79.2%). The top models are clustered within 1.1 points, suggesting this benchmark is nearing saturation for frontier models.
54 models have been evaluated on SWE-bench Pro. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Pro contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.
About SWE-bench Pro
Year
2025
Tasks
1,865 repository problems
Format
Repository task completion
Difficulty
Long-horizon professional engineering
The authors assembled 1,865 problems from 41 repositories across public, held-out, and commercial splits. Agents receive a repository and issue, then produce a patch that must pass the evaluation tests without breaking existing behavior.
BenchLM freshness & provenance
Version
SWE-bench Pro 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does SWE-bench Pro measure?
SWE-bench Pro gives an agent a repository and issue description, then checks whether its patch passes new tests without breaking existing behavior. The full benchmark contains 1,865 problems from 41 repositories across public, held-out, and commercial splits, with work that can span multiple files and long execution horizons.
Are SWE-bench Pro scores directly comparable?
Only when the split and evaluation setup match. Public, held-out, and commercial tasks are different pools, while scaffold, tool budget, retry policy, and token budget can change pass rates. This page preserves exact published rows, but it does not pretend every provider ran the same harness.
Should SWE-bench Pro decide which coding agent to use?
No. The benchmark covers realistic repository work, but OpenAI's July 2026 audit estimated that about 30% of the public tasks are broken and retracted its earlier adoption recommendation. Use SWE-bench Pro alongside LiveCodeBench, other repository evaluations, and a workload-specific trial instead of treating one score as a procurement decision.
Compare Top Models on SWE-bench Pro
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
One email each week. Unsubscribe anytime.