Benchmark profile
Software Engineering Benchmark Verified (SWE-bench Verified)
A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.
Data verifiedClaude Mythos 5 leads the SWE-bench Verified leaderboard on BenchLM's July 2026 update with 95.5%, ahead of Claude Fable 5 (95%) and Claude Opus 4.8 (88.6%), across 58 tracked models.
Top models on SWE-bench Verified — July 20, 2026
As of July 20, 2026, Claude Mythos 5 leads the SWE-bench Verified leaderboard with 95.5% , followed by Claude Fable 5 (95%) and Claude Opus 4.8 (88.6%).
Claude Mythos 5
Anthropic
claude-mythos-5
Claude Fable 5
Anthropic
claude-fable-5
Claude Opus 4.8
Anthropic
claude-opus-4-8
Leaderboard (58 models)
ScoreAccording to BenchLM.ai, Claude Mythos 5 leads the SWE-bench Verified benchmark with a score of 95.5%, followed by Claude Fable 5 (95%) and Claude Opus 4.8 (88.6%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.
58 models have been evaluated on SWE-bench Verified. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, SWE-bench Verified contributes 16% of the category score, so strong performance here directly affects a model's overall ranking.
About SWE-bench Verified
Year
2024
Tasks
500 verified issues
Format
Code patch generation
Difficulty
Professional software engineering
SWE-bench Verified is the gold standard for evaluating AI coding agents on real-world software engineering tasks. Each task requires understanding codebases, writing patches, and passing test suites.
BenchLM freshness & provenance
Version
SWE-bench Verified 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does SWE-bench Verified measure?
A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.
Which model scores highest on SWE-bench Verified?
Claude Mythos 5 by Anthropic currently leads with a score of 95.5% on SWE-bench Verified.
How many models are evaluated on SWE-bench Verified?
58 AI models have been evaluated on SWE-bench Verified on BenchLM.
Compare Top Models on SWE-bench Verified
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.