# SWE-bench Pro

> A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

Canonical page: https://benchlm.ai/benchmarks/swe-bench-pro

- Category: [Coding](/coding)
- Last updated: September 10, 2026

## About SWE-bench Pro

- Year: 2025
- Tasks: 1,865 repository problems
- Format: Repository task completion
- Difficulty: Long-horizon professional engineering
- Paper: [SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?](https://arxiv.org/abs/2509.16941)

The authors assembled 1,865 problems from 41 repositories across public, held-out, and commercial splits. Agents receive a repository and issue, then produce a patch that must pass the evaluation tests without breaking existing behavior.

SWE-bench Pro is currently weighted in BenchLM's scoring formula. The Coding category carries 20% of the overall score, and SWE-bench Pro contributes 25% of that category score.

## Leaderboard (68 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 81.2% |
| 2 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 80.3% |
| 3 | [Claude Fable 5](/models/claude-fable) | Anthropic | 80% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 79.2% |
| 5 | [Sakana Fugu-Ultra](/models/sakana-fugu-ultra) | Sakana AI | 73.7% |
| 6 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 69.2% |
| 7 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 67.7% |
| 8 | [Hy4 preview](/models/hy4-preview) | Tencent | 65.7% |
| 9 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 65.1% |
| 10 | [Grok 4.5](/models/grok-4-5) | xAI | 64.7% |
| 11 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 64.6% |
| 12 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 64.3% |
| 13 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 63.4% |
| 14 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 63.2% |
| 15 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 62.7% |
| 16 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 62.5% |
| 17 | [Ornith-1.0-397B](/models/ornith-1-0-397b) | DeepReinforce AI | 62.2% |
| 18 | [GLM-5.2](/models/glm-5-2) | Z.AI | 62.1% |
| 19 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 61.7% |
| 20 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 61.5% |
| 21 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 61% |
| 22 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 60.6% |
| 23 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 59.6% |
| 24 | [Laguna S 2.1](/models/laguna-s-2-1) | Poolside | 59.4% |
| 25 | [MiniMax M3](/models/minimax-m3) | MiniMax | 59% |
| 26 | [Sakana Fugu](/models/sakana-fugu) | Sakana AI | 59% |
| 27 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 58.6% |
| 28 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 58.6% |
| 29 | [GLM-5.1](/models/glm-5-1) | Z.AI | 58.4% |
| 30 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 57.7% |
| 31 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 57.6% |
| 32 | [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) | Alibaba | 57.3% |
| 33 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 57.2% |
| 34 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 57.1% |
| 35 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 56.8% |
| 36 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 56.6% |
| 37 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 56.6% |
| 38 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 56.3% |
| 39 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 56.2% |
| 40 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 56.1% |
| 41 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 55.9% |
| 42 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 55.6% |
| 43 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 55.4% |
| 44 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 55.1% |
| 45 | [GLM-5](/models/glm-5) | Z.AI | 55.1% |
| 46 | [Inkling](/models/inkling) | Thinking Machines Lab | 54.3% |
| 47 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 54.2% |
| 48 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 53.5% |
| 49 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 53.4% |
| 50 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 52.8% |
| 51 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 52.6% |
| 52 | [Muse Spark](/models/muse-spark) | Meta | 52.4% |
| 53 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 51.8% |
| 54 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 51.2% |
| 55 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 50.9% |
| 56 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 50.7% |
| 57 | [Ornith-1.0-35B](/models/ornith-1-0-35b) | DeepReinforce AI | 50.4% |
| 58 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 49.5% |
| 59 | [Laguna M.1](/models/laguna-m-1) | Poolside | 49.2% |
| 60 | [Laguna XS 2.1](/models/laguna-xs-2-1) | Poolside | 47.6% |
| 61 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 47.5% |
| 62 | [Laguna XS.2](/models/laguna-xs-2) | Poolside | 46.3% |
| 63 | [Ornith-1.0-9B](/models/ornith-1-0-9b) | DeepReinforce AI | 42.9% |
| 64 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 40.6% |
| 65 | [Granite 4.2 30B](/models/granite-4-2-30b) | IBM | 33.3% |
| 66 | [LLaDA2.2-flash](/models/llada2-2-flash) | InclusionAI | 30.1% |
| 67 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 19.1% |
| 68 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 14.4% |

## FAQ

### What does SWE-bench Pro measure?

SWE-bench Pro gives an agent a repository and issue description, then checks whether its patch passes new tests without breaking existing behavior. The full benchmark contains 1,865 problems from 41 repositories across public, held-out, and commercial splits, with work that can span multiple files and long execution horizons.

### Are SWE-bench Pro scores directly comparable?

Only when the split and evaluation setup match. Public, held-out, and commercial tasks are different pools, while scaffold, tool budget, retry policy, and token budget can change pass rates. This page preserves exact published rows, but it does not pretend every provider ran the same harness.

### Should SWE-bench Pro decide which coding agent to use?

No. The benchmark covers realistic repository work, but OpenAI's July 2026 audit estimated that about 30% of the public tasks are broken and retracted its earlier adoption recommendation. Use SWE-bench Pro alongside LiveCodeBench, other repository evaluations, and a workload-specific trial instead of treating one score as a procurement decision.

## Compare Top Models on SWE-bench Pro

- [Claude Fable 5.1 vs Claude Mythos 5](/compare/claude-fable-5-1-vs-claude-mythos-5)
- [Claude Mythos 5 vs Claude Fable 5](/compare/claude-fable-vs-claude-mythos-5)
- [Claude Fable 5 vs Claude Opus 5](/compare/claude-fable-vs-claude-opus-5)
- [Claude Opus 5 vs Sakana Fugu-Ultra](/compare/claude-opus-5-vs-sakana-fugu-ultra)
