Best LLM for Coding (October 2026): SWE-bench & LiveCodeBench Ranked
Data refreshed:
Programming and software development
Claude Sonnet 5.5 leads coding on BenchLM's October 2026 rankings with a score of 84, ahead of Claude Opus 5.5 (83.3) and Claude Fable 5.1 (77.3). Supported and Estimated labels show the evidence maturity behind each position.
Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.
Coding leaders change with every release. Get the releases and price changes that move this list. Follow model changes
- Data refreshed
- October 10, 2026
- Ranked
- 146 of 889 models
- Supported / Estimated
- 81 / 65
- Scored evidence
- 11 of 26 benchmarks
26 tracked benchmarks
HumanEval, SWE-bench Verified, LiveCodeBench, LiveCodeBench Pro, FLTEval, SWE-bench Pro, SWE-Rebench, SWE Multilingual, CursorBench 3.2, CursorBench 4.0, Multi-SWE Bench, VIBE-Pro, NL2Repo, Vibe Code Bench, React Native Evals, SWE-bench Verified*, Spider 2.0-Lite, Bug Hunt Bench, PostTrainBench v1.1
Evidence set: HumanEval, SWE-bench Verified, LiveCodeBench, LiveCodeBench Pro, FLTEval, SWE-bench Pro, SWE-Rebench, SWE Multilingual, CursorBench 3.2, CursorBench 4.0, Multi-SWE Bench, VIBE-Pro, NL2Repo, Vibe Code Bench, React Native Evals, SWE-bench Verified*, Spider 2.0-Lite, Bug Hunt Bench, PostTrainBench v1.1
Best Coding picks
BenchLM summaries for coding plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
SWE-bench Pro & LiveCodeBench Leaderboard
Primary score: BenchAlign coding score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Filters
| Parameters (B) | |||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1 | Not reported | Closed | 84.0% | 84.02 | — | — | — | — | — | 81.3% | — | 90.3% | — | 55.5% | — | — | — | — | — | — | — | 51.3 fixes | — |
2 | Not reported | Closed | 83.3% | 83.27 | — | — | — | — | — | 89.9% | — | 93.9% | — | 57.8% | — | — | — | — | — | — | — | 41.7 fixes | 49.3% |
3 | Not reported | Closed | 77.3% | 77.32 | — | — | — | — | — | 81.2% | — | 89.1% | 73.4% | 51.8% | — | — | — | — | — | — | — | 43.0 fixes | 40.2% |
4 | Not reported | Closed | 76.1% | 76.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 45.0 fixes | 44.3% |
5 | Not reported | Closed | 72.4% | 72.41 | — | — | — | — | — | — | — | — | — | — | — | — | — | 91.90% | — | — | — | — | 45.3% |
6 | Not reported | Closed | 71.6% | 71.58 | — | 95% | — | — | — | 80% | — | — | 70.5% | — | — | — | — | — | — | — | — | — | — |
7 | Not reported | Closed | 71.1% | 71.09 | — | 96% | — | — | — | 79.2% | — | 89.5% | 70.0% | 46.6% | — | — | — | — | — | — | — | 27.0 fixes | 35.0% |
8 | Not reported | Closed | 71.0% | 70.98 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 44.3 fixes | — |
9 | Not reported | Closed | 68.7% | 68.68 | — | — | — | — | — | 64.6% | — | — | 67.2% | 41.7% | — | — | — | — | — | — | — | 42.0 fixes | 36.2% |
10 | Not reported | Closed | 66.1% | 66.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
11 | Not reported | Open | 64.3% | 64.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 22.7 fixes | — |
12 | Not reported | Closed | 63.5% | 63.54 | — | — | — | — | — | 63.4% | — | — | 64.9% | 41.3% | — | — | — | — | — | — | — | — | — |
13 | Not reported | Closed | 62.7% | 62.66 | — | — | — | — | — | 62.7% | — | — | 61.1% | 35.9% | — | — | — | — | — | — | — | — | — |
14 | Not reported | Closed | 62.5% | 62.45 | — | — | — | — | — | — | — | — | 69.2% | 39.6% | — | — | — | — | — | — | — | 18.0 fixes | — |
15 | Not reported | Closed | 62.1% | 62.07 | — | — | — | — | — | 58.6% | — | — | 58.4% | — | — | — | — | 69.85% | 84.7% | — | — | — | 27.2% |
16 | Not reported | Open | 61.7% | 61.7 | — | — | — | — | — | — | — | — | — | — | — | — | 65.4% | — | — | — | — | 21.7 fixes | — |
17 | Not reported | Open | 61.0% | 61.02 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
18 | Not reported | Pending | 61.0% | 61.01 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
19 | Not reported | Closed | 60.8% | 60.77 | — | 88.6% | — | — | — | 69.2% | — | 84.4% | 62.3% | — | — | — | — | — | — | — | — | — | 32.9% |
20 | Not reported | Closed | 60.6% | 60.56 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 21.5 fixes | — |
21 | Not reported | Closed | 59.8% | 59.77 | — | — | — | — | — | — | — | — | 70.8% | 41.4% | — | — | — | — | — | — | — | 27.0 fixes | — |
22 | Not reported | Pending | 59.7% | 59.74 | — | — | — | — | — | — | — | — | 60.8% | — | — | — | — | — | — | — | — | 21.0 fixes | 32.0% |
23 | Not reported | Closed | 58.6% | 58.57 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
24 | Not reported | Closed | 58.3% | 58.29 | — | 85.2% | — | — | — | 63.2% | — | 78.3% | 61.5% | 34.1% | — | — | — | — | — | — | — | — | — |
25 | Not reported | Open | 57.4% | 57.38 | — | — | — | — | — | 65.7% | — | 82.9% | — | — | — | — | 58.9% | — | — | — | — | — | — |
Query current rankings and available benchmark results from your scripts or AI assistant. Start with 1,000 free reads each month.
See queries and coverageTop AI models for Coding — October 2026
As of October 2026, Claude Sonnet 5.5 leads the BenchAlign coding leaderboard with a score of 84.0, followed by Claude Opus 5.5 (83.3) and Claude Fable 5.1 (77.3). BenchLM is currently showing 81 Supported and 65 Estimated models in this category.
Ranks #1 on the current coding board with a Supported evidence label.
Ranks #2 on the current coding board with a Supported evidence label.
Ranks #3 on the current coding board with a Supported evidence label.
What changed
Claude Sonnet 5.5 ranks #1 at 84.0 with a Supported evidence label.
Claude Opus 5.5 ranks #2 at 83.3 with a Supported evidence label.
Claude Fable 5.1 ranks #3 at 77.3 with a Supported evidence label.
Top models by benchmark
Real-world GitHub issues from popular Python repos, human-verified subset(10% of category score)
Score in Context
What these scores mean
BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.
A ranking shows order, not distance. The Claude Fable 5.1 and GPT-6 Astra comparison shows how far apart two frontier models sit on each shared benchmark.
Comparing Claude releases? Read Fable 5 against Fable 5.1 and Opus 5 against Opus 5.5 to inspect each pair's coding evidence and listed API prices before testing on your repository.
Known limitations
The table orders point estimates. Conditional score ranges describe uncertainty under the scoring assumptions; their coverage after category changes and their ability to establish rank confidence have not been validated. Supported and Estimated describe the evidence behind each row.
How we weight
This lens combines external category signals with admitted benchmark protocols. Where category evidence is thin, the score pools toward the model's general capability, estimated without human-preference leaderboards. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| SWE-bench Pro | 26% | Scored | Harder frontier coding-agent benchmark for real software engineering work |
| DeepSWE | 15% | Scored | A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories. |
| CursorBench 3.2 | 10% | Scored | Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks from real Cursor sessions, 3.2 task set (frozen). |
| SciCode | 10% | Scored | SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research. |
| FrontierSWE v2 | 8% | Scored | A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation. |
| LiveCodeBench (Vals) | 8% | Scored | Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks. |
| AA-SciCode | 5% | Scored | An Artificial Analysis SciCode score. |
| SWE Multilingual | 5% | Scored | A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages. |
| SWE-Rebench | 5% | Scored | Continuously updated software-engineering benchmark using fresh GitHub issues published after each model release window. |
| FrontierCode 1.1 Main | 4% | Scored | Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics. |
| VulcanBench v3 | 3% | Scored | An open software-engineering benchmark built from real merged post-cutoff pull requests across Python, Rust, TypeScript, JavaScript, and Go repositories. |
| HumanEval | — | Display only | Python programming problems with test cases |
| SWE-bench Verified | — | Display only | Real-world GitHub issues from popular Python repos, human-verified subset |
| LiveCodeBench | — | Display only | Continuously updated competitive programming problems to prevent contamination |
| LiveCodeBench Pro | — | Display only | A harder competitive-programming benchmark family with quarter-specific public leaderboards. |
| FLTEval | — | Display only | Lean 4 formal verification and proof-engineering benchmark built from realistic FLT project pull requests |
| CursorBench 4.0 | — | Display only | Cursor's current first-party benchmark for ambiguous, multi-file coding-agent tasks from real Cursor sessions, 4.0 task set. |
| Multi-SWE Bench | — | Display only | A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem. |
| VIBE-Pro | — | Display only | A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks. |
| NL2Repo | — | Display only | A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes. |
| Vibe Code Bench | — | Display only | Vals.ai benchmark for building complete web applications from natural language specifications. |
| React Native Evals | — | Display only | An open benchmark for AI coding agents on framework-specific React Native implementation tasks covering real app behavior, architecture, and constraint adherence. |
| SWE-bench Verified* | — | Display only | Display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart. |
| Spider 2.0-Lite | — | Display only | A text-to-SQL benchmark over realistic warehouse-scale schemas, reported by Interfaze for model comparison. |
| Bug Hunt Bench | — | Display only | Blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories. |
| PostTrainBench v1.1 | — | Display only | Post-training four base language models across seven weighted benchmarks, with ten hours and one H100 per run. |
About Coding benchmarks
Python programming problems with test cases
Questions
Which LLM is best for coding?
The model in the #1 row of the live leaderboard above is BenchLM's current best LLM for coding. Rankings are recomputed on every data refresh from a weighted blend of SWE-bench Pro (real GitHub issues) and LiveCodeBench (contamination-resistant competitive programming), so the answer box at the top of this page always names the current leader and its score rather than a snapshot that can go stale.
What is the best LLM for coding right now?
Right now the top three coding models are shown in the answer box and Top ranked panel on this page, updated with each leaderboard refresh. The current leaders separate themselves on SWE-bench Pro, the hardest widely run software engineering benchmark, where a few points of difference typically decide whether a model can resolve a multi-file GitHub issue end to end. Check the live table for today's exact ordering and scores.
What is the best free LLM for coding?
Free usually means one of two things: a free chat tier for a proprietary model, or open weights you can download and run yourself. Most frontier coding models offer rate-limited free tiers in their chat apps, while open-weight models cost nothing to self-host beyond compute. For the strongest no-cost option, start with the highest-ranked open-weight model on this leaderboard, then compare it in our best open-source LLM ranking.
What is the best open source LLM for coding?
The best open-source coding model is the highest-ranked row marked Open Weight on the leaderboard above. Open-weight models now sit within a few points of the proprietary frontier on SWE-bench Pro and LiveCodeBench, and can be self-hosted, fine-tuned, and run without per-token API pricing. Our best open-source LLM page ranks them across every category and is the fastest way to find the current open-weight leader.
How do you benchmark an LLM's coding ability?
By executing the model's code, not by grading it subjectively. SWE-bench Pro hands models real GitHub issues and counts a task solved only when the generated patch passes the repository's own test suite. LiveCodeBench scores freshly published competitive-programming problems to rule out training-data contamination. DeepSWE adds long-horizon software engineering with agent harnesses. Each benchmark page states whether it counts toward the coding score and with what weight.
Coding leaderboard updates
Know which model codes best — before your team picks the wrong one.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.
Related
Hosting Platforms for AI Apps
Compare the deployment layer for the app around your coding model.
Deploy an AI App
Move from a model choice to a production URL.
Best LLMs Overall
Top models ranked across all benchmark categories.
Best Open-Weight Models
Top open-source models for code generation and debugging.
Agentic Benchmarks
How models perform on autonomous coding agent tasks.
AI Cost Calculator
Compare pricing across models for coding workloads.