Best AI Models in 2026 — Overall Rankings
The best AI model right now is Claude Fable 5.1. It leads the September 2026 ranking with a score of 83, ahead of Claude Fable 5 (82.6) and Claude Opus 5 (82.4). Each row uses the public BenchAlign v5 contract and shows its evidence status.
Citable stat 230 AI models are ranked as of September 2, 2026; Claude Fable 5.1 leads at 83/100.
What the ranking says today
I would start a broad evaluation with Claude Fable 5.1, Claude Fable 5, and Claude Opus 5. They are the first three rows under the same public scoring contract; the order is a shortlist, not a purchase order.
Operator receipt. I regenerated the active artifact against the reviewed 283-model registry on July 15, 2026. The public lane has been rebuilt on every data refresh since; as of September 2, 2026 it contains 230 eligible rows, led by Claude Fable 5.1 at 83.
Honest limit. BenchAlign v5 projects a common score from uneven public evidence. An “estimated” badge and a wide interval mean the row is useful for discovery but weak evidence for a close call. Replay the finalists on your own tasks, latency budget, and safety constraints.
Operator review by Glevd
About this ranking
Last verified: September 2, 2026
BenchLM.ai now distinguishes provisional overall ranking from verified overall ranking. The provisional score is a normalized weighted average across 8 benchmark categories: agentic (22%), coding (20%), reasoning (17%), knowledge (12%), multimodal & grounded (12%), multilingual (7%), instruction following (5%), and math (5%), using non-generated benchmark coverage plus bounded external consensus calibration. The verified leaderboard is stricter and only counts sourced benchmark rows. Each score includes a confidence indicator (1-4 dots) based on how much sourced coverage supports it. Display-only benchmarks — including MMLU, OpenBookQA, HumanEval, FLTEval, BBH, LisanBench, and older AIME/HMMT variants — remain visible for context but do not affect ranking.
This is the public BenchAlign v5 overall lane. Scores combine the approved agentic and coding projections; evidence labels distinguish supported rows from estimates.
The active lane starts with Claude Fable 5.1, followed by Claude Fable 5 and Claude Opus 5. All rows use the same BenchAlign v5 projection contract. Evidence badges matter as much as small point gaps because the public source coverage is uneven.
Open-weight and proprietary models share this lane, but deployment terms, latency, and operating cost are not part of the BenchAlign score. Use the ranking to choose a test set, then compare the finalists under the constraints that will exist in production.
This ranking uses the public BenchAlign v5 overall contract. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
Claude Fable 5.1 leads the live overall ranking at 83 with Estimated evidence.
Claude Fable 5 ranks #2 at 82.6 with Supported evidence.
Claude Opus 5 ranks #3 at 82.4 with Supported evidence.
Full Rankings (230 models)
72.1
BenchAlign v5
90% interval 62.21–81.95
68.7
BenchAlign v5
90% interval 58.80–78.54
66.7
BenchAlign v5
90% interval 59.06–74.26
65.3
BenchAlign v5
90% interval 53.76–76.78
65.1
BenchAlign v5
90% interval 52.93–77.30
64.7
BenchAlign v5
90% interval 51.13–78.18
63.6
BenchAlign v5
90% interval 57.95–69.23
62.3
BenchAlign v5
90% interval 50.83–73.85
61.2
BenchAlign v5
90% interval 51.36–71.10
61.2
BenchAlign v5
90% interval 51.35–71.09
61.2
BenchAlign v5
90% interval 51.35–71.09
61.2
BenchAlign v5
90% interval 51.35–71.09
61.2
BenchAlign v5
90% interval 51.35–71.09
61.1
BenchAlign v5
90% interval 49.54–72.57
60.5
BenchAlign v5
90% interval 48.99–72.02
60.4
BenchAlign v5
90% interval 41.34–79.50
60.4
BenchAlign v5
90% interval 48.84–71.87
59.7
BenchAlign v5
90% interval 45.67–73.71
59.3
BenchAlign v5
90% interval 48.33–70.24
59.1
BenchAlign v5
90% interval 47.61–70.63
59
BenchAlign v5
90% interval 47.49–70.52
57.9
BenchAlign v5
90% interval 46.37–69.40
57.6
BenchAlign v5
90% interval 46.07–69.10
57.4
BenchAlign v5
90% interval 45.90–68.93
56.9
BenchAlign v5
90% interval 45.38–68.41
55.9
BenchAlign v5
90% interval 41.80–70.02
55.5
BenchAlign v5
90% interval 43.95–66.98
55.4
BenchAlign v5
90% interval 43.90–66.93
54.7
BenchAlign v5
90% interval 43.19–66.22
53.9
BenchAlign v5
90% interval 42.35–65.37
53.8
BenchAlign v5
90% interval 42.27–65.30
53.4
BenchAlign v5
90% interval 43.50–63.24
52.9
BenchAlign v5
90% interval 33.10–72.69
51.9
BenchAlign v5
90% interval 40.34–63.37
51.2
BenchAlign v5
90% interval 39.64–62.67
51
BenchAlign v5
90% interval 24.75–77.27
51
BenchAlign v5
90% interval 39.44–62.47
50.8
BenchAlign v5
90% interval 39.24–62.27
50.6
BenchAlign v5
90% interval 40.71–60.45
50.2
BenchAlign v5
90% interval 38.69–61.72
49.4
BenchAlign v5
90% interval 37.90–60.92
48.6
BenchAlign v5
90% interval 37.06–60.09
48.1
BenchAlign v5
90% interval 30.81–65.33
45.9
BenchAlign v5
90% interval 34.37–57.40
45.2
BenchAlign v5
90% interval 33.68–56.71
45.1
BenchAlign v5
90% interval 42.12–48.16
44.7
BenchAlign v5
90% interval 33.16–56.19
44.4
BenchAlign v5
90% interval 32.85–55.88
43.7
BenchAlign v5
90% interval 32.17–55.19
43.1
BenchAlign v5
90% interval 31.57–54.60
42.2
BenchAlign v5
90% interval 38.89–45.43
41.1
BenchAlign v5
90% interval 29.62–52.65
41
BenchAlign v5
90% interval 29.52–52.55
40.3
BenchAlign v5
90% interval 28.75–51.77
39.9
BenchAlign v5
90% interval 28.36–51.39
39.7
BenchAlign v5
90% interval 28.21–51.24
35.4
BenchAlign v5
90% interval 16.38–54.32
34.4
BenchAlign v5
90% interval 17.63–51.09
26.6
BenchAlign v5
90% interval 16.73–36.46
15.8
BenchAlign v5
90% interval 12.78–18.87
15.1
BenchAlign v5
90% interval 12.20–18.02
Key Takeaways
The top model is Claude Fable 5.1 by Anthropic with a BenchAlign v5 score of 83 and Estimated evidence.
The best open-weight model is Qwen3.8 Max at position #6.
230 models are included in this ranking.
Score in Context
What these scores mean
The overall score is a weighted average across 8 benchmark categories. Agentic (22%), coding (20%), and reasoning (17%) carry the most weight. A 5-point gap in overall score is meaningful — it reflects consistent performance differences across multiple domains.
Known limitations
The overall score compresses 8 categories into one number. Two models with the same overall score can have very different strengths — one might lead coding while the other leads reasoning. Always check category scores for your specific use case.
Best AI Models Overall FAQ
What is the best AI model right now?
The model in the #1 row of the leaderboard above is the best AI model right now on BenchLM’s weighted rankings — the answer box at the top of this page names it with its current score. Rankings are recomputed on every data refresh across agentic, coding, reasoning, knowledge, multimodal, multilingual, instruction-following, and math benchmarks, so the leader can change as new models ship.
What is the smartest AI model?
"Smartest" depends on the yardstick. For raw reasoning and knowledge, check the reasoning and knowledge category columns on the leaderboard above; for the broadest definition of capability, the overall #1 is the best single answer. Reasoning-class models typically lead on hard-science benchmarks like GPQA Diamond and Humanity’s Last Exam, while non-reasoning models can still win on speed-sensitive everyday tasks.
What is the best LLM overall?
BenchLM ranks every tracked LLM by a weighted overall score — agentic work counts 22%, coding 20%, and reasoning 17%, with knowledge, multimodal, multilingual, instruction following, and math making up the rest. The current best LLM overall is the top row of this page’s leaderboard, with confidence dots showing how much sourced benchmark coverage supports its score.
How are these AI models ranked?
Each model’s overall score is a normalized weighted average across 8 benchmark categories, using non-generated benchmark coverage plus bounded external consensus calibration. A stricter verified leaderboard counts only sourced benchmark rows. Models without enough non-generated coverage are tracked but unranked, so a high position here reflects real, sourced evidence rather than a single headline benchmark.
Is the best AI model worth paying for?
Usually only for the slice of work where quality failures are expensive. The top-ranked models carry premium per-token prices, while models a few points lower often cost 2–5x less. Check the price-vs-performance view and the per-provider API pricing hubs to find the cheapest model that clears your quality bar, then route only high-stakes tasks to the leader.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.