OpenAI01
- Agentic
- 71
- Coding
- 74
- Reasoning
- 90
- Multimodal
- 83
- Knowledge
- 87
- Multilingual
- Not measured
- Instruction following
- Not measured
- Math
- 85
Supportedenough independent evidence
Details for GPT-6 AstraThe models you run will change. Get every release, price change and retirement with its source and date. Follow model changes
A ranking is a shortlist. Open the evidence before you treat it as a decision. 71 Supported and 125 Estimated models of 513 tracked — 455 benchmarks, pricing, and runtime data.
Snapshot September 23, 2026
OpenAI01
Supportedenough independent evidence
Details for GPT-6 AstraAnthropic02
Estimatedwider uncertainty
Details for Claude Opus 5.5Anthropic03
Supportedenough independent evidence
Details for Claude Fable 5.1Price context: GPT-6 Sol retains 93% of the top score with an output price 80% lower.
Benchmarks, pricing, runtime signals, and context window in one table. Filter state syncs to the URL so every view is shareable. Supported and Estimated labels show how much independent evidence backs each position.
Supported means we have enough independent evidence to stand behind the row. Estimated means the rank is visible and the uncertainty is wider.
Showing 25 of 507 models. Row order follows the BenchAlign score.
AG Agentic CO Coding RE Reasoning MM Multimodal & Grounded KN Knowledge ML Multilingual IF Instruction Following MA Math
Source-linked · the table lists 507 of 513 tracked models; 196 carry a BenchAlign rank
How the score is builtThe BenchLM LLM leaderboard 2026 ranks 196+ models and tracks 513+ large language models side by side across 455 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning. Whether you need the best AI models 2026 has to offer for agentic workflows, math, multilingual tasks, or instruction following, our AI benchmark comparison tables make it easy to see how GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open-source models stack up on both benchmarks and operator tradeoffs like price and context. Supported and Estimated labels separate evidence strength from score, so incomplete reporting does not automatically remove a model from the comparison.
Reading the table
Embed the live leaderboardDataset journalData verified
Move from the overall ranking into the capability or workload that actually decides your model choice.
Every model on this board will change. Radar keeps the record: releases, price changes, retirements, and API changes, each with its source and its date, matched to the models you declare.
Google · Sep 23, 2026
Anthropic · Sep 22, 2026
Xiaomi · Sep 22, 2026
Choose any two ranked models. The comparison opens with the decision, exact units, and evidence coverage.
Current models that clear the ranking, freshness, and evidence thresholds for each decision.
Current models that clear rank, freshness and evidence thresholds
Qualification methodQualification methodRank orders the evidence; it does not make the decision. Every row here opens to its sources so you can disagree with us.
A model without a public, independent row stays unranked, and says so. Vendor-run numbers are shown, labeled, and kept out of the rank.
Scoring bugs, missing coverage, and re-weightings are dated and linked, because a ranking you cannot audit is a launch post with a table.
The desk
Editorial policyTrack which release led its month, how provider depth changed, and whether the crown moved.
Best ranked model released in Sep 2026
Six-month release record
Explore the release timelineEach model's overall score is a normalized weighted average of category averages. Within each category, benchmark results are normalized to a common scale and combined using weights that favor harder, less-saturated evaluations.
Confidence remains separate from score. It shows how much public benchmark evidence supports a result, while display-only benchmarks stay visible for context without receiving ranking weight.
Data comes from OpenBench, official model papers, and public leaderboards. External consensus signals are bounded calibration inputs and are not exposed in exported data.
Read the complete methodologyTerminal-Bench 3.0 · Terminal-Bench 4.0 · AA AutomationBench · AA EnterpriseOps-Gym · AA Harvey LAB · AA Tau3 Banking · Terminal-Bench 2.0 · Terminal-Bench 2.1 · BrowseComp · HLE w/ tools · GDPval-AA · OSWorld-Verified · OSWorld 2.0 · JobBench · MCP Atlas · Toolathlon · Toolathlon-Verified · AutomationBench · Agents' Last Exam · DeepSearchQA · BFCL v4 · τ³-bench results · Terminal-Bench 2.1 (Vals)
23 weighted · 70 context-only
SWE-Rebench · SWE-bench Pro · VulcanBench v3 · FrontierCode 1.1 Main · SWE Multilingual · FrontierSWE v2 · SciCode · AA-SciCode · LiveCodeBench (Vals) · DeepSWE
10 weighted · 53 context-only
LongBench v2 · MRCRv2 · ARC-AGI-2 · ARC-AGI-3 · AA-LCR · CritPt
6 weighted · 24 context-only
MMMU-Pro · AA-MMMU-Pro · OfficeQA Pro · CharXiv
4 weighted · 65 context-only
GPQA · SuperGPQA · MMLU-Pro · HLE · AA-GPQA Diamond · AA-HLE · AA-Omniscience Accuracy · HLE w/o tools · GPQA Diamond (Vals) · MMLU-Pro (Vals)
10 weighted · 44 context-only
MMLU-ProX
1 weighted · 13 context-only
IFBench · AA-IFBench
2 weighted · 2 context-only
AIME26 · HMMT Feb 2026 · FrontierMath v2 (Tiers 1-3) · FrontierMath v2 (Tier 4) · USAMO 2026
5 weighted · 29 context-only
Scoring methodology · BenchAlign v5.6
Inspect method