Skip to main content
BenchLM

The models you run will change. Get every release, price change and retirement with its source and date. Follow model changes

Live evidenceSeptember 23, 2026

LLM Leaderboard & AI Model Benchmarks — September 2026

A ranking is a shortlist. Open the evidence before you treat it as a decision. 71 Supported and 125 Estimated models of 513 tracked — 455 benchmarks, pricing, and runtime data.

Search models and benchmarks

Current BenchAlign leaders

Snapshot September 23, 2026

Method
  1. OpenAI logo
    GPT-6 AstraSupported

    OpenAI01

    Score88.4790% interval 84.092.9
    Agentic
    71
    Coding
    74
    Reasoning
    90
    Multimodal
    83
    Knowledge
    87
    Multilingual
    Not measured
    Instruction following
    Not measured
    Math
    85

    Supportedenough independent evidence

    Details for GPT-6 Astra
  2. Anthropic logo

    Anthropic02

    Score88.4590% interval 76.9100.0
    Agentic
    89
    Coding
    87
    Reasoning
    79
    Multimodal
    89
    Knowledge
    90
    Multilingual
    Not measured
    Instruction following
    Not measured
    Math
    Not measured

    Estimatedwider uncertainty

    Details for Claude Opus 5.5
  3. Anthropic logo

    Anthropic03

    Score83.3490% interval 78.788.0
    Agentic
    79
    Coding
    81
    Reasoning
    79
    Multimodal
    Not measured
    Knowledge
    87
    Multilingual
    Not measured
    Instruction following
    Not measured
    Math
    Not measured

    Supportedenough independent evidence

    Details for Claude Fable 5.1

Price context: GPT-6 Sol retains 93% of the top score with an output price 80% lower.

Data refreshed
September 23, 2026
Supported
71 models
Ranking method
BenchAlign v5.6

The LLM leaderboard — September 2026 rankings

Benchmarks, pricing, runtime signals, and context window in one table. Filter state syncs to the URL so every view is shareable. Supported and Estimated labels show how much independent evidence backs each position.

Supported means we have enough independent evidence to stand behind the row. Estimated means the rank is visible and the uncertainty is wider.

Supportedenough independent evidenceEstimatedwider uncertainty

AG Agentic CO Coding RE Reasoning MM Multimodal & Grounded KN Knowledge ML Multilingual IF Instruction Following MA Math

Source-linked · the table lists 507 of 513 tracked models; 196 carry a BenchAlign rank

How the score is built

What the table covers

The BenchLM LLM leaderboard 2026 ranks 196+ models and tracks 513+ large language models side by side across 455 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning. Whether you need the best AI models 2026 has to offer for agentic workflows, math, multilingual tasks, or instruction following, our AI benchmark comparison tables make it easy to see how GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open-source models stack up on both benchmarks and operator tradeoffs like price and context. Supported and Estimated labels separate evidence strength from score, so incomplete reporting does not automatically remove a model from the comparison.

What changed in September 2026

  1. GPT-6 Astra (OpenAI) holds the #1 spot on the BenchAlign leaderboard at 88.47.
  2. 36 model releases tracked in September 2026.
  3. Google climbed 1 spot in the provider race.
  4. 71 Supported and 125 Estimated models appear in the public leaderboard.

Dataset journalData verified

Radar · latest confirmed releases

3 source-linked updates

Every model on this board will change. Radar keeps the record: releases, price changes, retirements, and API changes, each with its source and its date, matched to the models you declare.

Build a head-to-head comparison

Choose any two ranked models. The comparison opens with the decision, exact units, and evidence coverage.

Decision-ready picks

Current models that clear the ranking, freshness, and evidence thresholds for each decision.

DecisionModelCurrent measureProvenance
  • Best open weightHighest-scoring evidence-qualified current open-weight model.Qwen3.8 MaxAlibaba72.02overall scoreSupported · 3 source families
  • Best near-frontier value93% of the leading score at a lower output price.GPT-6 SolOpenAI$10output / 1M tokensEstimated · 1 source family
  • Fastest measuredFastest externally measured model that also clears the ranking and evidence thresholds.Gemini 3.5 Flash-LiteGoogle357tokens / secArtificial Analysis · updated 2026-09-23
  • Largest useful contextLargest context among evidence-qualified current models retaining at least half of the leading score.GPT-6 AstraOpenAI1.05Mcontext windowSupported · 2 source families

Current models that clear rank, freshness and evidence thresholds

Qualification method

How we work

A ranking is a shortlist.

Rank orders the evidence; it does not make the decision. Every row here opens to its sources so you can disagree with us.

We do not invent scores to fill a table.

A model without a public, independent row stays unranked, and says so. Vendor-run numbers are shown, labeled, and kept out of the rank.

When we were wrong, the fix is on the page.

Scoring bugs, missing coverage, and re-weightings are dated and linked, because a ranking you cannot audit is a launch post with a table.

The AI race, read as a record instead of a spectacle

Track which release led its month, how provider depth changed, and whether the crown moved.

Months
6
Releases
230
Lead changes
5

Best ranked model released in Sep 2026

GPT-6 AstraOpenAI
Score88.47
RankProviderTop-three avg.Models
01Anthropic8420
02OpenAI83.735
03Google67.718

Six-month release record

Explore the release timeline

Scoring methodology

Scoring methodology8 weighted categories, 61 ranking benchmarks, and 300 display-only records.Inspect method
Category weights
100%
Scored categories
8
Ranking benchmarks
61
Display-only
300

Each model's overall score is a normalized weighted average of category averages. Within each category, benchmark results are normalized to a common scale and combined using weights that favor harder, less-saturated evaluations.

Confidence remains separate from score. It shows how much public benchmark evidence supports a result, while display-only benchmarks stay visible for context without receiving ranking weight.

Data comes from OpenBench, official model papers, and public leaderboards. External consensus signals are bounded calibration inputs and are not exposed in exported data.

Read the complete methodology
Category and weightBenchmarks that affect the scoreCoverage
  • Agentic22%

    Terminal-Bench 3.0 · Terminal-Bench 4.0 · AA AutomationBench · AA EnterpriseOps-Gym · AA Harvey LAB · AA Tau3 Banking · Terminal-Bench 2.0 · Terminal-Bench 2.1 · BrowseComp · HLE w/ tools · GDPval-AA · OSWorld-Verified · OSWorld 2.0 · JobBench · MCP Atlas · Toolathlon · Toolathlon-Verified · AutomationBench · Agents' Last Exam · DeepSearchQA · BFCL v4 · τ³-bench results · Terminal-Bench 2.1 (Vals)

    23 weighted · 70 context-only

  • Coding20%

    SWE-Rebench · SWE-bench Pro · VulcanBench v3 · FrontierCode 1.1 Main · SWE Multilingual · FrontierSWE v2 · SciCode · AA-SciCode · LiveCodeBench (Vals) · DeepSWE

    10 weighted · 53 context-only

  • Reasoning17%

    LongBench v2 · MRCRv2 · ARC-AGI-2 · ARC-AGI-3 · AA-LCR · CritPt

    6 weighted · 24 context-only

  • Multimodal12%

    MMMU-Pro · AA-MMMU-Pro · OfficeQA Pro · CharXiv

    4 weighted · 65 context-only

  • Knowledge12%

    GPQA · SuperGPQA · MMLU-Pro · HLE · AA-GPQA Diamond · AA-HLE · AA-Omniscience Accuracy · HLE w/o tools · GPQA Diamond (Vals) · MMLU-Pro (Vals)

    10 weighted · 44 context-only

  • Multilingual7%

    MMLU-ProX

    1 weighted · 13 context-only

  • Instruction Following5%

    IFBench · AA-IFBench

    2 weighted · 2 context-only

  • Math5%

    AIME26 · HMMT Feb 2026 · FrontierMath v2 (Tiers 1-3) · FrontierMath v2 (Tier 4) · USAMO 2026

    5 weighted · 29 context-only

Scoring methodology · BenchAlign v5.6

Inspect method