Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live evidenceRankings, cost, context, and source coverage in one inspection surface.September 18, 2026
LLM leaderboard, September 2026

LLM Leaderboard & AI Model Benchmarks — September 2026

A ranking is a shortlist. Open the evidence before you treat it as a decision. 130 Supported and 100 Estimated models of 496 tracked — 446 benchmarks, pricing, and runtime data.

Search models and benchmarks
Data refreshed
September 18, 2026
Supported
130 models
Ranking method
BenchAlign v5.2

Radar · latest confirmed releases

3 source-linked updates

Every model on this board will change. Radar keeps the record: releases, price changes, retirements, and API changes, each with its source and its date, matched to the models you declare.

Decision-ready picks

Current models that clear the ranking, freshness, and evidence thresholds for each decision.

  • Best open weightHighest-scoring evidence-qualified current open-weight model.Qwen3.8 MaxAlibaba73.17overall scoreSupported · 3 source families
  • Best near-frontier value97% of the leading score at a lower output price.Claude Opus 5Anthropic$25output / 1M tokensSupported · 3 source families
  • Fastest measuredFastest externally measured model that also clears the ranking and evidence thresholds.Gemini 3.5 Flash-LiteGoogle358tokens / secArtificial Analysis · updated 2026-09-18
  • Largest useful contextLargest context among evidence-qualified current models retaining at least half of the leading score.GPT-6 AstraOpenAI1.05Mcontext windowEstimated · 2 source families
Qualification method

The LLM leaderboard — September 2026 rankings

Benchmarks, pricing, runtime signals, and context window in one table. Filter state syncs to the URL so every view is shareable. Supported and Estimated labels show how much independent evidence backs each position.

AG AgenticCO CodingRE ReasoningMM Multimodal & GroundedKN KnowledgeML MultilingualIF Instruction FollowingMA Math

Reading the table

What the table covers

The BenchLM LLM leaderboard 2026 ranks 230+ models and tracks 496+ large language models side by side across 446 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning. Whether you need the best AI models 2026 has to offer for agentic workflows, math, multilingual tasks, or instruction following, our AI benchmark comparison tables make it easy to see how GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open-source models stack up on both benchmarks and operator tradeoffs like price and context. Supported and Estimated labels separate evidence strength from score, so incomplete reporting does not automatically remove a model from the comparison.

Embed the live leaderboard

Dataset journal

What changed in September 2026

  • Claude Fable 5.1 (Anthropic) holds the #1 spot on the BenchAlign leaderboard at 84.74.
  • 26 model releases tracked in September 2026.
  • Google climbed 1 spot in the provider race.
  • 130 Supported and 100 Estimated models appear in the public leaderboard.
Data verified

The desk

How we work

A ranking is a shortlist.

Rank orders the evidence; it does not make the decision. Every row here opens to its sources so you can disagree with us.

We do not invent scores to fill a table.

A model without a public, independent row stays unranked, and says so. Vendor-run numbers are shown, labeled, and kept out of the rank.

When we were wrong, the fix is on the page.

Scoring bugs, missing coverage, and re-weightings are dated and linked, because a ranking you cannot audit is a launch post with a table.

The evidence map

Explore benchmarks by category

Move from the overall ranking into the capability or workload that actually decides your model choice.

Build a head-to-head comparison

Choose any two ranked models. The comparison opens with the decision, exact units, and evidence coverage.

Six-month release record

The AI race, read as a record instead of a spectacle

Track which release led its month, how provider depth changed, and whether the crown moved.

Months
6
Releases
220
Lead changes
5
Explore the release timeline

Best ranked model released in Sep 2026

Claude Fable 5.1Anthropic
Score84.74
RankProviderTop-three avg.Models
01Anthropic82.720
02OpenAI78.439
03Google71.519

Scoring methodology

Scoring methodology8 weighted categories, 37 ranking benchmarks, and 320 display-only records.Inspect method

Each model's overall score is a normalized weighted average of category averages. Within each category, benchmark results are normalized to a common scale and combined using weights that favor harder, less-saturated evaluations.

Confidence remains separate from score. It shows how much public benchmark evidence supports a result, while display-only benchmarks stay visible for context without receiving ranking weight.

Data comes from OpenBench, official model papers, and public leaderboards. External consensus signals are bounded calibration inputs and are not exposed in exported data.

Read the complete methodology
Category weights
100%
Scored categories
8
Ranking benchmarks
37
Display-only
320
  • Agentic22%

    Terminal-Bench 2.0 · BrowseComp · OSWorld-Verified · OSWorld 2.0 · AutomationBench

    5 weighted · 88 context-only

  • Coding20%

    SWE-bench Verified · SWE-Rebench · LiveCodeBench · SWE-bench Pro · SWE Multilingual · SciCode · LiveCodeBench (Vals)

    7 weighted · 58 context-only

  • Reasoning17%

    LongBench v2 · MRCRv2 · ARC-AGI-2 · ARC-AGI-3 · AA-LCR

    5 weighted · 25 context-only

  • Multimodal12%

    MMMU-Pro · AA-MMMU-Pro · OfficeQA Pro · CharXiv

    4 weighted · 64 context-only

  • Knowledge12%

    GPQA · SuperGPQA · MMLU-Pro · HLE · AA-Omniscience Accuracy · SimpleQA · HLE w/o tools · MMLU-Pro (Vals)

    8 weighted · 42 context-only

  • Multilingual7%

    MMLU-ProX

    1 weighted · 14 context-only

  • Instruction Following5%

    IFBench · AA-IFBench

    2 weighted · 2 context-only

  • Math5%

    AIME26 · HMMT Feb 2026 · FrontierMath v2 (Tiers 1-3) · FrontierMath v2 (Tier 4) · USAMO 2026

    5 weighted · 27 context-only