Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief
Live evidenceRankings, cost, context, and source coverage in one inspection surface.September 4, 2026

LLM leaderboard, September 2026

LLM Leaderboard & AI Model Benchmarks — September 2026

Compare frontier AI models by quality, cost, and context. 130 Supported and 102 Estimated models among 417 tracked LLMs — 422 benchmarks, real pricing, and runtime data in one place.

Search models and benchmarks
Data refreshed
September 4, 2026
Supported
130 models
Ranking method
BenchAlign v5.2

Radar · latest confirmed releases

3 source-linked updates

Free Radar Brief ranks up to five confirmed market changes and keeps the original source attached. No card required.

Decision-ready picks

Current models that clear the ranking, freshness, and evidence thresholds for each decision.

  • Best open weightHighest-scoring evidence-qualified current open-weight model.Qwen3.8 MaxAlibaba72.43overall scoreSupported · 3 source families
  • Best near-frontier value95% of the leading score at a lower output price.Gemini 3.8 FlashGoogle$3.75output / 1M tokensSupported · 2 source families
  • Fastest measuredFastest externally measured model that also clears the ranking and evidence thresholds.LFM2.5-2.6BLiquidAI595tokens / secArtificial Analysis · updated 2026-09-04
  • Largest useful contextLargest context among evidence-qualified current models retaining at least half of the leading score.Grok 4.20xAI2Mcontext windowSupported · 2 source families
Qualification method

The LLM leaderboard — September 2026 rankings

Benchmarks, pricing, runtime signals, and context window in one table. Filter state syncs to the URL so every view is shareable. Supported and Estimated labels show how much independent evidence backs each position.

AG AgenticCO CodingRE ReasoningMM Multimodal & GroundedKN KnowledgeML MultilingualIF Instruction FollowingMA Math

Reading the table

Use rank to shortlist, then inspect the evidence

The BenchLM LLM leaderboard 2026 ranks 232+ models and tracks 417+ large language models side by side across 422 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning. Whether you need the best AI models 2026 has to offer for agentic workflows, math, multilingual tasks, or instruction following, our AI benchmark comparison tables make it easy to see how GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open-source models stack up on both benchmarks and operator tradeoffs like price and context. Supported and Estimated labels separate evidence strength from score, so incomplete reporting does not automatically remove a model from the comparison.

Embed the live leaderboard

Dataset journal

What changed in September 2026

  • Claude Fable 5.1 (Anthropic) holds the #1 spot on the BenchAlign leaderboard at 82.95.
  • 8 model releases tracked in September 2026.
  • Google climbed 1 spot in the provider race.
  • 130 Supported and 102 Estimated models appear in the public leaderboard.
Data verified

Build a head-to-head comparison

Choose any two ranked models. The comparison opens with the decision, exact units, and evidence coverage.

Six-month release record

The AI race, read as a record instead of a spectacle

Track which release led its month, how provider depth changed, and whether the crown moved.

Months
6
Releases
149
Lead changes
5
Explore the release timeline

Best ranked model released in Sep 2026

Claude Fable 5.1Anthropic
Score82.95
RankProviderTop-three avg.Models
01Anthropic81.520
02OpenAI7839
03Google72.919

The evidence map

Explore benchmarks by category

Move from the overall ranking into the capability or workload that actually decides your model choice.

Scoring methodology

Scoring methodology8 weighted categories, 37 ranking benchmarks, and 306 display-only records.Inspect method

Each model's overall score is a normalized weighted average of category averages. Within each category, benchmark results are normalized to a common scale and combined using weights that favor harder, less-saturated evaluations.

Confidence remains separate from score. It shows how much public benchmark evidence supports a result, while display-only benchmarks stay visible for context without receiving ranking weight.

Data comes from OpenBench, official model papers, and public leaderboards. External consensus signals are bounded calibration inputs and are not exposed in exported data.

Read the complete methodology
Category weights
100%
Scored categories
8
Ranking benchmarks
37
Display-only
306
  • Agentic22%

    Terminal-Bench 2.0 · BrowseComp · OSWorld-Verified · OSWorld 2.0 · AutomationBench

    5 weighted · 83 context-only

  • Coding20%

    SWE-bench Verified · SWE-Rebench · LiveCodeBench · SWE-bench Pro · SWE Multilingual · SciCode · LiveCodeBench (Vals)

    7 weighted · 55 context-only

  • Reasoning17%

    LongBench v2 · MRCRv2 · ARC-AGI-2 · ARC-AGI-3 · AA-LCR

    5 weighted · 24 context-only

  • Multimodal12%

    MMMU-Pro · AA-MMMU-Pro · OfficeQA Pro · CharXiv

    4 weighted · 61 context-only

  • Knowledge12%

    GPQA · SuperGPQA · MMLU-Pro · HLE · AA-Omniscience Accuracy · SimpleQA · HLE w/o tools · MMLU-Pro (Vals)

    8 weighted · 42 context-only

  • Multilingual7%

    MMLU-ProX

    1 weighted · 12 context-only

  • Instruction Following5%

    IFBench · AA-IFBench

    2 weighted · 2 context-only

  • Math5%

    AIME26 · HMMT Feb 2026 · FrontierMath v2 (Tiers 1-3) · FrontierMath v2 (Tier 4) · USAMO 2026

    5 weighted · 27 context-only