Skip to main content
BenchLM
Data

LLM Leaderboard & AI Model Benchmarks — October 2026

As of October 6, 2026, the BenchLM AI leaderboard ranks 214 models. Claude Opus 5.5 and GPT-6 Astra tie under the 90% interval rule. Open the evidence before choosing.

Search models and benchmarks

Current BenchAlign leaders

Snapshot October 6, 2026

Method
  1. Anthropic logo

    Anthropic01

    Score86.3990% interval 80.5–92.3
    Agentic
    88
    Coding
    84
    Reasoning
    82
    Multimodal
    89
    Knowledge
    88
    Multilingual
    Not measured
    Instruction following
    Not measured
    Math
    Not measured

    Supportedenough independent evidence

    Details for Claude Opus 5.5
  2. OpenAI logo
    GPT-6 AstraSupported

    OpenAI02

    Score84.990% interval 79.9–90.0
    Agentic
    71
    Coding
    77
    Reasoning
    90
    Multimodal
    Not measured
    Knowledge
    85
    Multilingual
    Not measured
    Instruction following
    Not measured
    Math
    Not measured

    Supportedenough independent evidence

    Details for GPT-6 Astra
  3. Anthropic logo

    Anthropic03

    Score83.990% interval 77.1–90.7
    Agentic
    68
    Coding
    85
    Reasoning
    79
    Multimodal
    Not measured
    Knowledge
    81
    Multilingual
    Not measured
    Instruction following
    Not measured
    Math
    Not measured

    Supportedenough independent evidence

    Details for Claude Sonnet 5.5

Price context: Claude Sonnet 5.5 retains 97% of the top score with an output price 50% lower.

Data refreshed
October 6, 2026
Supported
83 models
Ranking method
BenchAlign v5.8

The LLM leaderboard — October 2026 rankings

Score, model size, context, price, and benchmark detail in one table. Filter state syncs to the URL so every view is shareable. Supported and Estimated labels show how much independent evidence backs each position.

Supported means we have enough independent evidence to stand behind the row. Estimated means the rank is visible and the uncertainty is wider.

Supportedenough independent evidenceEstimatedwider uncertainty

Parameters show total billions; active counts appear when a source reports them. Size filters use total parameters: Tiny ≤4B, Small >4–<40B, Medium 40–150B, and Large >150B.

AG Agentic CO Coding RE Reasoning MM Multimodal & Grounded KN Knowledge ML Multilingual IF Instruction Following MA Math

Showing 25 of 882 models · 214 ranked

Source-linked · the table lists all 882 tracked models; 214 carry a BenchAlign rank

How the score is built

What the table covers

The BenchLM LLM leaderboard 2026 ranks 214+ models and tracks 882+ large language models side by side across 613 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning. Whether you need the best AI models 2026 has to offer for agentic workflows, math, multilingual tasks, or instruction following, our AI benchmark comparison tables make it easy to see how GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open-source models stack up on both benchmarks and operator tradeoffs like price and context. Supported and Estimated labels separate evidence strength from score, so incomplete reporting does not automatically remove a model from the comparison.

What changed in October 2026

  1. Claude Opus 5.5 (Anthropic) holds the #1 spot on the BenchAlign leaderboard at 86.39.
  2. 10 model releases tracked in October 2026.
  3. 83 Supported and 131 Estimated models appear in the public leaderboard.

Dataset journalData verified

Radar · latest confirmed releases

3 source-linked updates

Every model on this board will change. Radar keeps the record: releases, price changes, retirements, and API changes, each with its source and its date, matched to the models you declare.

Build a head-to-head comparison

Choose any two ranked models. The comparison opens with the decision, exact units, and evidence coverage.

Decision-ready picks

Current models that clear the ranking, freshness, and evidence thresholds for each decision.

DecisionModelCurrent measureProvenance
  • Best open weightHighest-scoring evidence-qualified current open-weight model.MiMo-V2.6-ProXiaomi74.17overall scoreEstimated · 2 source families
  • Best near-frontier value97% of the leading score at a lower output price.Claude Sonnet 5.5Anthropic$10output / 1M tokensSupported · 3 source families
  • Fastest measuredFastest externally measured model that also clears the ranking and evidence thresholds.Mercury 2.5Inception691tokens / secArtificial Analysis · updated 2026-10-06
  • Largest useful contextLargest context among evidence-qualified current models retaining at least half of the leading score.GPT-6 AstraOpenAI1.05Mcontext windowSupported · 3 source families

Current models that clear rank, freshness and evidence thresholds

Qualification method

How we work

A ranking is a shortlist.

Rank orders the evidence; it does not make the decision. Every row here opens to its sources so you can disagree with us.

We do not invent scores to fill a table.

A model without a public, independent row stays unranked, and says so. Vendor-run numbers are shown, labeled, and kept out of the rank.

When we were wrong, the fix is on the page.

Scoring bugs, missing coverage, and re-weightings are dated and linked, because a ranking you cannot audit is a launch post with a table.

The AI race, read as a record instead of a spectacle

Track which release led its month, how provider depth changed, and whether the crown moved.

Months
6
Releases
219
Lead changes
5

Best ranked model released in Oct 2026

Mistral Large 4Mistral
Score53.69
RankProviderTop-three avg.Models
01Anthropic8421
02OpenAI81.436
03Google74.219

Six-month release record

Explore the release timeline

Scoring methodology

Scoring methodology8 weighted categories, 61 ranking benchmarks, and 312 display-only records.Inspect method
Category weights
100%
Scored categories
8
Ranking benchmarks
61
Display-only
312

Each model's overall score is a normalized weighted average of category averages. Within each category, benchmark results are normalized to a common scale and combined using weights that favor harder, less-saturated evaluations.

Confidence remains separate from score. It shows how much public benchmark evidence supports a result, while display-only benchmarks stay visible for context without receiving ranking weight.

Benchmark scores come from official provider tables, model cards, release announcements, and benchmark-native public leaderboards. External consensus signals are bounded calibration inputs and are not exposed in exported data.

Read the complete methodology
Category and weightBenchmarks that affect the scoreCoverage
  • Agentic22%

    Terminal-Bench 3.0 · Terminal-Bench 4.0 · AA AutomationBench · AA EnterpriseOps-Gym · AA Harvey LAB · AA Tau3 Banking · Terminal-Bench 2.0 · Terminal-Bench 2.1 · BrowseComp · HLE w/ tools · GDPval-AA · OSWorld-Verified · OSWorld 2.0 · JobBench · MCP Atlas · Toolathlon · Toolathlon-Verified · AutomationBench · Agents' Last Exam · DeepSearchQA · BFCL v4 · τ³-bench results · Terminal-Bench 2.1 (Vals)

    23 weighted · 76 context-only

  • Coding20%

    SWE-Rebench · SWE-bench Pro · VulcanBench v3 · FrontierCode 1.1 Main · SWE Multilingual · FrontierSWE v2 · SciCode · AA-SciCode · LiveCodeBench (Vals) · DeepSWE

    10 weighted · 53 context-only

  • Reasoning17%

    LongBench v2 · MRCRv2 · ARC-AGI-2 · ARC-AGI-3 · AA-LCR · CritPt

    6 weighted · 26 context-only

  • Multimodal12%

    MMMU-Pro · AA-MMMU-Pro · OfficeQA Pro · CharXiv

    4 weighted · 65 context-only

  • Knowledge12%

    GPQA · SuperGPQA · MMLU-Pro · HLE · AA-GPQA Diamond · AA-HLE · AA-Omniscience Accuracy · HLE w/o tools · GPQA Diamond (Vals) · MMLU-Pro (Vals)

    10 weighted · 47 context-only

  • Multilingual7%

    MMLU-ProX

    1 weighted · 13 context-only

  • Instruction Following5%

    IFBench · AA-IFBench

    2 weighted · 3 context-only

  • Math5%

    AIME26 · HMMT Feb 2026 · FrontierMath v2 (Tiers 1-3) · FrontierMath v2 (Tier 4) · USAMO 2026

    5 weighted · 29 context-only

Scoring methodology · BenchAlign v5.8

Inspect method