The desk
How we work
A ranking is a shortlist.
Rank orders the evidence; it does not make the decision. Every row here opens to its sources so you can disagree with us.
We do not invent scores to fill a table.
A model without a public, independent row stays unranked, and says so. Vendor-run numbers are shown, labeled, and kept out of the rank.
When we were wrong, the fix is on the page.
Scoring bugs, missing coverage, and re-weightings are dated and linked, because a ranking you cannot audit is a launch post with a table.
Affiliate partnerships never influence scores, rankings, or model coverage. See the affiliate disclosure for details.
Evidence strength is separate from score
Supported means we have enough independent evidence to stand behind the row. Estimated means the rank is visible and the uncertainty is wider. Both remain visible in one ranking.
Freshness is explicit
Every benchmark now carries BenchLM freshness metadata: version, refresh cadence, staleness state, saturation state, and whether the benchmark is weighted or display-only.
Independent signals share one scale
BenchAlign calibrates benchmark protocols, human-preference signals, and external composite indices before combining them. Raw percentages are never averaged across tests with different difficulty. Runtime metrics remain separate and were updated 2026-09-23.
Six pillars and familiar evidence lenses
The product ontology has six pillars: Agentic Work, Software Engineering, Reasoning and Mathematics, Knowledge and Factuality, Multimodal and Documents, and Communication and Language. Existing category URLs remain indexable evidence lenses beneath those pillars.
Agentic
23 scored benchmarks and 70 display-only benchmarks.
Coding
10 scored benchmarks and 53 display-only benchmarks.
Reasoning
6 scored benchmarks and 24 display-only benchmarks.
Multimodal
4 scored benchmarks and 65 display-only benchmarks.
Knowledge
10 scored benchmarks and 47 display-only benchmarks.
Multilingual
1 scored benchmark and 13 display-only benchmarks.
Instruction Following
2 scored benchmarks and 2 display-only benchmarks.
Math
5 scored benchmarks and 29 display-only benchmarks.
Missing data increases uncertainty
BenchAlign estimates capability from the evidence a model does have instead of assigning zero to missing tests or averaging only the tests a provider chose to report. Sparse models can rank, but they receive an Estimated label and a wider uncertainty interval until independent evidence accumulates. A model that rests on a single external source, on correlated sources of one class, or on sources of two classes where one carries more than half the weight is pulled part of the way toward a relevant prior; on the overall board, when two or more sources are present, the pull closes a tenth of the gap. The first choice is a reviewed relative whose own position is carried by evidence, such as its predecessor, a lower tier, or a lower-effort profile. For a newer release, that prior is the predecessor’s score plus the typical gain we measure between newer releases and their predecessors where both positions are Supported. Without a relative, the prior is the model’s own connected benchmark estimate, then the middle of the roster. A single-source position is not pulled at all when broad benchmark evidence lands within 3 points of it. An Estimated position also cannot sit more than 3 points above what the model’s own broad benchmark results imply, even when those results are provider-reported, because published results skew optimistic. This check only lowers a score, by at most 3 points. A Pro or higher-effort version runs the same model with more compute, and its results usually cover only the hard reasoning tests where that compute pays off, so its lead over the base model counts only in the share of category weight where it has results; elsewhere the base model’s evidence stands. A position that rests only on preference leaderboards, with no benchmark composite and no connected benchmark evidence, is withheld. Lineage is a prior, never a floor: a newer release can rank below the model it replaces when its own evidence says so. Models the catalog does not publish are left out of scoring and cannot move a visible row. Ranking order uses unrounded scores.
Method bench-align-v5.7-2026-09-24 · current input 53d6ae3ab86f82d2
Benchmarks by category
A scored benchmark moves at least one public ranking. On Agentic, Coding, and Knowledge, a benchmark's percentage is its share of that category's reference weight, a relative weight in the calibrated model rather than a fixed share of any score. A benchmark stored under two categories, such as Terminal-Bench 2, is scored where the model uses it. The other categories rank by a weighted category score, and their percentages are shares of that score.
Agentic
93 tracked benchmarks
Scored
Display only
Subfamilies
Coding
63 tracked benchmarks
Scored
Display only
Multimodal
69 tracked benchmarks
Display only
Subfamilies
Knowledge
57 tracked benchmarks
Scored
Display only
Subfamilies
Multilingual
14 tracked benchmarks
Instruction Following
4 tracked benchmarks
Scored
Display only
Math
34 tracked benchmarks
Display only
Future tracked families
BenchLM tracks a small number of important benchmark families that are intentionally not weighted yet because exact-source density and cross-model coverage are still too thin for defensible ranking use.
Calibration approach
External consensus signals
Overall capability starts from independent external leaderboard families, such as benchmark composite indices and human-preference leaderboards, each weighted by its uncertainty. When several families are available, none may carry more than half of the external part of the score, and preference leaderboards together carry at most a quarter of it. Direct benchmark evidence then adds a correction of at most 3 points, scaled by how independent that evidence is. An overall position is Supported when two external families from different evidence classes agree and no single source carries more than half, or when two families are confirmed by broad independent benchmark evidence that lands within 3 points of them. A preference leaderboard plus one index is not Supported unless benchmarks confirm it. A model with no external source is scored from its benchmark evidence; its overall position stays Estimated, and only a category position can be Supported by broad independent benchmarks alone. A category position is also Supported when its external category sources span two evidence classes, or when an external category source is backed by independent benchmark evidence from at least two benchmark families. Every other position is Estimated, and Supported and Estimated rows share one ordering. Agentic and Coding combine external category signals with benchmark evidence. Knowledge has no external category source, so its Supported rows come from broad independent benchmark evidence. Without an external category source, a category position needs two benchmark families and a published overall position; a single benchmark is not a category measurement. Where category evidence is thin, all three lenses pool toward the model’s general capability, estimated without preference leaderboards.
Runtime metrics
Runtime metrics (tokens/sec, time-to-first-token) stay separate from ranking and are shown as operational metadata only. They do not affect overall or category scores.
Source refresh: 2026-09-23
BenchLM defaults and caveats
BenchLM uses benchmark freshness as a product layer, not as a claim about an official benchmark maintainer. A benchmark marked Current means BenchLM still treats it as a strong differentiator. A benchmark marked Stale means it is still useful for context but is no longer relied on heavily to separate frontier models.
Public BenchLM benchmark tables default to exact-source rows only. Generated benchmark values remain excluded from BenchAlign evidence.
Overall, Agentic, Coding, and Knowledge use BenchAlign v5.7. We rebuild the rankings from current inputs on every build rather than freezing them at launch. Other category pages remain evidence lenses on the legacy category score until their own validation is complete.
BenchLM does not estimate runtime metrics when no sourced runtime snapshot is available. The leaderboard and model pages show N/A instead. Pricing sort uses the average of input and output token price so models can be ranked on a single cost column while the full input/output pair remains visible.
Last benchmark dataset refresh: September 27, 2026. For raw benchmark exploration, use the benchmark directory. For current provider rollups, use provider pages.