Skip to main content
BenchLM
Citable dataset

LLM Benchmark Statistics (2026)

Updated September 29, 2026 · Auto-generated from BenchLM's live dataset on every data refresh

As of September 29, 2026, BenchLM tracks 486 LLM benchmarks across 10 categories for 523 AI models.

Benchmarks tracked

As of September 29, 2026, BenchLM tracks 486 LLM benchmarks across 10 categories for 523 AI models.

486 across 10 categories

Benchmark saturation rate (top score ≥ 90/100)

37% of the 246 percentage-scaled benchmarks with meaningful coverage on BenchLM are saturated — the top model already scores 90 or higher — as of September 29, 2026.

37% (90 of 246)

Models with a BenchAlign score

210 of the 523 AI models tracked by BenchLM (40%) receive a BenchAlign score as of September 29, 2026; lower-support positions are marked Estimated.

210 of 523

SWE-bench Pro leader

As of September 29, 2026, Anthropic's Claude Opus 5.5 leads the SWE-bench Pro rows tracked by BenchLM at 89.9%.

Claude Opus 5.5 (89.9%)

SWE-bench Verified leader

As of September 29, 2026, Anthropic's Claude Opus 5 leads the SWE-bench Verified rows tracked by BenchLM at 96%.

Claude Opus 5 (96%)

LiveCodeBench leader

As of September 29, 2026, Alibaba's Qwen3.7 Max leads the LiveCodeBench rows tracked by BenchLM at 91.6%.

Qwen3.7 Max (91.6%)

Methodology & sources

Saturation is computed over percentage-scaled benchmarks where at least 3 tracked models have scores; a benchmark counts as saturated when the top model scores 90 or higher. BenchAlign v5.7 keeps sparse rows visible, calibrates evidence by source, and marks lower-support positions Estimated.

Cite these statistics

Every number on this page is generated from BenchLM's live dataset and refreshed with each data update. Link any statistic directly using its anchor, or cite the page as:

BenchLM.ai, "LLM Statistics" (September 29, 2026), https://benchlm.ai/stats/benchmarks

Questions

How many LLM benchmarks are there?

BenchLM tracks 486 LLM benchmarks across 10 categories as of September 29, 2026. The broader ecosystem is larger, but these are the benchmarks with usable, sourced scores across models.

How many LLM benchmarks are saturated?

37% of percentage-scaled benchmarks with meaningful coverage on BenchLM (90 of 246) are saturated, meaning the top model already scores 90 or higher, as of September 29, 2026.

Which model leads the current coding benchmarks?

As of September 29, 2026, Claude Opus 5.5 leads SWE-bench Pro at 89.9%, Claude Opus 5 leads SWE-bench Verified at 96%, and Qwen3.7 Max leads LiveCodeBench at 91.6%. These are separate protocols, so their percentages should not be compared across columns.

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.