Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Citable dataset

LLM Benchmark Statistics (2026)

Updated August 15, 2026 · Auto-generated from BenchLM's live dataset on every data refresh

As of August 15, 2026, BenchLM tracks 437 LLM benchmarks across 10 categories for 394 AI models.

Benchmarks tracked

As of August 15, 2026, BenchLM tracks 437 LLM benchmarks across 10 categories for 394 AI models.

437 across 10 categories

Benchmark saturation rate (top score ≥ 90/100)

37% of the 210 percentage-scaled benchmarks with meaningful coverage on BenchLM are saturated — the top model already scores 90 or higher — as of August 15, 2026.

37% (77 of 210)

Models with a BenchAlign score

218 of the 394 AI models tracked by BenchLM (55%) receive a BenchAlign score as of August 15, 2026; lower-support positions are marked Estimated.

218 of 394

SWE-bench Pro leader

As of August 15, 2026, Anthropic's Claude Mythos 5 leads the SWE-bench Pro rows tracked by BenchLM at 80.3%.

Claude Mythos 5 (80.3%)

SWE-bench Verified leader

As of August 15, 2026, Anthropic's Claude Opus 5 leads the SWE-bench Verified rows tracked by BenchLM at 96%.

Claude Opus 5 (96%)

LiveCodeBench leader

As of August 15, 2026, Alibaba's Qwen3.7 Max leads the LiveCodeBench rows tracked by BenchLM at 91.6%.

Qwen3.7 Max (91.6%)

Methodology & sources

Saturation is computed over percentage-scaled benchmarks where at least 3 tracked models have scores; a benchmark counts as saturated when the top model scores 90 or higher. BenchAlign v5 keeps sparse rows visible, calibrates evidence by source, and marks lower-support positions Estimated.

Cite these statistics

Every number on this page is generated from BenchLM's live dataset and refreshed with each data update. Link any statistic directly using its anchor, or cite the page as:

BenchLM.ai, "LLM Statistics" (August 15, 2026), https://benchlm.ai/stats/benchmarks

Frequently Asked Questions

How many LLM benchmarks are there?

BenchLM tracks 437 LLM benchmarks across 10 categories as of August 15, 2026. The broader ecosystem is larger, but these are the benchmarks with usable, sourced scores across models.

How many LLM benchmarks are saturated?

37% of percentage-scaled benchmarks with meaningful coverage on BenchLM (77 of 210) are saturated, meaning the top model already scores 90 or higher, as of August 15, 2026.

Which model leads the current coding benchmarks?

As of August 15, 2026, Claude Mythos 5 leads SWE-bench Pro at 80.3%, Claude Opus 5 leads SWE-bench Verified at 96%, and Qwen3.7 Max leads LiveCodeBench at 91.6%. These are separate protocols, so their percentages should not be compared across columns.

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.