Citable dataset
LLM Benchmark Statistics (2026)
Updated August 15, 2026 · Auto-generated from BenchLM's live dataset on every data refresh
As of August 15, 2026, BenchLM tracks 437 LLM benchmarks across 10 categories for 394 AI models.
Key statistics
- As of August 15, 2026, BenchLM tracks 437 LLM benchmarks across 10 categories for 394 AI models.
- 37% of the 210 percentage-scaled benchmarks with meaningful coverage on BenchLM are saturated — the top model already scores 90 or higher — as of August 15, 2026.
- 218 of the 394 AI models tracked by BenchLM (55%) receive a BenchAlign score as of August 15, 2026; lower-support positions are marked Estimated.
- As of August 15, 2026, Anthropic's Claude Mythos 5 leads the SWE-bench Pro rows tracked by BenchLM at 80.3%.
- As of August 15, 2026, Anthropic's Claude Opus 5 leads the SWE-bench Verified rows tracked by BenchLM at 96%.
- As of August 15, 2026, Alibaba's Qwen3.7 Max leads the LiveCodeBench rows tracked by BenchLM at 91.6%.
As of August 15, 2026, BenchLM tracks 437 LLM benchmarks across 10 categories for 394 AI models.
437 across 10 categories
37% of the 210 percentage-scaled benchmarks with meaningful coverage on BenchLM are saturated — the top model already scores 90 or higher — as of August 15, 2026.
37% (77 of 210)
218 of the 394 AI models tracked by BenchLM (55%) receive a BenchAlign score as of August 15, 2026; lower-support positions are marked Estimated.
218 of 394
As of August 15, 2026, Anthropic's Claude Mythos 5 leads the SWE-bench Pro rows tracked by BenchLM at 80.3%.
Claude Mythos 5 (80.3%)
As of August 15, 2026, Anthropic's Claude Opus 5 leads the SWE-bench Verified rows tracked by BenchLM at 96%.
Claude Opus 5 (96%)
Methodology & sources
Saturation is computed over percentage-scaled benchmarks where at least 3 tracked models have scores; a benchmark counts as saturated when the top model scores 90 or higher. BenchAlign v5 keeps sparse rows visible, calibrates evidence by source, and marks lower-support positions Estimated.
Cite these statistics
Every number on this page is generated from BenchLM's live dataset and refreshed with each data update. Link any statistic directly using its anchor, or cite the page as:
BenchLM.ai, "LLM Statistics" (August 15, 2026), https://benchlm.ai/stats/benchmarksFrequently Asked Questions
How many LLM benchmarks are there?
BenchLM tracks 437 LLM benchmarks across 10 categories as of August 15, 2026. The broader ecosystem is larger, but these are the benchmarks with usable, sourced scores across models.
How many LLM benchmarks are saturated?
37% of percentage-scaled benchmarks with meaningful coverage on BenchLM (77 of 210) are saturated, meaning the top model already scores 90 or higher, as of August 15, 2026.
Which model leads the current coding benchmarks?
As of August 15, 2026, Claude Mythos 5 leads SWE-bench Pro at 80.3%, Claude Opus 5 leads SWE-bench Verified at 96%, and Qwen3.7 Max leads LiveCodeBench at 91.6%. These are separate protocols, so their percentages should not be compared across columns.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.