# LLM Benchmark Statistics (2026)

> As of September 30, 2026, BenchLM tracks 486 LLM benchmarks across 10 categories for 523 AI models. Updated September 30, 2026; every number regenerates from BenchLM's live dataset on each data refresh.

## Key statistics

- As of September 30, 2026, BenchLM tracks 486 LLM benchmarks across 10 categories for 523 AI models.
- 37% of the 246 percentage-scaled benchmarks with meaningful coverage on BenchLM are saturated — the top model already scores 90 or higher — as of September 30, 2026.
- 210 of the 523 AI models tracked by BenchLM (40%) receive a BenchAlign score as of September 30, 2026; lower-support positions are marked Estimated.
- As of September 30, 2026, Anthropic's Claude Opus 5.5 leads the SWE-bench Pro rows tracked by BenchLM at 89.9%.
- As of September 30, 2026, Anthropic's Claude Opus 5 leads the SWE-bench Verified rows tracked by BenchLM at 96%.
- As of September 30, 2026, Alibaba's Qwen3.7 Max leads the LiveCodeBench rows tracked by BenchLM at 91.6%.

## Methodology

Saturation is computed over percentage-scaled benchmarks where at least 3 tracked models have scores; a benchmark counts as saturated when the top model scores 90 or higher. BenchAlign v5.7 keeps sparse rows visible, calibrates evidence by source, and marks lower-support positions Estimated.

## FAQ

### How many LLM benchmarks are there?

BenchLM tracks 486 LLM benchmarks across 10 categories as of September 30, 2026. The broader ecosystem is larger, but these are the benchmarks with usable, sourced scores across models.

### How many LLM benchmarks are saturated?

37% of percentage-scaled benchmarks with meaningful coverage on BenchLM (90 of 246) are saturated, meaning the top model already scores 90 or higher, as of September 30, 2026.

### Which model leads the current coding benchmarks?

As of September 30, 2026, Claude Opus 5.5 leads SWE-bench Pro at 89.9%, Claude Opus 5 leads SWE-bench Verified at 96%, and Qwen3.7 Max leads LiveCodeBench at 91.6%. These are separate protocols, so their percentages should not be compared across columns.

## Sources

- [Benchmark directory](https://benchlm.ai/benchmarks)
- [SWE-bench Pro results](https://benchlm.ai/benchmarks/swe-bench-pro)
- [SWE-bench Verified results](https://benchlm.ai/benchmarks/swe-bench-verified)
- [LiveCodeBench results](https://benchlm.ai/benchmarks/livecodebench)
- [Benchmark confidence & contamination](https://benchlm.ai/benchmark-confidence)
- [Methodology](https://benchlm.ai/methodology)
- [Benchmarks JSON (machine-readable)](https://benchlm.ai/data/benchmarks.json)

Cite as: BenchLM.ai, "LLM Statistics" (September 30, 2026), https://benchlm.ai/stats/benchmarks

Canonical page: https://benchlm.ai/stats/benchmarks
