Citable dataset
Local LLM Hardware Statistics (2026)
Updated September 10, 2026 · Auto-generated from BenchLM's live dataset on every data refresh
As of September 10, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.
Key statistics
- As of September 10, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.
- As of September 10, 2026, 2 of the 12 ranked local-model receipts on BenchLM use one 24GB NVIDIA RTX 4090 as the reference configuration.
- As of September 10, 2026, Google's Gemma 4 31B is the highest-scoring model in BenchLM's documented single-GPU subset at 52.6/100, with a 24GB RTX 4090 reference setup.
- As of September 10, 2026, 6 of 12 ranked models in BenchLM's deployment subset (50%) use an OSI-approved license.
As of September 10, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.
12 models
As of September 10, 2026, 2 of the 12 ranked local-model receipts on BenchLM use one 24GB NVIDIA RTX 4090 as the reference configuration.
2 models
Methodology & sources
This page joins the current ranked open-weight cohort to BenchLM's reviewed self-host catalog, last checked 2026-06-12. A “single-GPU” row has a documented reference configuration using at most 32GB of GPU or unified memory. These receipts are deployment examples, not a claim that every runtime, quantization, or context length will fit.
Cite these statistics
Every number on this page is generated from BenchLM's live dataset and refreshed with each data update. Link any statistic directly using its anchor, or cite the page as:
BenchLM.ai, "LLM Statistics" (September 10, 2026), https://benchlm.ai/stats/local-llm-hardwareFrequently Asked Questions
How many ranked local LLMs fit on one 24GB GPU?
As of September 10, 2026, 2 of the 12 ranked models with reviewed deployment receipts use one 24GB RTX 4090 as their reference setup. This count covers the documented quantization and configuration only; longer contexts and different runtimes can require more memory.
What is the highest-scoring LLM documented for one GPU?
As of September 10, 2026, Google's Gemma 4 31B leads BenchLM's documented single-GPU subset at 52.6/100. Its reference row uses one 24GB RTX 4090. The score measures benchmark capability, not local throughput, latency, or the quality loss from a specific quantization.
Does open-weight mean open source?
No. As of September 10, 2026, 6 of 12 ranked models in BenchLM's deployment subset use an OSI-approved license. “Open-weight” only means the weights are downloadable; community and custom licenses can restrict commercial use, redistribution, or derived models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.