Citable dataset
Local LLM Hardware Statistics (2026)
Updated July 27, 2026 · Auto-generated from BenchLM's live dataset on every data refresh
As of July 27, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.
Key statistics
- As of July 27, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.
- As of July 27, 2026, 2 of the 12 ranked local-model receipts on BenchLM use one 24GB NVIDIA RTX 4090 as the reference configuration.
- As of July 27, 2026, Google's Gemma 4 31B is the highest-scoring model in BenchLM's documented single-GPU subset at 60.2/100, with a 24GB RTX 4090 reference setup.
- As of July 27, 2026, 6 of 12 ranked models in BenchLM's deployment subset (50%) use an OSI-approved license.
As of July 27, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.
12 models
As of July 27, 2026, 2 of the 12 ranked local-model receipts on BenchLM use one 24GB NVIDIA RTX 4090 as the reference configuration.
2 models
Methodology & sources
This page joins the current ranked open-weight cohort to BenchLM's reviewed self-host catalog, last checked 2026-06-12. A “single-GPU” row has a documented reference configuration using at most 32GB of GPU or unified memory. These receipts are deployment examples, not a claim that every runtime, quantization, or context length will fit.
Cite these statistics
Every number on this page is generated from BenchLM's live dataset and refreshed with each data update. Link any statistic directly using its anchor, or cite the page as:
BenchLM.ai, "LLM Statistics" (July 27, 2026), https://benchlm.ai/stats/local-llm-hardwareFrequently Asked Questions
How many ranked local LLMs fit on one 24GB GPU?
As of July 27, 2026, 2 of the 12 ranked models with reviewed deployment receipts use one 24GB RTX 4090 as their reference setup. This count covers the documented quantization and configuration only; longer contexts and different runtimes can require more memory.
What is the highest-scoring LLM documented for one GPU?
As of July 27, 2026, Google's Gemma 4 31B leads BenchLM's documented single-GPU subset at 60.2/100. Its reference row uses one 24GB RTX 4090. The score measures benchmark capability, not local throughput, latency, or the quality loss from a specific quantization.
Does open-weight mean open source?
No. As of July 27, 2026, 6 of 12 ranked models in BenchLM's deployment subset use an OSI-approved license. “Open-weight” only means the weights are downloadable; community and custom licenses can restrict commercial use, redistribution, or derived models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
One email each week. Unsubscribe anytime.