Skip to main content

Citable dataset

Local LLM Hardware Statistics (2026)

Updated July 27, 2026 · Auto-generated from BenchLM's live dataset on every data refresh

As of July 27, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.

Ranked open-weight models with deployment receipts

As of July 27, 2026, BenchLM documents license, quantization, and reference hardware for 12 ranked open-weight models.

12 models

Models with a documented single-24GB-GPU setup

As of July 27, 2026, 2 of the 12 ranked local-model receipts on BenchLM use one 24GB NVIDIA RTX 4090 as the reference configuration.

2 models

Highest-scoring documented single-GPU model

As of July 27, 2026, Google's Gemma 4 31B is the highest-scoring model in BenchLM's documented single-GPU subset at 60.2/100, with a 24GB RTX 4090 reference setup.

Gemma 4 31B (60.2/100)

OSI-approved licenses in the deployment subset

As of July 27, 2026, 6 of 12 ranked models in BenchLM's deployment subset (50%) use an OSI-approved license.

50% (6 of 12)

Methodology & sources

This page joins the current ranked open-weight cohort to BenchLM's reviewed self-host catalog, last checked 2026-06-12. A “single-GPU” row has a documented reference configuration using at most 32GB of GPU or unified memory. These receipts are deployment examples, not a claim that every runtime, quantization, or context length will fit.

Cite these statistics

Every number on this page is generated from BenchLM's live dataset and refreshed with each data update. Link any statistic directly using its anchor, or cite the page as:

BenchLM.ai, "LLM Statistics" (July 27, 2026), https://benchlm.ai/stats/local-llm-hardware

Frequently Asked Questions

How many ranked local LLMs fit on one 24GB GPU?

As of July 27, 2026, 2 of the 12 ranked models with reviewed deployment receipts use one 24GB RTX 4090 as their reference setup. This count covers the documented quantization and configuration only; longer contexts and different runtimes can require more memory.

What is the highest-scoring LLM documented for one GPU?

As of July 27, 2026, Google's Gemma 4 31B leads BenchLM's documented single-GPU subset at 60.2/100. Its reference row uses one 24GB RTX 4090. The score measures benchmark capability, not local throughput, latency, or the quality loss from a specific quantization.

Does open-weight mean open source?

No. As of July 27, 2026, 6 of 12 ranked models in BenchLM's deployment subset use an OSI-approved license. “Open-weight” only means the weights are downloadable; community and custom licenses can restrict commercial use, redistribution, or derived models.

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

One email each week. Unsubscribe anytime.