BenchLM recommendation
Best Open Source LLMs in 2026
As of July 20, 2026, the top model in best open source llms on the BenchLM leaderboard is MiniMax M3 with a score of 69.8.
Last verified: July 20, 2026
This page is the canonical open-weight ranking. It uses the same public BenchAlign v5 overall lane as the main leaderboard, then filters to downloadable model weights. The score compares measured capability; it does not decide whether a license is permissive, a model fits your hardware, or self-hosting beats an API on cost. Use the linked decision guide for those deployment questions.
Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
Bottom line: use the live table above for capability order, then check evidence status, license terms, memory requirements, and serving cost before choosing a deployment.
MiniMax M3 leads this ranking with a score of 69.8, followed by GLM-5.1 (67.7) and Inkling (67.5). The top three are separated by just a few points — any of them would perform well for this use case.
All models in this ranking are open-weight, meaning they can be self-hosted for maximum control and cost efficiency.
This ranking is based on provisional overall weighted scores across BenchLM.ai's scoring formula tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
MiniMax M3
MiniMax · 1M
Current public open-weight overall leader with Supported evidence.
GLM-5.1
Z.AI · 203K
Supported top-tier open-weight row with broad coverage.
Inkling
Thinking Machines Lab · 1M
What changed
MiniMax M3 leads the public open-weight overall lane with Supported evidence.
GLM-5.1 follows on the live overall table with Supported evidence.
GLM-5 rounds out the current top three open-weight rows.
How to choose
Best open-weight model overall?
Start with the live #1, then check its evidence label and deployment constraints
Coding-focused open model?
Use the open rows on the calibrated coding leaderboard
Self-hosting economics?
Estimate hardware and API break-even before choosing a model family
Chinese-lab open models?
See the Chinese model rankings — most frontier open weights come from these labs
Full Rankings (86 models)
59.5
BenchAlign v5
90% interval 47.98–71.01
58.2
BenchAlign v5
90% interval 46.64–69.66
58
BenchAlign v5
90% interval 46.50–69.53
55.9
BenchAlign v5
90% interval 44.43–67.46
55.5
BenchAlign v5
90% interval 43.95–66.98
54
BenchAlign v5
90% interval 42.44–65.47
53.4
BenchAlign v5
90% interval 34.72–72.14
52.3
BenchAlign v5
90% interval 40.28–64.36
51
BenchAlign v5
90% interval 39.46–62.48
50.3
BenchAlign v5
90% interval 38.77–61.79
50.1
BenchAlign v5
90% interval 38.57–61.60
49.3
BenchAlign v5
90% interval 37.83–60.86
48.5
BenchAlign v5
90% interval 37.03–60.06
44.2
BenchAlign v5
90% interval 32.73–55.76
43.9
BenchAlign v5
90% interval 32.35–55.38
42.6
BenchAlign v5
90% interval 31.07–54.09
40.4
BenchAlign v5
90% interval 28.92–51.95
40.4
BenchAlign v5
90% interval 28.87–51.90
39.5
BenchAlign v5
90% interval 28.00–51.03
39.1
BenchAlign v5
90% interval 27.55–50.58
34.7
BenchAlign v5
90% interval 18.46–50.94
Key Takeaways
The top model is MiniMax M3 by MiniMax with a BenchAlign v5 score of 69.8 and Supported evidence.
The best open-weight model is MiniMax M3 at position #1.
86 models are included in this ranking.
Score in Context
What these scores mean
Open-weight models are ranked by the same public BenchAlign v5 overall score as proprietary models. Supported and Estimated describe the evidence behind each position; they are not separate leaderboards.
Known limitations
Open weight is not the same as OSI-approved open source. The ranking does not score license restrictions, memory requirements, serving throughput, fine-tuning support, or the engineering cost of operating the model.
Best Open Source LLMs FAQ
What is the best open source LLM right now?
The live table above is the ranking owner and recomputes from the public BenchAlign v5 artifact. MiniMax M3 leads the July 14 snapshot with Supported evidence; use the table rather than a copied winner sentence when the data changes.
Are open source LLMs as good as GPT or Claude?
They can match proprietary models on individual tasks, but the July 14 unified overall ranking still shows a double-digit gap between the open-weight and proprietary leaders. Deployment control, privacy, and serving economics can still make an open model the better operational choice.
What is the best open source LLM for coding?
Use the open-weight rows on the live coding leaderboard. Coding order differs from overall order, and a single HumanEval or LiveCodeBench result should not be treated as a complete coding verdict.
Can I run these models locally?
The smaller open-weight rows run on a single consumer GPU with 4-bit quantization; frontier-size models need multi-GPU rigs or high-memory Apple Silicon. The local LLM guide breaks the rankings down by VRAM tier, and the Ollama guide includes pull commands and size estimates per model.
Explore More
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.