Modified MIT
Community/custom
1000B MoE (32B active)
INT4, FP8, Q4
8× NVIDIA H100 (80GB) · 640GB total
The best open-source LLM by current open-weight benchmark score is MiMo-V2.6-Pro. It leads the September 2026 open-weight ranking at 74.2, ahead of Qwen3.8 Max (71.4) and GLM-5.3 (65.1). The license directory below separates OSI-approved licenses from community terms.
Citable stat 94 open-weight models are ranked as of September 29, 2026; MiMo-V2.6-Pro leads at 74.2/100.
Bottom line: use the live table above for capability order, then check evidence status, license terms, memory requirements, and serving cost before choosing a deployment.
This ranking moves when models ship. Get the releases, price changes and retirements that affect your shortlist. Follow model changes
63.9
BenchAlign v5.7
90% interval 52.41–75.43
62.6
BenchAlign v5.7
90% interval 52.73–72.47
60.8
BenchAlign v5.7
90% interval 49.29–72.31
55.5
BenchAlign v5.7
90% interval 48.47–62.50
55.3
BenchAlign v5.7
90% interval 43.82–66.84
54.1
BenchAlign v5.7
90% interval 44.00–64.18
52.3
BenchAlign v5.7
90% interval 42.45–62.19
52.3
BenchAlign v5.7
90% interval 42.45–62.19
52.3
BenchAlign v5.7
90% interval 42.45–62.19
52.3
BenchAlign v5.7
90% interval 42.45–62.19
52.2
BenchAlign v5.7
90% interval 42.35–62.09
50.3
BenchAlign v5.7
90% interval 41.87–58.66
44.5
BenchAlign v5.7
90% interval 33.00–56.03
42.4
BenchAlign v5.7
90% interval 30.89–53.91
39.9
BenchAlign v5.7
90% interval 19.23–60.48
38.3
BenchAlign v5.7
90% interval 15.86–60.75
36.1
BenchAlign v5.7
90% interval 24.57–47.60
34.1
BenchAlign v5.7
90% interval 13.54–54.72
33.7
BenchAlign v5.7
90% interval 22.23–45.25
31.9
BenchAlign v5.7
90% interval 21.98–41.72
31.4
BenchAlign v5.7
90% interval 19.84–42.87
31.1
BenchAlign v5.7
90% interval 19.61–42.64
30.7
BenchAlign v5.7
90% interval 19.18–42.21
30.6
BenchAlign v5.7
90% interval 19.06–42.09
30.5
BenchAlign v5.7
90% interval 19.94–40.98
27.1
BenchAlign v5.7
90% interval 10.89–43.38
21.2
BenchAlign v5.7
90% interval 3.76–38.60
18.9
BenchAlign v5.7
90% interval 9.05–28.79
MiMo-V2.6-Pro leads the live open-weight ranking at 74.2 with Estimated evidence.
Qwen3.8 Max ranks #2 at 71.4 with Supported evidence.
GLM-5.3 ranks #3 at 65.1 with Estimated evidence.
The top model is MiMo-V2.6-Pro by Xiaomi with a BenchAlign v5.7 score of 74.2 and Estimated evidence.
The best open-weight model is MiMo-V2.6-Pro at position #1.
94 models are included in this ranking.
Open-weight models are ranked by the same public BenchAlign v5.7 overall score as proprietary models. Supported and Estimated describe the evidence behind each position; they are not separate leaderboards.
Open weight is not the same as OSI-approved open source. The ranking does not score license restrictions, memory requirements, serving throughput, fine-tuning support, or the engineering cost of operating the model.
Ranking data as of September 29, 2026
This page is the canonical open-weight ranking. It uses the same public BenchAlign v5.7 overall lane as the main leaderboard, then filters to downloadable model weights. The score compares measured capability; it does not decide whether a license is permissive, a model fits your hardware, or self-hosting beats an API on cost. Use the linked decision guide for those deployment questions.
This is the public BenchAlign v5.7 overall lane filtered to open-weight models. Scores measure capability; evidence labels and score intervals show how much confidence to place in close comparisons.
The open-weight slice starts with MiMo-V2.6-Pro, followed by Qwen3.8 Max and GLM-5.3. All rows use the public BenchAlign v5.7 projection contract. Evidence badges and score intervals matter as much as small point gaps because public source coverage is uneven.
Every ranked model publishes downloadable weights, but that does not make every license OSI-approved or every deployment practical. Use the directory below to check license terms, quantization, and reference hardware before shortlisting a model.
This ranking uses the public BenchAlign v5.7 overall contract filtered to open-weight models. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.
Deployment evidence
The performance table ranks every eligible open-weight row. This smaller directory includes only models with a deployment record in the self-host catalog, so license and hardware claims remain auditable instead of being filled from family names.
Showing 12 of 12 deployment-documented rows
Deployment catalog checked 2026-06-12.
Modified MIT
Community/custom
1000B MoE (32B active)
INT4, FP8, Q4
8× NVIDIA H100 (80GB) · 640GB total
MIT
OSI-approved
744B MoE (40B active)
FP8, Q4, Q2
8× NVIDIA H100 (80GB) · 640GB total
Moonshot
Community/custom
120B dense
FP8, Q4, Q2
4× NVIDIA A100 (80GB) · 320GB total
Modified MIT
Community/custom
1000B MoE (32B active)
INT4, FP8, Q4
8× NVIDIA H100 (80GB) · 640GB total
Apache-2.0
OSI-approved
27B dense
FP8, Q4
1× NVIDIA RTX 4090 (24GB) · 24GB total
Apache-2.0
OSI-approved
31B dense
BF16, Q4
1× NVIDIA RTX 4090 (24GB) · 24GB total
MIT
OSI-approved
671B MoE (37B active)
FP8, Q4
8× NVIDIA H100 (80GB) · 640GB total
Apache-2.0
OSI-approved
119B MoE (22B active)
FP8, Q4
1× NVIDIA H100 (80GB) · 80GB total
MIT
OSI-approved
671B MoE (37B active)
FP8, Q4
8× NVIDIA H100 (80GB) · 640GB total
Qwen Research
Community/custom
72B dense
FP8, Q4
1× NVIDIA H100 (80GB) · 80GB total
Llama Community
Community/custom
109B MoE (17B active)
FP8, Q4
1× NVIDIA H100 (80GB) · 80GB total
Llama Community
Community/custom
400B MoE (17B active)
FP8, Q4
2× NVIDIA A100 (80GB) · 160GB total
This table is the sourced deployment subset, not the complete performance ranking. A missing row means the deployment catalog is incomplete, not that the model cannot be self-hosted.
The #1 row of the open-weight table above is the current answer, and the table recomputes from current ranking data on every build. The order can move when new evidence lands, so check the date on the ranking and the evidence label beside the leader before treating it as settled.
They can match proprietary models on individual tasks. The overall gap moves with every release, so compare the first row of the ranking above with the top of the overall leaderboard. Deployment control, privacy, and serving economics can still make an open model the better operational choice.
Use the open-weight rows on the live coding leaderboard. Coding order differs from overall order, and a single HumanEval or LiveCodeBench result should not be treated as a complete coding verdict.
The smaller open-weight rows run on a single consumer GPU with 4-bit quantization; frontier-size models need multi-GPU rigs or high-memory Apple Silicon. The local LLM guide breaks the rankings down by VRAM tier, and the Ollama guide includes pull commands and size estimates per model.
The current open-weight top tier comes almost entirely from Chinese labs — DeepSeek, Zhipu (GLM), Moonshot (Kimi), Alibaba (Qwen), and MiniMax — with Meta and Mistral behind on the live ranking. Family strengths differ by task, so check the per-category leaderboards and the Chinese model rankings rather than picking by lab reputation alone.
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.