Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief
BenchLM recommendation

Open-Source LLM Leaderboard 2026

Data verified

The best open-source LLM by current open-weight benchmark score is Qwen3.8 Max. It leads the September 2026 open-weight ranking at 78.7, ahead of Hy4 preview (78.3) and Qwen3.8-27B (72). The license directory below separates OSI-approved licenses from community terms.

Citable stat 106 open-weight models are ranked as of September 2, 2026; Qwen3.8 Max leads at 78.7/100.

Bottom line: use the live table above for capability order, then check evidence status, license terms, memory requirements, and serving cost before choosing a deployment.

About this ranking

Last verified: September 2, 2026

This page is the canonical open-weight ranking. It uses the same public BenchAlign v5 overall lane as the main leaderboard, then filters to downloadable model weights. The score compares measured capability; it does not decide whether a license is permissive, a model fits your hardware, or self-hosting beats an API on cost. Use the linked decision guide for those deployment questions.

This is the public BenchAlign v5 overall lane filtered to open-weight models. Scores measure capability; evidence labels and score intervals show how much confidence to place in close comparisons.

The open-weight slice starts with Qwen3.8 Max, followed by Hy4 preview and Qwen3.8-27B. All rows use the public BenchAlign v5 projection contract. Evidence badges and score intervals matter as much as small point gaps because public source coverage is uneven.

Every ranked model publishes downloadable weights, but that does not make every license OSI-approved or every deployment practical. Use the directory below to check license terms, quantization, and reference hardware before shortlisting a model.

This ranking uses the public BenchAlign v5 overall contract filtered to open-weight models. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

Deployment evidence

Compare license, size, and deployment

The performance table ranks every eligible open-weight row. This smaller directory includes only models with a deployment record in the self-host catalog, so license and hardware claims remain auditable instead of being filled from family names.

Open weight is an access category, not a license verdict. “OSI-approved” appears only when the deployment catalog records MIT or Apache 2.0. Community and custom licenses stay separate, even when their weights are downloadable.

Showing 12 of 12 deployment-documented rows

Deployment catalog checked 2026-06-12.

  • ModelGLM-5.1

    Z.AI · supported

    Score66.8
    License

    MIT

    OSI-approved

    Parameters

    744B MoE (40B active)

    FP8, Q4, Q2

    Context203K
    Reference hardware

    8× NVIDIA H100 (80GB) · 640GB total

  • ModelKimi K2.6

    Moonshot AI · estimated

    Score59.2
    License

    Modified MIT

    Community/custom

    Parameters

    1000B MoE (32B active)

    INT4, FP8, Q4

    Context256K
    Reference hardware

    8× NVIDIA H100 (80GB) · 640GB total

  • ModelKimi K2.5

    Moonshot AI · supported

    Score58.9
    License

    Moonshot

    Community/custom

    Parameters

    120B dense

    FP8, Q4, Q2

    Context256K
    Reference hardware

    4× NVIDIA A100 (80GB) · 320GB total

This table is the sourced deployment subset, not the complete performance ranking. A missing row means the deployment catalog is incomplete, not that the model cannot be self-hosted.

What changed

Qwen3.8 Max leads the live open-weight ranking at 78.7 with Supported evidence.

Hy4 preview ranks #2 at 78.3 with Estimated evidence.

Qwen3.8-27B ranks #3 at 72 with Supported evidence.

How to choose

Full Rankings (106 models)

1
Qwen3.8 Max
Alibaba·Open Weight·1M

78.7

BenchAlign v5

Supported

90% interval 76.51–80.80

2
Hy4 preview
Tencent·Open Weight·1M

78.3

BenchAlign v5

Estimated

90% interval 68.45–88.19

3
Qwen3.8-27B
Alibaba·Open Weight·262K

72

BenchAlign v5

Supported

90% interval 69.78–74.15

4
dots3-note Preview
Dots Studio·Open Weight·512K

68.7

BenchAlign v5

Estimated

90% interval 58.80–78.54

5
MiniMax M3
MiniMax·Open Weight·1M

68.5

BenchAlign v5

Supported

90% interval 62.55–74.34

6
Hy3
Tencent·Open Weight·256K

67.8

BenchAlign v5

Supported

90% interval 58.82–76.68

7
Ornith-1.5-397B
Ornith AI·Open Weight·262K

67.4

BenchAlign v5

Estimated

90% interval 57.52–77.26

8
GLM-5.3
Z.AI·Open Weight·1M

67

BenchAlign v5

Estimated

90% interval 55.49–78.52

9
GLM-5.2
Z.AI·Open Weight·1M

66.9

BenchAlign v5

Estimated

90% interval 52.78–81.03

10
GLM-5.1
Z.AI·Open Weight·203K

66.8

BenchAlign v5

Supported

90% interval 55.94–77.67

11
Inkling
Thinking Machines Lab·Open Weight·1M

66.7

BenchAlign v5

Supported

90% interval 59.06–74.26

12
GLM-5
Z.AI·Open Weight·200K

65.6

BenchAlign v5

Supported

90% interval 53.85–77.43

13
Inkling-Small
Thinking Machines Lab·Open Weight·1M

63.6

BenchAlign v5

Supported

90% interval 57.95–69.23

14
MiniMax M2.7
MiniMax·Open Weight·200K

62.9

BenchAlign v5

Supported

90% interval 55.03–70.83

15
GLM-5.3-Flash
Z.AI·Open Weight·1M

61.7

BenchAlign v5

Estimated

90% interval 50.16–73.18

16
Agents-A1
InternScience·Open Weight·262K

61.2

BenchAlign v5

Estimated

90% interval 51.35–71.09

17
Agents-A1-F16-GGUF
InternScience·Open Weight·262K

61.2

BenchAlign v5

Estimated

90% interval 51.35–71.09

18
Agents-A1-FP8
InternScience·Open Weight·262K

61.2

BenchAlign v5

Estimated

90% interval 51.35–71.09

19
Agents-A1-Q4_K_M-GGUF
InternScience·Open Weight·262K

61.2

BenchAlign v5

Estimated

90% interval 51.35–71.09

20
Agents-A1-Q8_0-GGUF
InternScience·Open Weight·262K

61.2

BenchAlign v5

Estimated

90% interval 51.35–71.09

21
Qwen3.8-Flash-Next
Alibaba·Open Weight·262K

61.1

BenchAlign v5

Estimated

90% interval 49.54–72.57

22
GLM-5 (Reasoning)
Z.AI·Open Weight·200K

60.8

BenchAlign v5

Estimated

90% interval 49.27–72.30

23
GLM-4.7
Z.AI·Open Weight·200K

60.8

BenchAlign v5

Supported

90% interval 46.31–75.23

24
Qwen3.5 397B (Reasoning)
Alibaba·Open Weight·128K

60.5

BenchAlign v5

Estimated

90% interval 48.99–72.02

25
Gemma 4 31B
Google·Open Weight·256K

60.2

BenchAlign v5

Supported

90% interval 43.39–77.07

26
Qwen3.5-27B
Alibaba·Open Weight·262K

59.8

BenchAlign v5

Supported

90% interval 49.81–69.87

27
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

59.6

BenchAlign v5

Supported

90% interval 47.69–71.54

28
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

59.3

BenchAlign v5

Estimated

90% interval 48.33–70.24

29
Kimi K2.6
Moonshot AI·Open Weight·256K

59.2

BenchAlign v5

Estimated

90% interval 49.31–69.05

30
DeepSeek V3.2 (Thinking)
DeepSeek·Open Weight·128K

59.1

BenchAlign v5

Estimated

90% interval 47.61–70.63

31
Qwen3 235B 2507 (Reasoning)
Alibaba·Open Weight·128K

59

BenchAlign v5

Estimated

90% interval 47.49–70.52

32
Kimi K2.5
Moonshot AI·Open Weight·256K

58.9

BenchAlign v5

Supported

90% interval 50.17–67.57

33
Qwen3.5 397B
Alibaba·Open Weight·128K

58

BenchAlign v5

Estimated

90% interval 46.47–69.50

34
Gemma 4 26B A4B
Google·Open Weight·256K

57.3

BenchAlign v5

Supported

90% interval 39.11–75.40

35
Qwen3 235B 2507
Alibaba·Open Weight·128K

57

BenchAlign v5

Estimated

90% interval 45.46–68.49

36
Trinity-Large-Preview
Arcee AI·Open Weight·512K

56.9

BenchAlign v5

Estimated

90% interval 45.38–68.41

37
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

56.2

BenchAlign v5

Supported

90% interval 44.18–68.29

38
DeepSeek LLM 2.0
DeepSeek·Open Weight·128K

55.5

BenchAlign v5

Estimated

90% interval 43.95–66.98

39
DeepSeek V3.2
DeepSeek·Open Weight·128K

54.8

BenchAlign v5

Supported

90% interval 37.10–72.47

40
GLM-4.6
Z.AI·Open Weight·200K

54.5

BenchAlign v5

Supported

90% interval 35.71–73.30

41
Step 3.7 Flash
StepFun·Open Weight·256K

54.5

BenchAlign v5

Estimated

90% interval 42.94–65.97

42
Step 3.5 Flash
StepFun·Open Weight·256K

54.4

BenchAlign v5

Supported

90% interval 40.15–68.56

43
Nemotron 3 Nano 30B
NVIDIA·Open Weight·32K

53.9

BenchAlign v5

Estimated

90% interval 42.35–65.37

44
Ling 3.0 Flash
InclusionAI·Open Weight·262K

53.8

BenchAlign v5

Estimated

90% interval 42.27–65.30

45
Qwen3.6-27B
Alibaba·Open Weight·262K

53.7

BenchAlign v5

Estimated

90% interval 42.22–65.25

46
MiMo-V2-Flash
Xiaomi·Open Weight·256K

53.4

BenchAlign v5

Supported

90% interval 38.35–68.44

47
Ornith-1.5-35B-A3B
Ornith AI·Open Weight·262K

53.4

BenchAlign v5

Estimated

90% interval 43.50–63.24

48
DeepSeek V3.1
DeepSeek·Open Weight·128K

53.1

BenchAlign v5

Supported

90% interval 33.58–72.58

49
Qwen2.5-72B
Alibaba·Open Weight·128K

53.1

BenchAlign v5

Estimated

90% interval 41.54–64.57

50
DeepSeek V3.1 (Reasoning)
DeepSeek·Open Weight·128K

52.9

BenchAlign v5

Supported

90% interval 33.10–72.69

51
Llama 3.1 405B
Meta·Open Weight·128K

52.6

BenchAlign v5

Estimated

90% interval 41.09–64.12

52
Nemotron 3 Super 120B A12B
NVIDIA·Open Weight·256K

51.9

BenchAlign v5

Estimated

90% interval 40.34–63.37

53
Llama 3 70B
Meta·Open Weight·128K

51.7

BenchAlign v5

Estimated

90% interval 40.19–63.22

54
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

51.5

BenchAlign v5

Estimated

90% interval 39.98–63.01

55
DeepSeek Coder 2.0
DeepSeek·Open Weight·128K

51.2

BenchAlign v5

Estimated

90% interval 39.64–62.67

56
DeepSeek-R1
DeepSeek·Open Weight·128K

51.1

BenchAlign v5

Supported

90% interval 32.58–69.66

57
Nemotron 3 Super 100B
NVIDIA·Open Weight·1M

51

BenchAlign v5

Estimated

90% interval 39.44–62.47

58
DeepSeekMath V2
DeepSeek·Open Weight·128K

50.8

BenchAlign v5

Estimated

90% interval 39.24–62.27

59
Qwen2.5-1M
Alibaba·Open Weight·1M

50.8

BenchAlign v5

Estimated

90% interval 39.24–62.27

60
GLM-4.7-Flash
Z.AI·Open Weight·200K

50.5

BenchAlign v5

Supported

90% interval 36.49–64.52

61
Ornith-1.5-9B
Ornith AI·Open Weight·262K

50.2

BenchAlign v5

Estimated

90% interval 40.37–60.11

62
Ministral 3 14B (Reasoning)
Mistral·Open Weight·128K

50.2

BenchAlign v5

Estimated

90% interval 38.69–61.72

63
Mixtral 8x22B Instruct v0.1
Mistral·Open Weight·64K

49.4

BenchAlign v5

Estimated

90% interval 37.90–60.92

64
GPT-OSS 120B
OpenAI·Open Weight·128K

49.4

BenchAlign v5

Supported

90% interval 36.14–62.60

65
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

48.1

BenchAlign v5

Supported

90% interval 30.81–65.33

66
Command A+
Cohere·Open Weight·128K

47.8

BenchAlign v5

Estimated

90% interval 36.24–59.27

67
Gemma 4 12B
Google·Open Weight·256K

47.5

BenchAlign v5

Estimated

90% interval 36.00–59.03

68
Mistral Small 4
Mistral·Open Weight·256K

46.5

BenchAlign v5

Estimated

90% interval 35.02–58.05

69
Nemotron-4 15B
NVIDIA·Open Weight·32K

45.9

BenchAlign v5

Estimated

90% interval 34.37–57.40

70
Mistral 8x7B
Mistral·Open Weight·32K

45.8

BenchAlign v5

Estimated

90% interval 34.28–57.30

71
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

45.7

BenchAlign v5

Estimated

90% interval 35.84–55.58

72
Nemotron Ultra 253B
NVIDIA·Open Weight·32K

45.2

BenchAlign v5

Estimated

90% interval 33.68–56.71

73
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

44.7

BenchAlign v5

Estimated

90% interval 33.16–56.19

74
DeepSeek V3
DeepSeek·Open Weight·128K

44.5

BenchAlign v5

Supported

90% interval 25.28–63.79

75
Ling 2.6 Flash
InclusionAI·Open Weight·262K

44.4

BenchAlign v5

Estimated

90% interval 32.85–55.88

76
Hy3 Preview
Tencent·Open Weight·256K

43.8

BenchAlign v5

Estimated

90% interval 33.88–53.62

77
Gemma 4 E4B
Google·Open Weight·128K

43.6

BenchAlign v5

Estimated

90% interval 32.05–55.08

78
Sarvam 105B
Sarvam·Open Weight·128K

43.5

BenchAlign v5

Estimated

90% interval 31.94–54.97

79
LFM2.5-2.6B
LiquidAI·Open Weight·128K

43.1

BenchAlign v5

Estimated

90% interval 31.58–54.61

80
DeepSeek R1 Distill Qwen 32B
DeepSeek·Open Weight·128K

43.1

BenchAlign v5

Estimated

90% interval 31.57–54.60

81
Granite 4.2 8B
IBM·Open Weight·128K

42.8

BenchAlign v5

Supported

90% interval 32.82–52.83

82
Gemma 4 E2B
Google·Open Weight·128K

42.5

BenchAlign v5

Estimated

90% interval 31.01–54.04

83
GPT-OSS 20B
OpenAI·Open Weight·128K

42.4

BenchAlign v5

Supported

90% interval 27.20–57.61

84
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

42

BenchAlign v5

Estimated

90% interval 30.46–53.49

85
Gemma 3 27B
Google·Open Weight·32K

41.4

BenchAlign v5

Supported

90% interval 17.07–65.77

86
Sarvam 30B
Sarvam·Open Weight·64K

41.3

BenchAlign v5

Estimated

90% interval 29.78–52.81

87
Ministral 3 8B (Reasoning)
Mistral·Open Weight·128K

41.1

BenchAlign v5

Estimated

90% interval 29.62–52.65

88
Exaone 4.0 32B
LG AI Research·Open Weight·128K

41

BenchAlign v5

Estimated

90% interval 29.52–52.55

89
Mistral 7B v0.3
Mistral·Open Weight·32K

40.7

BenchAlign v5

Estimated

90% interval 29.13–52.16

90
Qwen2.5-VL-32B
Alibaba·Open Weight·32K

40.6

BenchAlign v5

Estimated

90% interval 29.09–52.11

91
Llama 4 Behemoth
Meta·Open Weight·32K

40.6

BenchAlign v5

Estimated

90% interval 29.04–52.07

92
Ministral 3 3B (Reasoning)
Mistral·Open Weight·128K

40.3

BenchAlign v5

Estimated

90% interval 28.75–51.77

93
Mistral 8x7B v0.2
Mistral·Open Weight·32K

39.9

BenchAlign v5

Estimated

90% interval 28.36–51.39

94
Exaone 4.0 1.2B
LG AI Research·Open Weight·128K

39.7

BenchAlign v5

Estimated

90% interval 28.21–51.24

95
Llama 4 Scout
Meta·Open Weight·10M

39.7

BenchAlign v5

Supported

90% interval 20.78–58.62

96
Granite-4.0-350M
IBM·Open Weight·32K

39.2

BenchAlign v5

Estimated

90% interval 27.68–50.71

97
Granite-4.0-H-350M
IBM·Open Weight·32K

39.2

BenchAlign v5

Estimated

90% interval 27.68–50.71

98
Qwen2.5 Coder 32B Instruct
Alibaba·Open Weight·128K

34.4

BenchAlign v5

Supported

90% interval 17.63–51.09

99
Ministral 3 14B
Mistral·Open Weight·128K

34.2

BenchAlign v5

Supported

90% interval 23.10–45.19

100
ZAYA1-8B
Zyphra·Open Weight·131K

31.2

BenchAlign v5

Estimated

90% interval 21.35–41.09

101

26.6

BenchAlign v5

Estimated

90% interval 16.73–36.46

102
Llama 4 Maverick
Meta·Open Weight·1M

22.9

BenchAlign v5

Supported

90% interval 15.29–30.43

103
Phi-4
Microsoft·Open Weight·16K

22.7

BenchAlign v5

Supported

90% interval 0.00–45.80

104
Ministral 3 8B
Mistral·Open Weight·128K

20.5

BenchAlign v5

Supported

90% interval 16.80–24.18

105
Ministral 3 3B
Mistral·Open Weight·128K

18

BenchAlign v5

Supported

90% interval 14.11–21.97

106
MiniCPM5-1B
OpenBMB·Open Weight·131K

11.1

BenchAlign v5

Estimated

90% interval 1.26–21.00

Key Takeaways

The top model is Qwen3.8 Max by Alibaba with a BenchAlign v5 score of 78.7 and Supported evidence.

The best open-weight model is Qwen3.8 Max at position #1.

106 models are included in this ranking.

Score in Context

What these scores mean

Open-weight models are ranked by the same public BenchAlign v5 overall score as proprietary models. Supported and Estimated describe the evidence behind each position; they are not separate leaderboards.

Known limitations

Open weight is not the same as OSI-approved open source. The ranking does not score license restrictions, memory requirements, serving throughput, fine-tuning support, or the engineering cost of operating the model.

Best Open Source LLMs FAQ

What is the best open source LLM right now?

The live table above is the ranking owner and recomputes from the public BenchAlign v5 artifact. MiniMax M3 leads the July 14 snapshot with Supported evidence; use the table rather than a copied winner sentence when the data changes.

Are open source LLMs as good as GPT or Claude?

They can match proprietary models on individual tasks, but the July 14 unified overall ranking still shows a double-digit gap between the open-weight and proprietary leaders. Deployment control, privacy, and serving economics can still make an open model the better operational choice.

What is the best open source LLM for coding?

Use the open-weight rows on the live coding leaderboard. Coding order differs from overall order, and a single HumanEval or LiveCodeBench result should not be treated as a complete coding verdict.

Can I run these models locally?

The smaller open-weight rows run on a single consumer GPU with 4-bit quantization; frontier-size models need multi-GPU rigs or high-memory Apple Silicon. The local LLM guide breaks the rankings down by VRAM tier, and the Ollama guide includes pull commands and size estimates per model.

Which labs make the best open-weight models?

The current open-weight top tier comes almost entirely from Chinese labs — DeepSeek, Zhipu (GLM), Moonshot (Kimi), Alibaba (Qwen), and MiniMax — with Meta and Mistral behind on the live ranking. Family strengths differ by task, so check the per-category leaderboards and the Chinese model rankings rather than picking by lab reputation alone.

Last updated: September 2, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.