Skip to main content

BenchLM recommendation

Best Open Source LLMs in 2026

Data verified

As of July 20, 2026, the top model in best open source llms on the BenchLM leaderboard is MiniMax M3 with a score of 69.8.

Last verified: July 20, 2026

This page is the canonical open-weight ranking. It uses the same public BenchAlign v5 overall lane as the main leaderboard, then filters to downloadable model weights. The score compares measured capability; it does not decide whether a license is permissive, a model fits your hardware, or self-hosting beats an API on cost. Use the linked decision guide for those deployment questions.

Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Bottom line: use the live table above for capability order, then check evidence status, license terms, memory requirements, and serving cost before choosing a deployment.

MiniMax M3 leads this ranking with a score of 69.8, followed by GLM-5.1 (67.7) and Inkling (67.5). The top three are separated by just a few points — any of them would perform well for this use case.

All models in this ranking are open-weight, meaning they can be self-hosted for maximum control and cost efficiency.

This ranking is based on provisional overall weighted scores across BenchLM.ai's scoring formula tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

MiniMax M3 leads the public open-weight overall lane with Supported evidence.

GLM-5.1 follows on the live overall table with Supported evidence.

GLM-5 rounds out the current top three open-weight rows.

How to choose

Full Rankings (86 models)

1
MiniMax M3
MiniMax·Open Weight·1M

69.8

BenchAlign v5

Supported

90% interval 65.53–73.98

2
GLM-5.1
Z.AI·Open Weight·203K

67.7

BenchAlign v5

Supported

90% interval 58.01–77.46

3
Inkling
Thinking Machines Lab·Open Weight·1M

67.5

BenchAlign v5

Supported

90% interval 61.11–73.97

4
GLM-5
Z.AI·Open Weight·200K

66.1

BenchAlign v5

Supported

90% interval 55.41–76.71

5
MiniMax M2.7
MiniMax·Open Weight·200K

64.1

BenchAlign v5

Supported

90% interval 57.83–70.40

6
GLM-5.2
Z.AI·Open Weight·1M

64

BenchAlign v5

Estimated

90% interval 48.35–79.58

7
GLM-4.7
Z.AI·Open Weight·200K

61.2

BenchAlign v5

Supported

90% interval 47.69–74.62

8
Gemma 4 31B
Google·Open Weight·256K

61.1

BenchAlign v5

Supported

90% interval 45.85–76.32

9
Qwen3.5-27B
Alibaba·Open Weight·262K

60.7

BenchAlign v5

Supported

90% interval 52.05–69.35

10
DeepSeek V4 Pro
DeepSeek·Open Weight·1M

60.7

BenchAlign v5

Supported

90% interval 42.29–79.04

11
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

60.6

BenchAlign v5

Supported

90% interval 50.16–70.96

12
GLM-5 (Reasoning)
Z.AI·Open Weight·200K

59.8

BenchAlign v5

Estimated

90% interval 48.26–71.29

13
Kimi K2.5
Moonshot AI·Open Weight·256K

59.7

BenchAlign v5

Supported

90% interval 52.50–66.83

14
Qwen3.5 397B (Reasoning)
Alibaba·Open Weight·128K

59.5

BenchAlign v5

Estimated

90% interval 47.98–71.01

15
DeepSeek V4 Flash
DeepSeek·Open Weight·1M

58.9

BenchAlign v5

Estimated

90% interval 47.36–70.39

16
DeepSeek V3.2 (Thinking)
DeepSeek·Open Weight·128K

58.2

BenchAlign v5

Estimated

90% interval 46.64–69.66

17
Qwen3 235B 2507 (Reasoning)
Alibaba·Open Weight·128K

58

BenchAlign v5

Estimated

90% interval 46.50–69.53

18
Gemma 4 26B A4B
Google·Open Weight·256K

58

BenchAlign v5

Supported

90% interval 41.18–74.74

19
Qwen3.5 397B
Alibaba·Open Weight·128K

57

BenchAlign v5

Estimated

90% interval 45.50–68.52

20
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

57

BenchAlign v5

Supported

90% interval 46.19–67.75

21
Kimi K2.6
Moonshot AI·Open Weight·256K

56.8

BenchAlign v5

Estimated

90% interval 46.92–66.66

22
Qwen3 235B 2507
Alibaba·Open Weight·128K

56

BenchAlign v5

Estimated

90% interval 44.50–67.53

23
Trinity-Large-Preview
Arcee AI·Open Weight·512K

55.9

BenchAlign v5

Estimated

90% interval 44.43–67.46

24
Hy3
Tencent·Open Weight·256K

55.6

BenchAlign v5

Estimated

90% interval 44.13–67.16

25
DeepSeek V4 Pro (High)
DeepSeek·Open Weight·1M

55.5

BenchAlign v5

Estimated

90% interval 43.95–66.98

26
DeepSeek V3.2
DeepSeek·Open Weight·128K

55.4

BenchAlign v5

Supported

90% interval 38.94–71.86

27
GLM-4.6
Z.AI·Open Weight·200K

55.1

BenchAlign v5

Supported

90% interval 37.46–72.79

28
Step 3.5 Flash
StepFun·Open Weight·256K

55.1

BenchAlign v5

Supported

90% interval 42.05–68.16

29
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

55

BenchAlign v5

Estimated

90% interval 42.87–67.12

30
DeepSeek LLM 2.0
DeepSeek·Open Weight·128K

54.5

BenchAlign v5

Estimated

90% interval 43.01–66.04

31
MiMo-V2-Flash
Xiaomi·Open Weight·256K

54.1

BenchAlign v5

Supported

90% interval 40.17–67.95

32
DeepSeek V4 Flash (High)
DeepSeek·Open Weight·1M

54

BenchAlign v5

Estimated

90% interval 42.44–65.47

33
Qwen3.6-27B
Alibaba·Open Weight·262K

53.8

BenchAlign v5

Estimated

90% interval 42.30–65.33

34
DeepSeek V3.1
DeepSeek·Open Weight·128K

53.6

BenchAlign v5

Supported

90% interval 35.22–72.05

35
DeepSeek V3.1 (Reasoning)
DeepSeek·Open Weight·128K

53.4

BenchAlign v5

Supported

90% interval 34.72–72.14

36
Nemotron 3 Nano 30B
NVIDIA·Open Weight·32K

52.9

BenchAlign v5

Estimated

90% interval 41.43–64.46

37
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

52.3

BenchAlign v5

Supported

90% interval 40.28–64.36

38
Qwen2.5-72B
Alibaba·Open Weight·128K

52.2

BenchAlign v5

Estimated

90% interval 40.64–63.67

39
Llama 3.1 405B
Meta·Open Weight·128K

51.7

BenchAlign v5

Estimated

90% interval 40.20–63.22

40
DeepSeek-R1
DeepSeek·Open Weight·128K

51.7

BenchAlign v5

Supported

90% interval 34.16–69.17

41
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

51.5

BenchAlign v5

Estimated

90% interval 39.95–62.98

42
GLM-4.7-Flash
Z.AI·Open Weight·200K

51.3

BenchAlign v5

Supported

90% interval 38.20–64.30

43
Nemotron 3 Super 120B A12B
NVIDIA·Open Weight·256K

51

BenchAlign v5

Estimated

90% interval 39.46–62.48

44
Step 3.7 Flash
StepFun·Open Weight·256K

50.9

BenchAlign v5

Estimated

90% interval 39.35–62.38

45
Llama 3 70B
Meta·Open Weight·128K

50.8

BenchAlign v5

Estimated

90% interval 39.31–62.34

46
DeepSeek Coder 2.0
DeepSeek·Open Weight·128K

50.3

BenchAlign v5

Estimated

90% interval 38.77–61.79

47
GPT-OSS 120B
OpenAI·Open Weight·128K

50.1

BenchAlign v5

Supported

90% interval 37.76–62.41

48
Nemotron 3 Super 100B
NVIDIA·Open Weight·1M

50.1

BenchAlign v5

Estimated

90% interval 38.57–61.60

49
DeepSeekMath V2
DeepSeek·Open Weight·128K

49.9

BenchAlign v5

Estimated

90% interval 38.37–61.40

50
Qwen2.5-1M
Alibaba·Open Weight·1M

49.9

BenchAlign v5

Estimated

90% interval 38.37–61.40

51
Ministral 3 14B (Reasoning)
Mistral·Open Weight·128K

49.3

BenchAlign v5

Estimated

90% interval 37.83–60.86

52
Mixtral 8x22B Instruct v0.1
Mistral·Open Weight·64K

48.5

BenchAlign v5

Estimated

90% interval 37.03–60.06

53
Command A+
Cohere·Open Weight·128K

47.5

BenchAlign v5

Estimated

90% interval 35.99–59.02

54
Gemma 4 12B
Google·Open Weight·256K

47.3

BenchAlign v5

Estimated

90% interval 35.77–58.80

55
Mistral Small 4
Mistral·Open Weight·256K

46.2

BenchAlign v5

Estimated

90% interval 34.72–57.75

56
Nemotron-4 15B
NVIDIA·Open Weight·32K

45.1

BenchAlign v5

Estimated

90% interval 33.57–56.60

57
Mistral 8x7B
Mistral·Open Weight·32K

45

BenchAlign v5

Estimated

90% interval 33.47–56.50

58
DeepSeek V3
DeepSeek·Open Weight·128K

45

BenchAlign v5

Supported

90% interval 26.51–63.43

59
Nemotron Ultra 253B
NVIDIA·Open Weight·32K

44.4

BenchAlign v5

Estimated

90% interval 32.89–55.92

60
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

44.2

BenchAlign v5

Estimated

90% interval 32.73–55.76

61
Ling 2.6 Flash
InclusionAI·Open Weight·262K

43.9

BenchAlign v5

Estimated

90% interval 32.35–55.38

62
Gemma 4 E4B
Google·Open Weight·128K

43.2

BenchAlign v5

Estimated

90% interval 31.69–54.72

63
Sarvam 105B
Sarvam·Open Weight·128K

43

BenchAlign v5

Estimated

90% interval 31.45–54.48

64
GPT-OSS 20B
OpenAI·Open Weight·128K

42.7

BenchAlign v5

Supported

90% interval 28.16–57.33

65
DeepSeek R1 Distill Qwen 32B
DeepSeek·Open Weight·128K

42.6

BenchAlign v5

Estimated

90% interval 31.07–54.09

66
Gemma 4 E2B
Google·Open Weight·128K

41.8

BenchAlign v5

Estimated

90% interval 30.30–53.33

67
Gemma 3 27B
Google·Open Weight·32K

41.6

BenchAlign v5

Supported

90% interval 17.74–65.40

68
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

41.4

BenchAlign v5

Estimated

90% interval 29.91–52.93

69
Sarvam 30B
Sarvam·Open Weight·64K

40.7

BenchAlign v5

Estimated

90% interval 29.19–52.22

70
Exaone 4.0 32B
LG AI Research·Open Weight·128K

40.4

BenchAlign v5

Estimated

90% interval 28.92–51.95

71
Ministral 3 8B (Reasoning)
Mistral·Open Weight·128K

40.4

BenchAlign v5

Estimated

90% interval 28.87–51.90

72
Mistral 7B v0.3
Mistral·Open Weight·32K

39.9

BenchAlign v5

Estimated

90% interval 28.39–51.41

73
Llama 4 Scout
Meta·Open Weight·10M

39.9

BenchAlign v5

Supported

90% interval 21.39–58.36

74
Qwen2.5-VL-32B
Alibaba·Open Weight·32K

39.9

BenchAlign v5

Estimated

90% interval 28.34–51.37

75
Llama 4 Behemoth
Meta·Open Weight·32K

39.8

BenchAlign v5

Estimated

90% interval 28.29–51.32

76
Ministral 3 3B (Reasoning)
Mistral·Open Weight·128K

39.5

BenchAlign v5

Estimated

90% interval 28.00–51.03

77
Mistral 8x7B v0.2
Mistral·Open Weight·32K

39.1

BenchAlign v5

Estimated

90% interval 27.62–50.65

78
Exaone 4.0 1.2B
LG AI Research·Open Weight·128K

39.1

BenchAlign v5

Estimated

90% interval 27.55–50.58

79
Granite-4.0-350M
IBM·Open Weight·32K

38.3

BenchAlign v5

Estimated

90% interval 26.80–49.83

80
Granite-4.0-H-350M
IBM·Open Weight·32K

38.3

BenchAlign v5

Estimated

90% interval 26.80–49.83

81
Qwen2.5 Coder 32B Instruct
Alibaba·Open Weight·128K

34.7

BenchAlign v5

Supported

90% interval 18.46–50.94

82
Ministral 3 14B
Mistral·Open Weight·128K

34.5

BenchAlign v5

Supported

90% interval 23.82–45.13

83
Llama 4 Maverick
Meta·Open Weight·1M

23.5

BenchAlign v5

Supported

90% interval 15.55–31.42

84
Phi-4
Microsoft·Open Weight·16K

22.7

BenchAlign v5

Supported

90% interval 0.00–45.80

85
Ministral 3 8B
Mistral·Open Weight·128K

21

BenchAlign v5

Supported

90% interval 17.13–24.79

86
Ministral 3 3B
Mistral·Open Weight·128K

18.3

BenchAlign v5

Supported

90% interval 14.40–22.15

Key Takeaways

The top model is MiniMax M3 by MiniMax with a BenchAlign v5 score of 69.8 and Supported evidence.

The best open-weight model is MiniMax M3 at position #1.

86 models are included in this ranking.

Score in Context

What these scores mean

Open-weight models are ranked by the same public BenchAlign v5 overall score as proprietary models. Supported and Estimated describe the evidence behind each position; they are not separate leaderboards.

Known limitations

Open weight is not the same as OSI-approved open source. The ranking does not score license restrictions, memory requirements, serving throughput, fine-tuning support, or the engineering cost of operating the model.

Best Open Source LLMs FAQ

What is the best open source LLM right now?

The live table above is the ranking owner and recomputes from the public BenchAlign v5 artifact. MiniMax M3 leads the July 14 snapshot with Supported evidence; use the table rather than a copied winner sentence when the data changes.

Are open source LLMs as good as GPT or Claude?

They can match proprietary models on individual tasks, but the July 14 unified overall ranking still shows a double-digit gap between the open-weight and proprietary leaders. Deployment control, privacy, and serving economics can still make an open model the better operational choice.

What is the best open source LLM for coding?

Use the open-weight rows on the live coding leaderboard. Coding order differs from overall order, and a single HumanEval or LiveCodeBench result should not be treated as a complete coding verdict.

Can I run these models locally?

The smaller open-weight rows run on a single consumer GPU with 4-bit quantization; frontier-size models need multi-GPU rigs or high-memory Apple Silicon. The local LLM guide breaks the rankings down by VRAM tier, and the Ollama guide includes pull commands and size estimates per model.

Last updated: July 20, 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.