Skip to main content

BenchLM recommendation

Best AI Models in 2026 — Overall Rankings

Data verified

Claude Mythos 5 leads all AI models on BenchLM's July 2026 rankings with a score of 83.9, ahead of Claude Fable 5 (83.7) and GPT-5.6 Sol (82). Each row uses the public BenchAlign v5 contract and shows its evidence status.

Last verified: July 20, 2026

BenchLM.ai now distinguishes provisional overall ranking from verified overall ranking. The provisional score is a normalized weighted average across 8 benchmark categories: agentic (22%), coding (20%), reasoning (17%), knowledge (12%), multimodal & grounded (12%), multilingual (7%), instruction following (5%), and math (5%), using non-generated benchmark coverage plus bounded external consensus calibration. The verified leaderboard is stricter and only counts sourced benchmark rows. Each score includes a confidence indicator (1-4 dots) based on how much sourced coverage supports it. Display-only benchmarks — including MMLU, OpenBookQA, HumanEval, FLTEval, BBH, LisanBench, and older AIME/HMMT variants — remain visible for context but do not affect ranking.

This is the public BenchAlign v5 overall lane. Scores combine the approved agentic and coding projections; evidence labels distinguish supported rows from estimates.

Bottom line: Claude Fable 5 leads overall, but GPT-5.4 and Claude Opus 4.6 are within striking distance — and significantly cheaper.

Operator review

What the ranking says today

I would start a broad evaluation with Claude Mythos 5, Claude Fable 5, and GPT-5.6 Sol. They are the first three rows under the same public scoring contract; the order is a shortlist, not a purchase order.

Operator receipt. I regenerated the active artifact against the reviewed 283-model registry on July 15, 2026. The public lane contains 200 eligible rows, led by Claude Mythos 5 at 83.9.

Honest limit. BenchAlign v5 projects a common score from uneven public evidence. An “estimated” badge and a wide interval mean the row is useful for discovery but weak evidence for a close call. Replay the finalists on your own tasks, latency budget, and safety constraints.

Reviewed by Glevd on July 15, 2026.

The active lane starts with Claude Mythos 5, followed by Claude Fable 5 and GPT-5.6 Sol. All rows use the same BenchAlign v5 projection contract. Evidence badges matter as much as small point gaps because the public source coverage is uneven.

Open-weight and proprietary models share this lane, but deployment terms, latency, and operating cost are not part of the BenchAlign score. Use the ranking to choose a test set, then compare the finalists under the constraints that will exist in production.

This ranking is based on the public BenchAlign v5 overall contract tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

Claude Fable 5 entered at #1 with the highest overall score on BenchLM.

GPT-5.4 holds a strong #2 across all categories.

Claude Opus 4.6 remains #3, the most consistent model across all 8 benchmark categories.

How to choose

Full Rankings (200 models)

1
Claude Mythos 5
Anthropic·Proprietary·1M+

83.9

BenchAlign v5

Supported

90% interval 79.92–87.94

2
Claude Fable 5
Anthropic·Proprietary·1M+

83.7

BenchAlign v5

Supported

90% interval 80.40–86.96

3
GPT-5.6 Sol
OpenAI·Proprietary·1M

82

BenchAlign v5

Supported

90% interval 78.09–85.84

4
Kimi K3
Moonshot AI·Pending·1.05M

81

BenchAlign v5

Supported

90% interval 77.87–84.04

5
Claude Opus 4.8
Anthropic·Proprietary·1M

78.3

BenchAlign v5

Supported

90% interval 75.32–81.36

6
Muse Spark 1.1
Meta·Proprietary·1M

77.4

BenchAlign v5

Supported

90% interval 73.62–81.26

7
Grok 4.5
xAI·Proprietary·500K

76.7

BenchAlign v5

Supported

90% interval 71.84–81.60

8
GPT-5.4
OpenAI·Proprietary·1.05M

74.2

BenchAlign v5

Supported

90% interval 71.24–77.25

9
GPT-5.5
OpenAI·Proprietary·1M

73.5

BenchAlign v5

Estimated

90% interval 64.31–82.72

10
Qwen3.7 Max
Alibaba·Proprietary·1M

72.8

BenchAlign v5

Supported

90% interval 66.59–79.08

11
GPT-5.6 Terra
OpenAI·Proprietary·1M

72.6

BenchAlign v5

Estimated

90% interval 62.70–82.44

12
Claude Opus 4.7
Anthropic·Proprietary·1M

71.9

BenchAlign v5

Supported

90% interval 60.15–83.72

13
Muse Spark
Meta·Proprietary·262K

71

BenchAlign v5

Supported

90% interval 62.62–79.46

14
MiMo-V2.5-Pro
Xiaomi·Proprietary·1M

70.2

BenchAlign v5

Supported

90% interval 62.93–77.44

15
MiniMax M3
MiniMax·Open Weight·1M

69.8

BenchAlign v5

Supported

90% interval 65.53–73.98

16
Claude Opus 4.6
Anthropic·Proprietary·1M

68.6

BenchAlign v5

Supported

90% interval 56.01–81.17

17
MiMo-V2-Pro
Xiaomi·Proprietary·1M

67.8

BenchAlign v5

Supported

90% interval 60.49–75.06

18
GLM-5.1
Z.AI·Open Weight·203K

67.7

BenchAlign v5

Supported

90% interval 58.01–77.46

19
Gemini 3 Pro
Google·Proprietary·2M

67.7

BenchAlign v5

Supported

90% interval 57.07–78.39

20
Inkling
Thinking Machines Lab·Open Weight·1M

67.5

BenchAlign v5

Supported

90% interval 61.11–73.97

21
Qwen3.7 Plus
Alibaba·Proprietary·1M

67.2

BenchAlign v5

Supported

90% interval 58.08–76.36

22
GPT-5.6 Luna
OpenAI·Proprietary·1M

67.2

BenchAlign v5

Estimated

90% interval 56.48–77.86

23
GPT-5.2 Pro
OpenAI·Proprietary·400K

67

BenchAlign v5

Supported

90% interval 56.80–77.22

24
GLM-5-Turbo
Z.AI·Proprietary·200K

66.9

BenchAlign v5

Supported

90% interval 57.48–76.31

25
GPT-5.4 nano
OpenAI·Proprietary·400K

66.8

BenchAlign v5

Supported

90% interval 56.69–76.88

26
GPT-5.3 Codex
OpenAI·Proprietary·400K

66.7

BenchAlign v5

Supported

90% interval 63.42–69.96

27
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

66.3

BenchAlign v5

Estimated

90% interval 56.40–76.14

28
GLM-5
Z.AI·Open Weight·200K

66.1

BenchAlign v5

Supported

90% interval 55.41–76.71

29
Claude Sonnet 5
Anthropic·Proprietary·1M

65.3

BenchAlign v5

Estimated

90% interval 50.50–80.15

30
Qwen3.6 Plus
Alibaba·Proprietary·1M

65.2

BenchAlign v5

Supported

90% interval 56.68–73.71

31
Grok 4.3
xAI·Proprietary·1M

65.1

BenchAlign v5

Supported

90% interval 55.65–74.56

32
Claude Sonnet 4.6
Anthropic·Proprietary·200K

65.1

BenchAlign v5

Supported

90% interval 53.11–77.04

33
Gemini 3.5 Flash
Google·Proprietary·1M

64.8

BenchAlign v5

Estimated

90% interval 53.79–75.71

34
Claude Opus 4.5
Anthropic·Proprietary·200K

64.2

BenchAlign v5

Supported

90% interval 51.25–77.18

35
Claude Opus 4.6 (Adaptive)
Anthropic·Proprietary·1M

64.2

BenchAlign v5

Estimated

90% interval 52.66–75.69

36
MiniMax M2.7
MiniMax·Open Weight·200K

64.1

BenchAlign v5

Supported

90% interval 57.83–70.40

37
GLM-5.2
Z.AI·Open Weight·1M

64

BenchAlign v5

Estimated

90% interval 48.35–79.58

38
GPT-5.5 Pro
OpenAI·Proprietary·1M

63.7

BenchAlign v5

Estimated

90% interval 49.00–78.38

39
GLM-5V-Turbo
Z.AI·Proprietary·200K

63.5

BenchAlign v5

Supported

90% interval 52.89–74.10

40
MiMo-V2-Omni
Xiaomi·Proprietary·262K

63.2

BenchAlign v5

Supported

90% interval 53.64–72.65

41
Gemini 3 Pro Deep Think
Google·Proprietary·2M

61.3

BenchAlign v5

Estimated

90% interval 49.79–72.82

42
GLM-4.7
Z.AI·Open Weight·200K

61.2

BenchAlign v5

Supported

90% interval 47.69–74.62

43
Gemma 4 31B
Google·Open Weight·256K

61.1

BenchAlign v5

Supported

90% interval 45.85–76.32

44
GPT-5.4 Pro
OpenAI·Proprietary·1.05M

60.9

BenchAlign v5

Estimated

90% interval 44.20–77.57

45
Qwen3.5-27B
Alibaba·Open Weight·262K

60.7

BenchAlign v5

Supported

90% interval 52.05–69.35

46
DeepSeek V4 Pro
DeepSeek·Open Weight·1M

60.7

BenchAlign v5

Supported

90% interval 42.29–79.04

47
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

60.6

BenchAlign v5

Supported

90% interval 50.16–70.96

48
Grok 4.1 Fast (Reasoning)
xAI·Proprietary·2M

60.5

BenchAlign v5

Supported

90% interval 47.82–73.21

49
Gemini 3 Flash
Google·Proprietary·1M

60.5

BenchAlign v5

Supported

90% interval 42.02–78.97

50
Grok 4
xAI·Proprietary·128K

60.4

BenchAlign v5

Supported

90% interval 51.10–69.75

51
Grok 4.1
xAI·Proprietary·1M

60

BenchAlign v5

Estimated

90% interval 48.45–71.48

52
GLM-5 (Reasoning)
Z.AI·Open Weight·200K

59.8

BenchAlign v5

Estimated

90% interval 48.26–71.29

53
Qwen 3.6 Max (preview)
Alibaba·Proprietary·256K

59.7

BenchAlign v5

Supported

90% interval 40.87–78.57

54
Kimi K2.5
Moonshot AI·Open Weight·256K

59.7

BenchAlign v5

Supported

90% interval 52.50–66.83

55
MiniMax M2.5
MiniMax·Proprietary·128K

59.5

BenchAlign v5

Supported

90% interval 52.22–66.83

56
Qwen3.5 397B (Reasoning)
Alibaba·Open Weight·128K

59.5

BenchAlign v5

Estimated

90% interval 47.98–71.01

57
Kimi K2.5 (Reasoning)
Moonshot AI·Proprietary·128K

59.4

BenchAlign v5

Estimated

90% interval 47.83–70.86

58
GPT-5.2-Codex
OpenAI·Proprietary·400K

59.1

BenchAlign v5

Supported

90% interval 55.18–63.01

59
GPT-5.2 Instant
OpenAI·Proprietary·128K

59

BenchAlign v5

Estimated

90% interval 47.49–70.52

60
GPT-5.3 Instant
OpenAI·Proprietary·400K

58.9

BenchAlign v5

Estimated

90% interval 47.38–70.41

61
DeepSeek V4 Flash
DeepSeek·Open Weight·1M

58.9

BenchAlign v5

Estimated

90% interval 47.36–70.39

62
MiMo-V2.5
Xiaomi·Proprietary·1M

58.6

BenchAlign v5

Estimated

90% interval 47.11–70.14

63
GPT-5 (high)
OpenAI·Proprietary·128K

58.6

BenchAlign v5

Estimated

90% interval 47.09–70.12

64
GPT-5.2
OpenAI·Proprietary·400K

58.4

BenchAlign v5

Estimated

90% interval 50.04–66.82

65
DeepSeek V3.2 (Thinking)
DeepSeek·Open Weight·128K

58.2

BenchAlign v5

Estimated

90% interval 46.64–69.66

66
Qwen3 235B 2507 (Reasoning)
Alibaba·Open Weight·128K

58

BenchAlign v5

Estimated

90% interval 46.50–69.53

67
Gemma 4 26B A4B
Google·Open Weight·256K

58

BenchAlign v5

Supported

90% interval 41.18–74.74

68
GLM-4.5
Z.AI·Proprietary·128K

57.6

BenchAlign v5

Estimated

90% interval 46.04–69.07

69
Claude Opus 4.5 Thinking
Anthropic·Proprietary·200K

57.4

BenchAlign v5

Estimated

90% interval 45.93–68.95

70
Gemini 2.5 Pro
Google·Proprietary·1M

57.3

BenchAlign v5

Supported

90% interval 40.11–74.39

71
Qwen3.5 397B
Alibaba·Open Weight·128K

57

BenchAlign v5

Estimated

90% interval 45.50–68.52

72
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

57

BenchAlign v5

Supported

90% interval 46.19–67.75

73
GPT-5.3-Codex-Spark
OpenAI·Proprietary·256K

56.9

BenchAlign v5

Estimated

90% interval 45.40–68.42

74
Kimi K2.6
Moonshot AI·Open Weight·256K

56.8

BenchAlign v5

Estimated

90% interval 46.92–66.66

75
GPT-5.4 mini
OpenAI·Proprietary·400K

56.8

BenchAlign v5

Estimated

90% interval 45.26–68.28

76
Grok 4 Fast (Reasoning)
xAI·Proprietary·2M

56.6

BenchAlign v5

Supported

90% interval 43.78–69.40

77
Claude Haiku 4.5
Anthropic·Proprietary·200K

56.6

BenchAlign v5

Estimated

90% interval 45.07–68.09

78
Qwen3 235B 2507
Alibaba·Open Weight·128K

56

BenchAlign v5

Estimated

90% interval 44.50–67.53

79
Trinity-Large-Preview
Arcee AI·Open Weight·512K

55.9

BenchAlign v5

Estimated

90% interval 44.43–67.46

80
Hy3
Tencent·Open Weight·256K

55.6

BenchAlign v5

Estimated

90% interval 44.13–67.16

81
DeepSeek V4 Pro (High)
DeepSeek·Open Weight·1M

55.5

BenchAlign v5

Estimated

90% interval 43.95–66.98

82
DeepSeek V3.2
DeepSeek·Open Weight·128K

55.4

BenchAlign v5

Supported

90% interval 38.94–71.86

83
Gemini 3.1 Pro
Google·Proprietary·1M

55.3

BenchAlign v5

Estimated

90% interval 39.46–71.13

84
GPT-5 (medium)
OpenAI·Proprietary·128K

55.2

BenchAlign v5

Supported

90% interval 51.75–58.55

85
GLM-4.6
Z.AI·Open Weight·200K

55.1

BenchAlign v5

Supported

90% interval 37.46–72.79

86
Step 3.5 Flash
StepFun·Open Weight·256K

55.1

BenchAlign v5

Supported

90% interval 42.05–68.16

87
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

55

BenchAlign v5

Estimated

90% interval 42.87–67.12

88
Grok 4.20
xAI·Proprietary·2M

54.7

BenchAlign v5

Estimated

90% interval 38.01–71.36

89
DeepSeek LLM 2.0
DeepSeek·Open Weight·128K

54.5

BenchAlign v5

Estimated

90% interval 43.01–66.04

90
GPT-5.1-Codex-Max
OpenAI·Proprietary·400K

54.5

BenchAlign v5

Estimated

90% interval 42.96–65.99

91
MiMo-V2-Flash
Xiaomi·Open Weight·256K

54.1

BenchAlign v5

Supported

90% interval 40.17–67.95

92
DeepSeek V4 Flash (High)
DeepSeek·Open Weight·1M

54

BenchAlign v5

Estimated

90% interval 42.44–65.47

93
Qwen3.6-27B
Alibaba·Open Weight·262K

53.8

BenchAlign v5

Estimated

90% interval 42.30–65.33

94
GPT-5.1
OpenAI·Proprietary·200K

53.7

BenchAlign v5

Estimated

90% interval 42.14–65.17

95
DeepSeek V3.1
DeepSeek·Open Weight·128K

53.6

BenchAlign v5

Supported

90% interval 35.22–72.05

96
Claude Sonnet 4.5
Anthropic·Proprietary·200K

53.6

BenchAlign v5

Estimated

90% interval 42.10–65.13

97
DeepSeek V3.1 (Reasoning)
DeepSeek·Open Weight·128K

53.4

BenchAlign v5

Supported

90% interval 34.72–72.14

98
Nemotron 3 Nano 30B
NVIDIA·Open Weight·32K

52.9

BenchAlign v5

Estimated

90% interval 41.43–64.46

99
GPT-5.1-Codex
OpenAI·Proprietary·400K

52.7

BenchAlign v5

Estimated

90% interval 41.21–64.23

100
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

52.3

BenchAlign v5

Supported

90% interval 40.28–64.36

101
Qwen2.5-72B
Alibaba·Open Weight·128K

52.2

BenchAlign v5

Estimated

90% interval 40.64–63.67

102
Llama 3.1 405B
Meta·Open Weight·128K

51.7

BenchAlign v5

Estimated

90% interval 40.20–63.22

103
DeepSeek-R1
DeepSeek·Open Weight·128K

51.7

BenchAlign v5

Supported

90% interval 34.16–69.17

104
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

51.5

BenchAlign v5

Estimated

90% interval 39.95–62.98

105
Mercury 2
Inception·Proprietary·128K

51.3

BenchAlign v5

Supported

90% interval 41.65–60.91

106
GLM-4.7-Flash
Z.AI·Open Weight·200K

51.3

BenchAlign v5

Supported

90% interval 38.20–64.30

107
Grok 4.1 Fast
xAI·Proprietary·1M

51.3

BenchAlign v5

Supported

90% interval 29.58–72.92

108
GPT-4.1
OpenAI·Proprietary·1M

51.1

BenchAlign v5

Supported

90% interval 31.76–70.46

109
Nemotron 3 Super 120B A12B
NVIDIA·Open Weight·256K

51

BenchAlign v5

Estimated

90% interval 39.46–62.48

110
Step 3.7 Flash
StepFun·Open Weight·256K

50.9

BenchAlign v5

Estimated

90% interval 39.35–62.38

111
Gemini 3.1 Flash-Lite
Google·Proprietary·1M

50.8

BenchAlign v5

Supported

90% interval 24.90–76.77

112
Llama 3 70B
Meta·Open Weight·128K

50.8

BenchAlign v5

Estimated

90% interval 39.31–62.34

113
Mistral Large 3
Mistral·Proprietary·128K

50.4

BenchAlign v5

Supported

90% interval 28.46–72.33

114
DeepSeek Coder 2.0
DeepSeek·Open Weight·128K

50.3

BenchAlign v5

Estimated

90% interval 38.77–61.79

115
Seed 1.6
ByteDance·Proprietary·256K

50.2

BenchAlign v5

Estimated

90% interval 38.72–61.74

116
GPT-OSS 120B
OpenAI·Open Weight·128K

50.1

BenchAlign v5

Supported

90% interval 37.76–62.41

117
Nemotron 3 Super 100B
NVIDIA·Open Weight·1M

50.1

BenchAlign v5

Estimated

90% interval 38.57–61.60

118
o4-mini (high)
OpenAI·Proprietary·200K

50

BenchAlign v5

Estimated

90% interval 38.47–61.50

119
DeepSeekMath V2
DeepSeek·Open Weight·128K

49.9

BenchAlign v5

Estimated

90% interval 38.37–61.40

120
Qwen2.5-1M
Alibaba·Open Weight·1M

49.9

BenchAlign v5

Estimated

90% interval 38.37–61.40

121
Seed-2.0-Lite
ByteDance·Proprietary·256K

49.8

BenchAlign v5

Estimated

90% interval 38.32–61.35

122
Ministral 3 14B (Reasoning)
Mistral·Open Weight·128K

49.3

BenchAlign v5

Estimated

90% interval 37.83–60.86

123
o1-preview
OpenAI·Proprietary·200K

49.1

BenchAlign v5

Supported

90% interval 30.24–67.97

124
Aion-2.0
Aion Labs·Proprietary·128K

48.7

BenchAlign v5

Estimated

90% interval 37.14–60.17

125
Mixtral 8x22B Instruct v0.1
Mistral·Open Weight·64K

48.5

BenchAlign v5

Estimated

90% interval 37.03–60.06

126
K-Exaone
LG AI Research·Proprietary·256K

48.5

BenchAlign v5

Estimated

90% interval 36.94–59.96

127
o3-pro
OpenAI·Proprietary·200K

48.3

BenchAlign v5

Supported

90% interval 43.06–53.60

128
Qwen3 Max
Alibaba·Proprietary·1M

48.2

BenchAlign v5

Estimated

90% interval 36.64–59.67

129
o1
OpenAI·Proprietary·200K

48.1

BenchAlign v5

Estimated

90% interval 36.59–59.62

130
Gemini 2.5 Flash
Google·Proprietary·1M

48.1

BenchAlign v5

Supported

90% interval 25.59–70.59

131
o3
OpenAI·Proprietary·200K

47.9

BenchAlign v5

Supported

90% interval 44.59–51.19

132
Claude 3.5 Sonnet
Anthropic·Proprietary·200K

47.7

BenchAlign v5

Estimated

90% interval 36.22–59.25

133
GLM-4.5-Air
Z.AI·Proprietary·128K

47.7

BenchAlign v5

Supported

90% interval 29.71–65.69

134
Qwen3.5 Flash
Alibaba·Proprietary·1M

47.7

BenchAlign v5

Supported

90% interval 24.83–70.47

135
Command A+
Cohere·Open Weight·128K

47.5

BenchAlign v5

Estimated

90% interval 35.99–59.02

136
o3-mini
OpenAI·Proprietary·200K

47.4

BenchAlign v5

Supported

90% interval 33.24–61.59

137
Gemma 4 12B
Google·Open Weight·256K

47.3

BenchAlign v5

Estimated

90% interval 35.77–58.80

138
Qwen3.5 Plus
Alibaba·Proprietary·1M

47.2

BenchAlign v5

Estimated

90% interval 33.40–61.01

139
GPT-5 nano
OpenAI·Proprietary·400K

46.4

BenchAlign v5

Estimated

90% interval 34.85–57.88

140
Mistral Small 4
Mistral·Open Weight·256K

46.2

BenchAlign v5

Estimated

90% interval 34.72–57.75

141
o1-pro
OpenAI·Proprietary·200K

45.9

BenchAlign v5

Estimated

90% interval 34.43–57.46

142
Claude 4.1 Opus
Anthropic·Proprietary·200K

45.9

BenchAlign v5

Supported

90% interval 42.91–48.84

143
Z-1
Z·Proprietary·128K

45.1

BenchAlign v5

Estimated

90% interval 33.62–56.65

144
Nemotron-4 15B
NVIDIA·Open Weight·32K

45.1

BenchAlign v5

Estimated

90% interval 33.57–56.60

145
Seed 1.6 Flash
ByteDance·Proprietary·256K

45.1

BenchAlign v5

Estimated

90% interval 33.57–56.60

146
Mistral 8x7B
Mistral·Open Weight·32K

45

BenchAlign v5

Estimated

90% interval 33.47–56.50

147
DeepSeek V3
DeepSeek·Open Weight·128K

45

BenchAlign v5

Supported

90% interval 26.51–63.43

148
Moonshot v1
Moonshot AI·Proprietary·128K

44.8

BenchAlign v5

Estimated

90% interval 33.28–56.31

149
Seed-2.0-Mini
ByteDance·Proprietary·256K

44.6

BenchAlign v5

Estimated

90% interval 33.08–56.11

150
Nemotron Ultra 253B
NVIDIA·Open Weight·32K

44.4

BenchAlign v5

Estimated

90% interval 32.89–55.92

151
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

44.2

BenchAlign v5

Estimated

90% interval 32.73–55.76

152
GPT-4.1 mini
OpenAI·Proprietary·1M

44.2

BenchAlign v5

Estimated

90% interval 32.67–55.70

153
GPT-5 mini
OpenAI·Proprietary·128K

43.9

BenchAlign v5

Supported

90% interval 40.89–46.97

154
Ling 2.6 Flash
InclusionAI·Open Weight·262K

43.9

BenchAlign v5

Estimated

90% interval 32.35–55.38

155
Gemma 4 E4B
Google·Open Weight·128K

43.2

BenchAlign v5

Estimated

90% interval 31.69–54.72

156
Mistral Medium 3
Mistral·Proprietary·128K

43.2

BenchAlign v5

Estimated

90% interval 31.69–54.71

157
Sarvam 105B
Sarvam·Open Weight·128K

43

BenchAlign v5

Estimated

90% interval 31.45–54.48

158
Claude 4 Sonnet
Anthropic·Proprietary·200K

42.8

BenchAlign v5

Supported

90% interval 39.82–45.75

159
GPT-OSS 20B
OpenAI·Open Weight·128K

42.7

BenchAlign v5

Supported

90% interval 28.16–57.33

160
DeepSeek R1 Distill Qwen 32B
DeepSeek·Open Weight·128K

42.6

BenchAlign v5

Estimated

90% interval 31.07–54.09

161
GPT-4.1 nano
OpenAI·Proprietary·1M

42.1

BenchAlign v5

Estimated

90% interval 30.54–53.57

162
Gemma 4 E2B
Google·Open Weight·128K

41.8

BenchAlign v5

Estimated

90% interval 30.30–53.33

163
Mistral Large 2
Mistral·Proprietary·128K

41.8

BenchAlign v5

Estimated

90% interval 30.26–53.28

164
Gemma 3 27B
Google·Open Weight·32K

41.6

BenchAlign v5

Supported

90% interval 17.74–65.40

165
GPT-4o
OpenAI·Proprietary·128K

41.5

BenchAlign v5

Supported

90% interval 22.74–60.24

166
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

41.4

BenchAlign v5

Estimated

90% interval 29.91–52.93

167
Solar Pro 2
Upstage·Proprietary·128K

41.2

BenchAlign v5

Estimated

90% interval 29.67–52.70

168
Claude 3 Opus
Anthropic·Proprietary·200K

41.1

BenchAlign v5

Supported

90% interval 23.95–58.30

169
Sarvam 30B
Sarvam·Open Weight·64K

40.7

BenchAlign v5

Estimated

90% interval 29.19–52.22

170
Exaone 4.0 32B
LG AI Research·Open Weight·128K

40.4

BenchAlign v5

Estimated

90% interval 28.92–51.95

171
Grok 3 [Beta]
xAI·Proprietary·128K

40.4

BenchAlign v5

Supported

90% interval 9.06–71.81

172
Ministral 3 8B (Reasoning)
Mistral·Open Weight·128K

40.4

BenchAlign v5

Estimated

90% interval 28.87–51.90

173
Mistral 7B v0.3
Mistral·Open Weight·32K

39.9

BenchAlign v5

Estimated

90% interval 28.39–51.41

174
Llama 4 Scout
Meta·Open Weight·10M

39.9

BenchAlign v5

Supported

90% interval 21.39–58.36

175
Qwen2.5-VL-32B
Alibaba·Open Weight·32K

39.9

BenchAlign v5

Estimated

90% interval 28.34–51.37

176
Llama 4 Behemoth
Meta·Open Weight·32K

39.8

BenchAlign v5

Estimated

90% interval 28.29–51.32

177
Ministral 3 3B (Reasoning)
Mistral·Open Weight·128K

39.5

BenchAlign v5

Estimated

90% interval 28.00–51.03

178
Mistral 8x7B v0.2
Mistral·Open Weight·32K

39.1

BenchAlign v5

Estimated

90% interval 27.62–50.65

179
Exaone 4.0 1.2B
LG AI Research·Open Weight·128K

39.1

BenchAlign v5

Estimated

90% interval 27.55–50.58

180
Grok Code Fast 1
xAI·Proprietary·256K

38.6

BenchAlign v5

Supported

90% interval 35.73–41.55

181
Granite-4.0-350M
IBM·Open Weight·32K

38.3

BenchAlign v5

Estimated

90% interval 26.80–49.83

182
Granite-4.0-H-350M
IBM·Open Weight·32K

38.3

BenchAlign v5

Estimated

90% interval 26.80–49.83

183
GPT-4o mini
OpenAI·Proprietary·128K

37.9

BenchAlign v5

Supported

90% interval 17.60–58.14

184
Claude 4.1 Opus Thinking
Anthropic·Proprietary·200K

36.6

BenchAlign v5

Supported

90% interval 16.51–56.59

185
Gemini 1.5 Pro
Google·Proprietary·2M

35.7

BenchAlign v5

Supported

90% interval 22.17–49.25

186
Qwen2.5 Coder 32B Instruct
Alibaba·Open Weight·128K

34.7

BenchAlign v5

Supported

90% interval 18.46–50.94

187
Ministral 3 14B
Mistral·Open Weight·128K

34.5

BenchAlign v5

Supported

90% interval 23.82–45.13

188
GPT-4 Turbo
OpenAI·Proprietary·128K

27.4

BenchAlign v5

Supported

90% interval 20.27–34.62

189
Kimi K2
Moonshot AI·Proprietary·128K

27.2

BenchAlign v5

Supported

90% interval 16.38–37.99

190
MiniMax M1 80k
MiniMax·Proprietary·80K

25.1

BenchAlign v5

Supported

90% interval 14.66–35.59

191
Llama 4 Maverick
Meta·Open Weight·1M

23.5

BenchAlign v5

Supported

90% interval 15.55–31.42

192
Phi-4
Microsoft·Open Weight·16K

22.7

BenchAlign v5

Supported

90% interval 0.00–45.80

193
Gemini 1.0 Pro
Google·Proprietary·32K

21.8

BenchAlign v5

Supported

90% interval 14.49–29.06

194
Claude 3 Haiku
Anthropic·Proprietary·200K

21.4

BenchAlign v5

Supported

90% interval 15.43–27.34

195
Ministral 3 8B
Mistral·Open Weight·128K

21

BenchAlign v5

Supported

90% interval 17.13–24.79

196
Nova Pro
Amazon·Proprietary·128K

20.3

BenchAlign v5

Supported

90% interval 17.08–23.56

197
LFM2-24B-A2B
LiquidAI·Proprietary·32K

18.9

BenchAlign v5

Supported

90% interval 15.87–21.87

198
Ministral 3 3B
Mistral·Open Weight·128K

18.3

BenchAlign v5

Supported

90% interval 14.40–22.15

199
LFM2.5-1.2B-Thinking
LiquidAI·Proprietary·32K

16.2

BenchAlign v5

Supported

90% interval 13.22–19.24

200
LFM2.5-1.2B-Instruct
LiquidAI·Proprietary·32K

15.5

BenchAlign v5

Supported

90% interval 12.61–18.43

Key Takeaways

The top model is Claude Mythos 5 by Anthropic with a BenchAlign v5 score of 83.9 and Supported evidence.

The best open-weight model is MiniMax M3 at position #15.

200 models are included in this ranking.

Score in Context

What these scores mean

The overall score is a weighted average across 8 benchmark categories. Agentic (22%), coding (20%), and reasoning (17%) carry the most weight. A 5-point gap in overall score is meaningful — it reflects consistent performance differences across multiple domains.

Known limitations

The overall score compresses 8 categories into one number. Two models with the same overall score can have very different strengths — one might lead coding while the other leads reasoning. Always check category scores for your specific use case.

Best AI Models Overall FAQ

What is the best AI model right now?

The model in the #1 row of the leaderboard above is the best AI model right now on BenchLM’s weighted rankings — the answer box at the top of this page names it with its current score. Rankings are recomputed on every data refresh across agentic, coding, reasoning, knowledge, multimodal, multilingual, instruction-following, and math benchmarks, so the leader can change as new models ship.

What is the smartest AI model?

"Smartest" depends on the yardstick. For raw reasoning and knowledge, check the reasoning and knowledge category columns on the leaderboard above; for the broadest definition of capability, the overall #1 is the best single answer. Reasoning-class models typically lead on hard-science benchmarks like GPQA Diamond and Humanity’s Last Exam, while non-reasoning models can still win on speed-sensitive everyday tasks.

What is the best LLM overall?

BenchLM ranks every tracked LLM by a weighted overall score — agentic work counts 22%, coding 20%, and reasoning 17%, with knowledge, multimodal, multilingual, instruction following, and math making up the rest. The current best LLM overall is the top row of this page’s leaderboard, with confidence dots showing how much sourced benchmark coverage supports its score.

How are these AI models ranked?

Each model’s overall score is a normalized weighted average across 8 benchmark categories, using non-generated benchmark coverage plus bounded external consensus calibration. A stricter verified leaderboard counts only sourced benchmark rows. Models without enough non-generated coverage are tracked but unranked, so a high position here reflects real, sourced evidence rather than a single headline benchmark.

Is the best AI model worth paying for?

Usually only for the slice of work where quality failures are expensive. The top-ranked models carry premium per-token prices, while models a few points lower often cost 2–5x less. Check the price-vs-performance view and the per-provider API pricing hubs to find the cheapest model that clears your quality bar, then route only high-stakes tasks to the leader.

Last updated: July 20, 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.