Skip to main content
BenchLM
Recommendation

Best Value LLM for Coding in 2026 — Cost-Adjusted Rankings

As of October 1, 2026, the top model in best value llm for coding on the BenchLM leaderboard is MiMo-V2.6-Flash with a score of 189.8.

Bottom line: value here is coding score per output dollar, so check each row's raw score before you shortlist it.

Ranking data as of

Full Rankings (91 models)

1
MiMo-V2.6-Flash
Xiaomi·Open Weight·1M

189.82

Score/$

Score: 53.2 · $0.28/1M

2
Mercury 2.5
Inception·Proprietary·260K

124.4

Score/$

Score: 18.7 · $0.15/1M

3
Ministral 3 3B
Mistral·Open Weight·128K

101.1

Score/$

Score: 10.1 · $0.1/1M

4
GPT-6 Luna
OpenAI·Proprietary·1.05M

100.36

Score/$

Score: 50.2 · $0.5/1M

5
MiMo-V2.6-Pro
Xiaomi·Open Weight·1M

78.15

Score/$

Score: 68 · $0.87/1M

6
Ministral 3 8B
Mistral·Open Weight·128K

77.6

Score/$

Score: 11.6 · $0.15/1M

7
DeepSeek V3.2
DeepSeek·Open Weight·128K

75.07

Score/$

Score: 31.5 · $0.42/1M

8
Qwen3.5 Flash
Alibaba·Proprietary·1M

66.83

Score/$

Score: 26.7 · $0.4/1M

9
Ministral 3 14B
Mistral·Open Weight·128K

65.3

Score/$

Score: 13.1 · $0.2/1M

10
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

52.82

Score/$

Score: 63.4 · $1.2/1M

11
DeepSeek V4.1 Flash
DeepSeek·Open Weight·1M

49.08

Score/$

Score: 58.9 · $1.2/1M

12
GPT-4.1 nano
OpenAI·Proprietary·1M

35.65

Score/$

Score: 14.3 · $0.4/1M

13
Mistral Small 4
Mistral·Open Weight·256K

35.58

Score/$

Score: 21.4 · $0.6/1M

14
MiniMax M3
MiniMax·Open Weight·1M

32.28

Score/$

Score: 38.7 · $1.2/1M

15
MiniMax M2.7
MiniMax·Open Weight·200K

30.05

Score/$

Score: 36.1 · $1.2/1M

16
MiniMax M2.5
MiniMax·Proprietary·128K

29.84

Score/$

Score: 35.8 · $1.2/1M

17
Step 3.7 Flash
StepFun·Open Weight·256K

28.49

Score/$

Score: 32.8 · $1.15/1M

18
Inkling-Small
Thinking Machines Lab·Open Weight·1M

27.81

Score/$

Score: 40 · $1.44/1M

19
GPT-5.4 nano
OpenAI·Proprietary·400K

24.93

Score/$

Score: 31.2 · $1.25/1M

20
Mercury 2
Inception·Proprietary·128K

24.55

Score/$

Score: 18.4 · $0.75/1M

21
Quasar 438B
Multiverse Computing·Proprietary·1M

24.2

Score/$

Score: 43.6 · $1.8/1M

22
Step 5 Preview
StepFun·Pending·1M

22.76

Score/$

Score: 61.5 · $2.7/1M

23
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

22.26

Score/$

Score: 20 · $0.9/1M

24
Celeris-1
Celeris·Proprietary·128K

19.09

Score/$

Score: 13.4 · $0.7/1M

25
GPT-4o mini
OpenAI·Proprietary·128K

18.55

Score/$

Score: 11.1 · $0.6/1M

26
Gemini 3.1 Flash-Lite
Google·Proprietary·1M

17.55

Score/$

Score: 26.3 · $1.5/1M

27
Gemini 3.8 Flash
Google·Proprietary·1M

16.98

Score/$

Score: 63.7 · $3.75/1M

28
DeepSeek V3
DeepSeek·Open Weight·128K

16.6

Score/$

Score: 18.3 · $1.1/1M

29
Gemini 3.7 Flash
Google·Proprietary·1M

15.82

Score/$

Score: 59.3 · $3.75/1M

30
Gemini 3.5 Flash-Lite
Google·Proprietary·1M

15.32

Score/$

Score: 38.3 · $2.5/1M

31
GPT-5 mini
OpenAI·Proprietary·128K

15.04

Score/$

Score: 30.1 · $2/1M

32
DeepSeek V3.2 (Thinking)
DeepSeek·Open Weight·128K

14.62

Score/$

Score: 32 · $2.19/1M

33
Gemini 3.6 Flash
Google·Proprietary·1M

14.31

Score/$

Score: 53.7 · $3.75/1M

34
Kimi K2.5
Moonshot AI·Open Weight·256K

13.4

Score/$

Score: 40.2 · $3/1M

35
Muse Spark 1.2
Meta·Proprietary·1M

12.96

Score/$

Score: 55.1 · $4.25/1M

36
Grok Code Fast 1
xAI·Proprietary·256K

12.89

Score/$

Score: 19.3 · $1.5/1M

37
GLM-5.2
Z.AI·Open Weight·1M

12.79

Score/$

Score: 56.3 · $4.4/1M

38
Grok Build 0.1
xAI·Proprietary·256K

12.64

Score/$

Score: 25.3 · $2/1M

39
GLM-5
Z.AI·Open Weight·200K

12.63

Score/$

Score: 40.4 · $3.2/1M

40
DeepSeek V4 Pro 0813
DeepSeek·Open Weight·1M

12.44

Score/$

Score: 49.3 · $3.96/1M

41
GPT-4.1 mini
OpenAI·Proprietary·1M

12.38

Score/$

Score: 19.8 · $1.6/1M

42
Gemini 3 Flash
Google·Proprietary·1M

12.11

Score/$

Score: 36.3 · $3/1M

43
Grok 4.3
xAI·Proprietary·1M

11.76

Score/$

Score: 29.4 · $2.5/1M

44
GLM-5.1
Z.AI·Open Weight·203K

11.54

Score/$

Score: 50.8 · $4.4/1M

45
Kimi K2.6
Moonshot AI·Open Weight·256K

11.52

Score/$

Score: 46.1 · $4/1M

46
Mistral Large 3
Mistral·Proprietary·128K

11.42

Score/$

Score: 17.1 · $1.5/1M

47
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

11.14

Score/$

Score: 44.6 · $4/1M

48
Grok 4.6
xAI·Proprietary·500K

10.24

Score/$

Score: 61.4 · $6/1M

49
Grok 4.5
xAI·Proprietary·500K

9.69

Score/$

Score: 58.1 · $6/1M

50
Grok 4.20
xAI·Proprietary·2M

9.23

Score/$

Score: 23.1 · $2.5/1M

51
GLM-5V-Turbo
Z.AI·Proprietary·200K

8.87

Score/$

Score: 35.5 · $4/1M

52
Claude Sonnet 5.5
Anthropic·Proprietary·1M

8.56

Score/$

Score: 85.6 · $10/1M

53
GPT-5.4 mini
OpenAI·Proprietary·400K

8.21

Score/$

Score: 37 · $4.5/1M

54
Inkling
Thinking Machines Lab·Open Weight·1M

7.78

Score/$

Score: 36.4 · $4.68/1M

55
Gemini 4 Argon
Google·Proprietary·

7.4

Score/$

Score: 74 · $10/1M

56
GPT-6 Sol
OpenAI·Proprietary·1.05M

6.19

Score/$

Score: 61.9 · $10/1M

57
Claude Sonnet 5
Anthropic·Proprietary·1M

5.98

Score/$

Score: 59.8 · $10/1M

58
GPT-6.1 Sol
OpenAI·Proprietary·1.05M

5.91

Score/$

Score: 59.1 · $10/1M

59
Gemini 3.5 Flash
Google·Proprietary·1M

5.82

Score/$

Score: 52.4 · $9/1M

60
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

5.36

Score/$

Score: 64.3 · $12/1M

61
Claude Opus 5.5
Anthropic·Proprietary·1M

4.17

Score/$

Score: 83.3 · $20/1M

62
Kimi K3
Moonshot AI·Pending·1.05M

4.09

Score/$

Score: 61.4 · $15/1M

63
GPT-5.3 Codex
OpenAI·Proprietary·400K

4.01

Score/$

Score: 56.1 · $14/1M

64
Claude Haiku 4.5
Anthropic·Proprietary·200K

3.92

Score/$

Score: 19.6 · $5/1M

65
GPT-5.1
OpenAI·Proprietary·200K

3.89

Score/$

Score: 38.9 · $10/1M

66
Gemini 3 Pro
Google·Proprietary·2M

3.79

Score/$

Score: 45.4 · $12/1M

67
GPT-5.1-Codex
OpenAI·Proprietary·400K

3.57

Score/$

Score: 35.7 · $10/1M

68
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

3.53

Score/$

Score: 70.5 · $20/1M

69
Gemini 3.1 Pro
Google·Proprietary·1M

3.53

Score/$

Score: 42.3 · $12/1M

70
Mistral Medium 3.5 128B
Mistral·Open Weight·256K

3.38

Score/$

Score: 25.4 · $7.5/1M

71
Gemini 1.5 Pro
Google·Proprietary·2M

3.36

Score/$

Score: 16.8 · $5/1M

72
GPT-5.2-Codex
OpenAI·Proprietary·400K

3.33

Score/$

Score: 46.6 · $14/1M

73
GPT-5.4
OpenAI·Proprietary·1.05M

3.31

Score/$

Score: 49.7 · $15/1M

74
Claude Sonnet 4.6
Anthropic·Proprietary·200K

3.14

Score/$

Score: 47.1 · $15/1M

75
Claude Opus 5
Anthropic·Proprietary·

2.88

Score/$

Score: 72.1 · $25/1M

76
GPT-5.2
OpenAI·Proprietary·400K

2.83

Score/$

Score: 39.7 · $14/1M

77
Claude Opus 4.8
Anthropic·Proprietary·1M

2.49

Score/$

Score: 62.3 · $25/1M

78
Gemini 2.5 Pro
Google·Proprietary·1M

2.43

Score/$

Score: 24.3 · $10/1M

79
Command A+
Cohere·Open Weight·128K

2.34

Score/$

Score: 23.4 · $10/1M

80
Claude Opus 4.7
Anthropic·Proprietary·1M

2.32

Score/$

Score: 58 · $25/1M

81
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

2.3

Score/$

Score: 57.6 · $25/1M

82
GPT-5.5
OpenAI·Proprietary·1M

2.09

Score/$

Score: 62.8 · $30/1M

83
Claude Opus 4.6
Anthropic·Proprietary·1M

1.98

Score/$

Score: 49.6 · $25/1M

84
Claude Opus 4.5
Anthropic·Proprietary·200K

1.71

Score/$

Score: 42.6 · $25/1M

85
Claude Fable 5.1
Anthropic·Proprietary·1M

1.6

Score/$

Score: 79.8 · $50/1M

86
GPT-6 Astra
OpenAI·Proprietary·1.05M

1.48

Score/$

Score: 74 · $50/1M

87
Claude Fable 5
Anthropic·Proprietary·1M+

1.46

Score/$

Score: 73 · $50/1M

88
o1
OpenAI·Proprietary·200K

0.51

Score/$

Score: 30.4 · $60/1M

89
GPT-4 Turbo
OpenAI·Proprietary·128K

0.44

Score/$

Score: 13.2 · $30/1M

90
o1-preview
OpenAI·Proprietary·200K

0.39

Score/$

Score: 23.4 · $60/1M

91
Claude 3 Opus
Anthropic·Proprietary·200K

0.19

Score/$

Score: 14 · $75/1M

How to choose

Key Takeaways

The best value model is MiMo-V2.6-Flash by Xiaomi with a provisional Score/$ ratio of 189.82 (score: 53.2, output: $0.28/1M tokens).

The best open-weight model is MiMo-V2.6-Flash at position #1.

91 models are included in this ranking.

Score in Context

What these scores mean

Value scores divide the weighted coding score by output token price (per 1M tokens). Higher means more capability per dollar. Models with no listed price are excluded.

Known limitations

Value rankings favor cheap models even if absolute performance is modest. A model scoring half as well at one-tenth the price wins on value — but may not meet your quality bar. Always check raw scores alongside value rankings.

About this ranking

Ranking data as of October 1, 2026

Raw benchmark scores only tell half the story. This ranking divides each model's weighted coding score by its output token price, surfacing models that deliver the most coding capability per dollar spent. A model scoring 70 at $1/1M tokens outranks one scoring 80 at $15/1M tokens here — because for the same budget you get far more coding work done. Use this alongside the standard coding leaderboard to find the sweet spot between performance and cost for your coding workflows.

Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

MiMo-V2.6-Flash leads this ranking with a score of 189.82, followed by Mercury 2.5 (124.4) and Ministral 3 3B (101.1). There is a significant gap between the leading models and the rest of the field.

The best open-weight option is MiMo-V2.6-Flash (ranked #1 with a score of 189.82). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional weighted averages across the scoring benchmarks in coding. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.

Questions

What is the best LLM for coding right now?

For raw coding capability, use the live coding leaderboard — it recomputes the order from calibrated coding evidence on every build, with evidence labels showing which positions are Supported versus Estimated. This page answers the follow-up question: which model delivers the most coding capability per dollar spent.

Is Claude or GPT better for coding?

Check the live coding leaderboard for the current order rather than a copied verdict — the top rows shift as new evidence lands. Evidence labels matter as much as rank: a Supported position slightly below an Estimated one is often the safer pick for production coding work.

What is the best open-source LLM for coding?

Use the open-weight rows on the coding leaderboard for capability order, and the open-source rankings for the full downloadable-weights picture. Open coding leaders trail the proprietary top tier on raw score but often win decisively on this page’s capability-per-dollar measure.

Which coding model gives the best value?

That is what this ranking measures: weighted coding score divided by output token price. A cheap model with a low coding score can rank high here, so read each row's raw coding score next to its price before you shortlist it.

Explore More

Last updated: October 1, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.