Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
BenchLM recommendation

Best Value LLM for Coding in 2026 — Cost-Adjusted Rankings

Data verified

As of September 15, 2026, the top model in best value llm for coding on the BenchLM leaderboard is Mercury 2.5 with a score of 314.8.

Bottom line: Gemini 3.1 Flash-Lite dominates coding value at $0.40/1M output. DeepSeek Coder 2.0 offers the best absolute coding performance per dollar among serious coding models.

About this ranking

Last verified: September 15, 2026

Raw benchmark scores only tell half the story. This ranking divides each model's weighted coding score by its output token price, surfacing models that deliver the most coding capability per dollar spent. A model scoring 70 at $1/1M tokens outranks one scoring 80 at $15/1M tokens here — because for the same budget you get far more coding work done. Use this alongside the standard coding leaderboard to find the sweet spot between performance and cost for your coding workflows.

Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Mercury 2.5 leads this ranking with a score of 314.8, followed by Laguna S 2.1 (230.05) and Ministral 3 3B (190). There is a significant gap between the leading models and the rest of the field.

The best open-weight option is Laguna S 2.1 (ranked #2 with a score of 230.05). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional weighted averages across the scoring benchmarks in coding. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

Gemini 3.1 Flash-Lite leads coding value — most coding capability per dollar spent.

DeepSeek Coder 2.0 best value among dedicated coding models with strong raw scores.

Gemini 2.5 Flash good coding value with broader model capabilities.

How to choose

Full Rankings (89 models)

1
Mercury 2.5
Inception·Proprietary·260K

314.8

Score/$

Score: 47.2 · $0.15/1M

2
Laguna S 2.1
Poolside·Open Weight·1M

230.05

Score/$

Score: 46 · $0.2/1M

3
Ministral 3 3B
Mistral·Open Weight·128K

190

Score/$

Score: 19 · $0.1/1M

4
Ministral 3 14B
Mistral·Open Weight·128K

151.2

Score/$

Score: 30.2 · $0.2/1M

5
Ministral 3 8B
Mistral·Open Weight·128K

140.87

Score/$

Score: 21.1 · $0.15/1M

6
DeepSeek V3.2
DeepSeek·Open Weight·128K

118.98

Score/$

Score: 50 · $0.42/1M

7
Qwen3.5 Flash
Alibaba·Proprietary·1M

116.95

Score/$

Score: 46.8 · $0.4/1M

8
GPT-4.1 nano
OpenAI·Proprietary·1M

77.65

Score/$

Score: 31.1 · $0.4/1M

9
Mistral Small 4
Mistral·Open Weight·256K

68.33

Score/$

Score: 41 · $0.6/1M

10
DeepSeek V4 Pro 0813
DeepSeek·Proprietary·1M

60.07

Score/$

Score: 52.3 · $0.87/1M

11
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

55.65

Score/$

Score: 66.8 · $1.2/1M

12
GPT-4o mini
OpenAI·Proprietary·128K

52.83

Score/$

Score: 31.7 · $0.6/1M

13
Step 3.7 Flash
StepFun·Open Weight·256K

41.09

Score/$

Score: 47.3 · $1.15/1M

14
MiniMax M3
MiniMax·Open Weight·1M

40.56

Score/$

Score: 48.7 · $1.2/1M

15
MiniMax M2.7
MiniMax·Open Weight·200K

40.48

Score/$

Score: 48.6 · $1.2/1M

16
MiniMax M2.5
MiniMax·Proprietary·128K

40.39

Score/$

Score: 48.5 · $1.2/1M

17
DeepSeek V3
DeepSeek·Open Weight·128K

35.87

Score/$

Score: 39.5 · $1.1/1M

18
Mercury 2
Inception·Proprietary·128K

32.8

Score/$

Score: 24.6 · $0.75/1M

19
Inkling-Small
Thinking Machines Lab·Open Weight·1M

30.42

Score/$

Score: 43.8 · $1.44/1M

20
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

30.41

Score/$

Score: 27.4 · $0.9/1M

21
GPT-5.4 nano
OpenAI·Proprietary·400K

29.77

Score/$

Score: 37.2 · $1.25/1M

22
Gemini 3.1 Flash-Lite
Google·Proprietary·1M

25.45

Score/$

Score: 38.2 · $1.5/1M

23
GPT-4.1 mini
OpenAI·Proprietary·1M

22.26

Score/$

Score: 35.6 · $1.6/1M

24
DeepSeek V3.2 (Thinking)
DeepSeek·Open Weight·128K

21.84

Score/$

Score: 47.8 · $2.19/1M

25
Composer 2.5
Cursor·Proprietary·200K

20.21

Score/$

Score: 50.5 · $2.5/1M

26
GPT-5 mini
OpenAI·Proprietary·128K

20.14

Score/$

Score: 40.3 · $2/1M

27
Grok Code Fast 1
xAI·Proprietary·256K

19.97

Score/$

Score: 30 · $1.5/1M

28
Composer 2
Cursor·Proprietary·200K

19.61

Score/$

Score: 49 · $2.5/1M

29
Grok Build 0.1
xAI·Proprietary·256K

19.55

Score/$

Score: 39.1 · $2/1M

30
Gemini 3.8 Flash
Google·Proprietary·1M

17.74

Score/$

Score: 66.5 · $3.75/1M

31
GLM-5
Z.AI·Open Weight·200K

17.52

Score/$

Score: 56.1 · $3.2/1M

32
Gemini 3.5 Flash-Lite
Google·Proprietary·1M

17.33

Score/$

Score: 43.3 · $2.5/1M

33
Kimi K2.5 (Reasoning)
Moonshot AI·Proprietary·128K

16.95

Score/$

Score: 50.8 · $3/1M

34
Mistral Large 3
Mistral·Proprietary·128K

16.95

Score/$

Score: 25.4 · $1.5/1M

35
Kimi K2.5
Moonshot AI·Open Weight·256K

16.85

Score/$

Score: 50.6 · $3/1M

36
Gemini 3.7 Flash
Google·Proprietary·1M

16.72

Score/$

Score: 62.7 · $3.75/1M

37
Muse Spark 1.2
Meta·Proprietary·1M

14.13

Score/$

Score: 60.1 · $4.25/1M

38
Qwen3.5 397B
Alibaba·Open Weight·128K

13.99

Score/$

Score: 50.4 · $3.6/1M

39
GLM-5.2
Z.AI·Open Weight·1M

13.85

Score/$

Score: 60.9 · $4.4/1M

40
Gemini 3 Flash
Google·Proprietary·1M

13.81

Score/$

Score: 41.4 · $3/1M

41
Grok 4.3
xAI·Proprietary·1M

13.48

Score/$

Score: 33.7 · $2.5/1M

42
GLM-5V-Turbo
Z.AI·Proprietary·200K

13.21

Score/$

Score: 52.8 · $4/1M

43
GLM-5.1
Z.AI·Open Weight·203K

12.8

Score/$

Score: 56.3 · $4.4/1M

44
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

12.73

Score/$

Score: 50.9 · $4/1M

45
Kimi K2.6
Moonshot AI·Open Weight·256K

12.7

Score/$

Score: 50.8 · $4/1M

46
Grok 4.6
xAI·Proprietary·500K

11.04

Score/$

Score: 66.2 · $6/1M

47
Grok 4.5
xAI·Proprietary·500K

10.35

Score/$

Score: 62.1 · $6/1M

48
o3-mini
OpenAI·Proprietary·200K

10.29

Score/$

Score: 45.3 · $4.4/1M

49
GPT-5.4 mini
OpenAI·Proprietary·400K

9.5

Score/$

Score: 42.8 · $4.5/1M

50
Inkling
Thinking Machines Lab·Open Weight·1M

8.79

Score/$

Score: 41.1 · $4.68/1M

51
Gemini 3.6 Flash
Google·Proprietary·1M

7.81

Score/$

Score: 58.5 · $7.5/1M

52
Gemini 1.5 Pro
Google·Proprietary·2M

6.84

Score/$

Score: 34.2 · $5/1M

53
Claude Sonnet 5
Anthropic·Proprietary·1M

6.4

Score/$

Score: 64 · $10/1M

54
Gemini 3.5 Flash
Google·Proprietary·1M

6.38

Score/$

Score: 57.4 · $9/1M

55
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

5.59

Score/$

Score: 67 · $12/1M

56
Claude Haiku 4.5
Anthropic·Proprietary·200K

5.48

Score/$

Score: 27.4 · $5/1M

57
Mistral Medium 3.5 128B
Mistral·Open Weight·256K

5.22

Score/$

Score: 39.2 · $7.5/1M

58
GPT-4.1
OpenAI·Proprietary·1M

5.17

Score/$

Score: 41.3 · $8/1M

59
Gemini 3 Pro
Google·Proprietary·2M

4.99

Score/$

Score: 59.9 · $12/1M

60
Grok 4.20
xAI·Proprietary·2M

4.73

Score/$

Score: 28.4 · $6/1M

61
Kimi K3
Moonshot AI·Pending·1.05M

4.52

Score/$

Score: 67.8 · $15/1M

62
GPT-5.1
OpenAI·Proprietary·200K

4.5

Score/$

Score: 45 · $10/1M

63
GPT-5.1-Codex
OpenAI·Proprietary·400K

4.49

Score/$

Score: 44.9 · $10/1M

64
GPT-5.3 Codex
OpenAI·Proprietary·400K

4.43

Score/$

Score: 62 · $14/1M

65
Command A+
Cohere·Open Weight·128K

4.36

Score/$

Score: 43.6 · $10/1M

66
Gemini 3.1 Pro
Google·Proprietary·1M

3.85

Score/$

Score: 46.2 · $12/1M

67
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

3.72

Score/$

Score: 74.4 · $20/1M

68
GPT-5.2-Codex
OpenAI·Proprietary·400K

3.68

Score/$

Score: 51.6 · $14/1M

69
GPT-5.4
OpenAI·Proprietary·1.05M

3.59

Score/$

Score: 53.9 · $15/1M

70
Claude Sonnet 4.6
Anthropic·Proprietary·200K

3.49

Score/$

Score: 52.3 · $15/1M

71
Claude Sonnet 4.5
Anthropic·Proprietary·200K

3.38

Score/$

Score: 50.7 · $15/1M

72
GPT-5.2
OpenAI·Proprietary·400K

3.32

Score/$

Score: 46.4 · $14/1M

73
Gemini 2.5 Pro
Google·Proprietary·1M

3.18

Score/$

Score: 31.8 · $10/1M

74
Claude Opus 5
Anthropic·Proprietary·

3.03

Score/$

Score: 75.7 · $25/1M

75
Claude 4 Sonnet
Anthropic·Proprietary·200K

2.67

Score/$

Score: 40.1 · $15/1M

76
Claude Opus 4.8
Anthropic·Proprietary·1M

2.66

Score/$

Score: 66.4 · $25/1M

77
Claude Opus 4.7
Anthropic·Proprietary·1M

2.51

Score/$

Score: 62.8 · $25/1M

78
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

2.31

Score/$

Score: 57.7 · $25/1M

79
Claude Opus 4.5
Anthropic·Proprietary·200K

2.27

Score/$

Score: 56.7 · $25/1M

80
Claude Opus 4.6
Anthropic·Proprietary·1M

2.26

Score/$

Score: 56.5 · $25/1M

81
GPT-5.5
OpenAI·Proprietary·1M

2.25

Score/$

Score: 67.6 · $30/1M

82
Claude Fable 5.1
Anthropic·Proprietary·1M

1.68

Score/$

Score: 83.9 · $50/1M

83
Claude Fable 5
Anthropic·Proprietary·1M+

1.54

Score/$

Score: 77 · $50/1M

84
GPT-6 Astra
OpenAI·Proprietary·1.05M

1.49

Score/$

Score: 74.5 · $50/1M

85
GPT-4 Turbo
OpenAI·Proprietary·128K

0.93

Score/$

Score: 27.8 · $30/1M

86
o1
OpenAI·Proprietary·200K

0.74

Score/$

Score: 44.4 · $60/1M

87
o1-preview
OpenAI·Proprietary·200K

0.74

Score/$

Score: 44.1 · $60/1M

88
Claude 4.1 Opus
Anthropic·Proprietary·200K

0.57

Score/$

Score: 42.9 · $75/1M

89
Claude 3 Opus
Anthropic·Proprietary·200K

0.45

Score/$

Score: 34 · $75/1M

Key Takeaways

The best value model is Mercury 2.5 by Inception with a provisional Score/$ ratio of 314.8 (score: 47.2, output: $0.15/1M tokens).

The best open-weight model is Laguna S 2.1 at position #2.

89 models are included in this ranking.

Score in Context

What these scores mean

Value scores divide the weighted coding score by output token price (per 1M tokens). Higher means more capability per dollar. Models with no listed price are excluded.

Known limitations

Value rankings favor cheap models even if absolute performance is modest. A model scoring half as well at one-tenth the price wins on value — but may not meet your quality bar. Always check raw scores alongside value rankings.

Best Value LLM for Coding FAQ

What is the best LLM for coding right now?

For raw coding capability, use the live coding leaderboard — it recomputes the order from calibrated coding evidence on every build, with evidence labels showing which positions are Supported versus Estimated. This page answers the follow-up question: which model delivers the most coding capability per dollar spent.

Is Claude or GPT better for coding?

Check the live coding leaderboard for the current order rather than a copied verdict — the top rows shift as new evidence lands. Evidence labels matter as much as rank: a Supported position slightly below an Estimated one is often the safer pick for production coding work.

What is the best open-source LLM for coding?

Use the open-weight rows on the coding leaderboard for capability order, and the open-source rankings for the full downloadable-weights picture. Open coding leaders trail the proprietary top tier on raw score but often win decisively on this page’s capability-per-dollar measure.

Which coding model gives the best value?

That is what this ranking measures: weighted coding score divided by output token price. The current value leaders pair near-frontier coding scores with prices an order of magnitude below frontier models — see the table above for the live order and the exact price used for each row.

Last updated: September 15, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.