Best Value LLM for Coding in 2026 — Cost-Adjusted Rankings
As of September 15, 2026, the top model in best value llm for coding on the BenchLM leaderboard is Mercury 2.5 with a score of 314.8.
Bottom line: Gemini 3.1 Flash-Lite dominates coding value at $0.40/1M output. DeepSeek Coder 2.0 offers the best absolute coding performance per dollar among serious coding models.
About this ranking
Last verified: September 15, 2026
Raw benchmark scores only tell half the story. This ranking divides each model's weighted coding score by its output token price, surfacing models that deliver the most coding capability per dollar spent. A model scoring 70 at $1/1M tokens outranks one scoring 80 at $15/1M tokens here — because for the same budget you get far more coding work done. Use this alongside the standard coding leaderboard to find the sweet spot between performance and cost for your coding workflows.
Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
Mercury 2.5 leads this ranking with a score of 314.8, followed by Laguna S 2.1 (230.05) and Ministral 3 3B (190). There is a significant gap between the leading models and the rest of the field.
The best open-weight option is Laguna S 2.1 (ranked #2 with a score of 230.05). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.
This ranking uses provisional weighted averages across the scoring benchmarks in coding. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
Gemini 3.1 Flash-Lite leads coding value — most coding capability per dollar spent.
DeepSeek Coder 2.0 best value among dedicated coding models with strong raw scores.
Gemini 2.5 Flash good coding value with broader model capabilities.
Full Rankings (89 models)
Key Takeaways
The best value model is Mercury 2.5 by Inception with a provisional Score/$ ratio of 314.8 (score: 47.2, output: $0.15/1M tokens).
The best open-weight model is Laguna S 2.1 at position #2.
89 models are included in this ranking.
Score in Context
What these scores mean
Value scores divide the weighted coding score by output token price (per 1M tokens). Higher means more capability per dollar. Models with no listed price are excluded.
Known limitations
Value rankings favor cheap models even if absolute performance is modest. A model scoring half as well at one-tenth the price wins on value — but may not meet your quality bar. Always check raw scores alongside value rankings.
Best Value LLM for Coding FAQ
What is the best LLM for coding right now?
For raw coding capability, use the live coding leaderboard — it recomputes the order from calibrated coding evidence on every build, with evidence labels showing which positions are Supported versus Estimated. This page answers the follow-up question: which model delivers the most coding capability per dollar spent.
Is Claude or GPT better for coding?
Check the live coding leaderboard for the current order rather than a copied verdict — the top rows shift as new evidence lands. Evidence labels matter as much as rank: a Supported position slightly below an Estimated one is often the safer pick for production coding work.
What is the best open-source LLM for coding?
Use the open-weight rows on the coding leaderboard for capability order, and the open-source rankings for the full downloadable-weights picture. Open coding leaders trail the proprietary top tier on raw score but often win decisively on this page’s capability-per-dollar measure.
Which coding model gives the best value?
That is what this ranking measures: weighted coding score divided by output token price. The current value leaders pair near-frontier coding scores with prices an order of magnitude below frontier models — see the table above for the live order and the exact price used for each row.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.