LLM Price vs Performance Chart
Compare the active public score against current API pricing. The shortlist favors the cheapest model that keeps at least 90% of the leader’s score; the efficiency frontier shows every undominated tradeoff.
Grok 4.5
Score 75.4 · $6.00/1M · 88% below the leader’s price
Claude Mythos 5
Score: 83.2 · $50.00/1M out
Availability restrictions apply
Ministral 3 3B
Score: 17.8 · $0.10/1M out
80 priced models match these filters. The blended view assumes three input tokens for every output token.
Lowest-cost score leaders (Overall)
Score/$ is shown as a narrow comparison aid, not a universal value verdict.
| # | Model | Score | Output $/1M | Score/$ |
|---|---|---|---|---|
| 1 | Ministral 3 3B Mistral | 17.8 | $0.10 | 178.0 |
| 2 | Ministral 3 8B Mistral | 20.3 | $0.15 | 135.3 |
| 3 | Ministral 3 14B Mistral | 33.9 | $0.20 | 169.5 |
| 4 | Step 3.5 Flash StepFun | 54.3 | $0.30 | 181.0 |
| 5 | DeepSeek V3.2 DeepSeek | 54.7 | $0.42 | 130.2 |
| 6 | DeepSeek V4 Pro 0813 DeepSeek | 61.2 | $0.87 | 70.3 |
| 7 | MiniMax M3 MiniMax | 68.7 | $1.20 | 57.3 |
| 8 | Grok 4.5 xAI | 75.4 | $6.00 | 12.6 |
| 9 | Gemini 3.6 Flash | 75.5 | $7.50 | 10.1 |
| 10 | Kimi K3 Moonshot AI | 80.5 | $15.00 | 5.4 |
Frequently Asked Questions
What is the LLM price-performance chart?
This chart plots each AI model by its benchmark score (vertical axis) against its API output price per million tokens (horizontal axis). Models in the upper-left quadrant offer the best value — high performance at low cost. The efficiency frontier line connects the best-value models at each price point.
What is the efficiency frontier?
The efficiency frontier (Pareto frontier) connects models where no other model offers both a higher score and a lower price. Models on this line represent the optimal price-performance tradeoff. If a model is below and to the right of the frontier, there exists a cheaper model with a better score.
Which LLM has the best price-to-performance ratio?
Grok 4.5 is the lowest-priced model that retains at least 90% of the current overall leader's score under these filters. We use this threshold instead of declaring the largest Score/$ ratio the universal winner, because normalized benchmark points are an index rather than units of completed work.
How are scores calculated?
The chart reads the active public ranking lane for each category and pairs that score with current catalog pricing. Overall, coding, and agentic views use the current BenchAlign methodology when enabled; other categories use the provisional composite. The 90% shortlist is a decision aid, not a claim that benchmark points convert directly into dollars.
Keep the value shortlist current
One weekly email when price changes or new benchmark evidence alter the models worth shortlisting.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.