LLM Price vs Performance Chart
Claude Sonnet 5.5 (Anthropic) is the lowest-priced model that keeps at least 90% of the overall leader's score: 83.3 points at $10.00 per million output tokens, 80% below GPT-6 Astra. Ministral 3 3B sets the price floor at $0.10/1M.
Compare the active public score against current API pricing. The shortlist favors the cheapest model that keeps at least 90% of the leader’s score; the efficiency frontier shows every undominated tradeoff.
- 90% shortlist
- Claude Sonnet 5.5Score 83.3 · $10.00/1M · 80% below the leader’s price
- Highest score
- GPT-6 AstraScore 88.7 · $50.00/1M out · availability restrictions apply
- Price floor
- Ministral 3 3BScore 19.3 · $0.10/1M out
90 priced models match these filters. The blended view assumes three input tokens for every output token.
Lowest-cost score leaders (Overall)
Score/$ is shown as a narrow comparison aid, not a universal value verdict.
| # | Model | Score | Output $/1M | Score/$ |
|---|---|---|---|---|
| 1 | Ministral 3 3B Mistral | 19.3 | $0.10 | 193.0 |
| 2 | Qwen3.7 Flash Alibaba | 49.1 | $0.13 | 377.7 |
| 3 | DeepSeek V3.2 DeepSeek | 50.9 | $0.42 | 121.2 |
| 4 | MiMo-V2.6-Pro Xiaomi | 75.5 | $0.87 | 86.8 |
| 5 | Claude Sonnet 5.5 Anthropic | 83.3 | $10.00 | 8.3 |
| 6 | Claude Opus 5.5 Anthropic | 87.7 | $20.00 | 4.4 |
| 7 | GPT-6 Astra OpenAI | 88.7 | $50.00 | 1.8 |
Inside one provider the tradeoff is easiest to see: Gemini 3.8 Flash or Gemini 3.1 Pro is a price-for-capability decision within Google's own line.
Questions
What is the LLM price-performance chart?
This chart plots each AI model by its benchmark score (vertical axis) against its API output price per million tokens (horizontal axis). Models in the upper-left quadrant offer the best value — high performance at low cost. The efficiency frontier line connects the best-value models at each price point.
What is the efficiency frontier?
The efficiency frontier (Pareto frontier) connects models where no other model offers both a higher score and a lower price. Models on this line represent the optimal price-performance tradeoff. If a model is below and to the right of the frontier, there exists a cheaper model with a better score.
Which LLM has the best price-to-performance ratio?
Claude Sonnet 5.5 is the lowest-priced model that retains at least 90% of the current overall leader's score under these filters. We use this threshold instead of declaring the largest Score/$ ratio the universal winner, because normalized benchmark points are an index rather than units of completed work.
How are scores calculated?
The chart reads the active public ranking lane for each category and pairs that score with current catalog pricing. Overall, coding, and agentic views use the current BenchAlign methodology when enabled; other categories use the provisional composite. The 90% shortlist is a decision aid, not a claim that benchmark points convert directly into dollars.
Keep the value shortlist current
One weekly email when price changes or new benchmark evidence alter the models worth shortlisting.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.