Skip to main content
BenchLM

LLM Price vs Performance Chart

Claude Sonnet 5.5 (Anthropic) is the lowest-priced model that keeps at least 90% of the overall leader's score: 83.3 points at $10.00 per million output tokens, 80% below GPT-6 Astra. Ministral 3 3B sets the price floor at $0.10/1M.

Pricing verified

Compare the active public score against current API pricing. The shortlist favors the cheapest model that keeps at least 90% of the leader’s score; the efficiency frontier shows every undominated tradeoff.

90% shortlist
Claude Sonnet 5.5
Score 83.3 · $10.00/1M · 80% below the leader’s price
Highest score
GPT-6 Astra
Score 88.7 · $50.00/1M out · availability restrictions apply
Price floor
Ministral 3 3B
Score 19.3 · $0.10/1M out
Score Axis
Cost Basis
Source Type
Price Range
Model Families

90 priced models match these filters. The blended view assumes three input tokens for every output token.

Efficiency Frontier

Lowest-cost score leaders (Overall)

Score/$ is shown as a narrow comparison aid, not a universal value verdict.

RankModelScore / cost / value
1
Ministral 3 3BMistral · frontier tradeoff
19.3 score$0.10 / 1M out193.0 score/$
2
Qwen3.7 FlashAlibaba · frontier tradeoff
49.1 score$0.13 / 1M out377.7 score/$
3
DeepSeek V3.2DeepSeek · frontier tradeoff
50.9 score$0.42 / 1M out121.2 score/$
4
MiMo-V2.6-ProXiaomi · frontier tradeoff
75.5 score$0.87 / 1M out86.8 score/$
5
Claude Sonnet 5.5Anthropic · within 90% of leader
83.3 score$10.00 / 1M out8.3 score/$
6
Claude Opus 5.5Anthropic · within 90% of leader
87.7 score$20.00 / 1M out4.4 score/$
7
GPT-6 AstraOpenAI · within 90% of leader
88.7 score$50.00 / 1M out1.8 score/$

Inside one provider the tradeoff is easiest to see: Gemini 3.8 Flash or Gemini 3.1 Pro is a price-for-capability decision within Google's own line.

Questions

What is the LLM price-performance chart?

This chart plots each AI model by its benchmark score (vertical axis) against its API output price per million tokens (horizontal axis). Models in the upper-left quadrant offer the best value — high performance at low cost. The efficiency frontier line connects the best-value models at each price point.

What is the efficiency frontier?

The efficiency frontier (Pareto frontier) connects models where no other model offers both a higher score and a lower price. Models on this line represent the optimal price-performance tradeoff. If a model is below and to the right of the frontier, there exists a cheaper model with a better score.

Which LLM has the best price-to-performance ratio?

Claude Sonnet 5.5 is the lowest-priced model that retains at least 90% of the current overall leader's score under these filters. We use this threshold instead of declaring the largest Score/$ ratio the universal winner, because normalized benchmark points are an index rather than units of completed work.

How are scores calculated?

The chart reads the active public ranking lane for each category and pairs that score with current catalog pricing. Overall, coding, and agentic views use the current BenchAlign methodology when enabled; other categories use the provisional composite. The 90% shortlist is a decision aid, not a claim that benchmark points convert directly into dollars.

Keep the value shortlist current

One weekly email when price changes or new benchmark evidence alter the models worth shortlisting.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.