LLM API Pricing Comparison and Cost Calculator
As of September 30, 2026, the cheapest LLM API is Qwen3.7 Flash at $0.03/$0.13 per million input/output tokens, of 169 paid models across 30 providers. The cheapest production-grade model (score 70+) is MiMo-V2.6-Pro at $0.43/$0.87. The cheapest frontier-tier model (score 80+) is Claude Sonnet 5.5 at $2.00/$10.00.
The cheapest API is rarely the cheapest model to run. Price per token below, price per answer in the calculator. The Score/$ column uses the public overall score per dollar of output — higher means better value.
Updated September 30, 2026. 169 paid models, 30 providers, per million tokens. Official API rates. Open-weight model costs depend on your infrastructure. How does LLM token pricing work?
Cheapest LLM APIs in 2026
The lowest token rate, the strongest score per output dollar, and the highest-scoring paid model answer different questions. Use the first for raw volume, the second for a shortlist, and the third when quality matters more than list price.
$0.1 in / $0.1 out per 1M tokens
Ministral 3 3B
Qwen3.7 Flash
Score/$ ratio: 377.3
GPT-6 Astra
Public score: 88.69 — $10 in / $50 out
LLM pricing comparison
Filter by provider, family, or model name, then sort the current API rates by input price, output price, context window, benchmark score, or score per output dollar. Cache rates remain blank unless the exact model row has a published cache-hit price.
* Score/$ = Overall benchmark score ÷ output price per million tokens. Higher is better. Free/open-weight models are excluded. “Price checked” is the registry observation date, not a claim that every provider changed its rate that day.
Best value by category
Below the top tier, the choice most teams actually face is Gemini 3.8 Flash or Claude Sonnet 5.5, compared on score and price.
Buying through a gateway or marketplace instead of a provider account? We paid seven of them to run the same workload and scored the invoices: OpenRouter alternatives, tested.
LLM API pricing calculator
Pick a starting workload or enter your own token counts. The result applies the current input, output, and published cache-hit rates from the pricing table below.
| Model | Per request | Per month | Input share | Output share | Cache treatment |
|---|---|---|---|---|---|
| DeepSeek V4.1 FlashDeepSeek | $0.0016 | $1.56 | $0.600 | $0.960 | Published: $0.006/1M |
| Grok 4.6xAI | $0.0088 | $8.80 | $4.00 | $4.80 | Published: $0.5/1M |
| Gemini 3.5 FlashGoogle | $0.010 | $10.20 | $3.00 | $7.20 | Published: $0.15/1M |
| Claude Sonnet 5Anthropic | $0.012 | $12.00 | $4.00 | $8.00 | Published: $0.2/1M |
| GPT-5.6 TerraOpenAI | $0.014 | $13.60 | $4.00 | $9.60 | Published: $0.2/1M |
Keep these assumptions with the decision. The estimate above is today’s list rates against the workload you typed; both move. Save it as a Radar cost scenario and the next review starts from the same numbers, with the source of every rate attached.
Estimate only. It excludes batch discounts, long-context surcharges, taxes, tools, images, and provider minimums. A cache share is applied only when the selected model has a published cache-hit rate. Pricing registry updated September 30, 2026.
The estimate depends on these assumptions
The rate card tells you what a token costs. Your workload decides how many paid tokens it takes to finish the job, and that is the number the calculator estimates. Keep the input, output and cache assumptions beside the estimate before you compare two models, because the same rate card produces a different bill for every workload.
The calculator applies one formula per model: uncached input tokens at the input rate, cached input tokens at the published cache-hit rate, output tokens at the output rate, each divided by one million. Requests per month multiplies everything. There is no hidden markup and no allowance for anything the provider bills separately.
What the calculation assumes
| Assumption | What the calculator does | What you should check |
|---|---|---|
| Input tokens per request | Multiplies by requests per month at the model’s input rate | Measure a real request, including the system prompt and any retrieved context |
| Output tokens per request | Multiplies by requests per month at the output rate | Output usually costs more per token than input; long answers move the bill faster than long prompts |
| Cache-hit share | Applies the published cache-hit rate to that share of input tokens | A share above zero only helps when your prompts repeat a stable prefix |
| Missing cache rate | Applies the standard input rate to every input token and shows “Standard input rate” in the cache column | A blank cache field means the registry has no published rate, not that caching is free or unavailable |
| Zero requests | Returns a zero estimate and says so | A zero bill describes zero traffic, not a free model |
| Blank, negative or non-numeric input | Uses 0 for the estimate and names the field to correct; your entry is kept | Re-enter the number; the estimate is only as good as the workload you typed |
One worked example, with invented rates
The two models below are hypothetical. Their rates are not a provider’s prices and do not appear in the table on this page. The workload is 20,000 requests a month, 1,500 input tokens and 400 output tokens per request, with 40% of input tokens served from cache.
| Hypothetical | Model A (hypothetical) | Model B (hypothetical) |
|---|---|---|
| Input rate per 1M tokens | $0.50 | $0.30 |
| Output rate per 1M tokens | $2.00 | $3.50 |
| Cache-hit rate per 1M tokens | $0.10 | Not published |
| Monthly input tokens | 30,000,000 | 30,000,000 |
| Monthly output tokens | 8,000,000 | 8,000,000 |
| Input cost | 18M × $0.50 + 12M × $0.10 = $10.20 | 30M × $0.30 = $9.00 (cache share ignored) |
| Output cost | 8M × $2.00 = $16.00 | 8M × $3.50 = $28.00 |
| Monthly estimate | $26.20 | $37.00 |
| Per request | $0.0013 | $0.0019 |
Model B has the lower input rate and the higher bill. Two things did that: output tokens cost more than input tokens on both cards, and Model B has no published cache rate, so the 40% cache share does nothing for it. Change the workload to 100 output tokens and the order can flip. That is the point of keeping the assumptions visible. We recomputed this table with the same function the calculator runs, so the figures match what the page would show for those inputs. They are arithmetic, not a saving anyone realised.
What the estimate leaves out
Retries, unless you enter them as extra requests. Reasoning tokens and tool charges that some models bill on top of the output count. Batch discounts, long-context surcharges, taxes, image inputs and provider minimums. Human review of the outputs, which is often the largest cost in a workload that needs an acceptable result rather than a response. The cheapest list rate is a shortlist, not a verdict: before you switch, run the candidates against a small fixed set of your own requests and count the outputs you would accept. The custom benchmark guide shows a repeatable way to get that number.
Confirm the setup you priced
The estimate is only valid for the exact model row. Copy the documented identifier from the model ID directory before it reaches production, and check the pricing history if the rate you priced was set more than a few weeks ago. The Radar placement on this page remains the way to follow published price changes for supported providers; it does not promise a cost alert for your own bill.
How to compare LLM API pricing
Start with the token mix your application actually sends. Monthly cost equals standard input, cached input, and output token charges across every request. A low input rate can still lose when the model produces long answers, misses the cache, or needs more retries to pass the same acceptance test.
Separate input from output
Output tokens usually cost more. Compare the ratio instead of averaging the two advertised rates.
Apply cache rates exactly
Use a cache discount only when the provider publishes one for that model and your prompts can reuse it.
Test cost against quality
Score/$ builds a shortlist. Your accepted result per dollar is the production number that matters.
A price only means something next to what it buys: GPT-6 Astra against Gemini 3.8 Flash sets the two side by side, benchmark by benchmark and rate by rate.
Pricing sources and field rules
Numeric rates come from provider-owned pricing pages, API documentation, or launch announcements for the exact model row. A missing cache field means the registry does not have an explicit cache-hit rate; it does not mean caching is unavailable or free. Tiered, reseller, and self-hosting rates stay separate.
LLM pricing updates and API price history
The live table is a current registry snapshot. Use the history view for dated model launches and price changes, or the token price index for a normalized view of how frontier inference cost moves over time.
LLM API pricing questions
How do I calculate LLM API cost?
LLM API cost equals input tokens multiplied by the input rate, plus output tokens multiplied by the output rate, divided by one million. Add cached input at the provider’s published cache rate. The calculator above applies that formula per request and per month across the models you select.
What is the cheapest LLM API in 2026?
As of September 30, 2026, Qwen3.7 Flash is the cheapest listed paid API at $0.03 input and $0.13 output per million tokens. The lowest list price is not automatically the lowest production cost: output volume, cache hits, retries, and task quality can change the result.
Which LLM has the best price-to-performance ratio?
On the current table, Qwen3.7 Flash has the highest public benchmark score per dollar of output. Score/$ is a shortlist metric, not a workload guarantee: it ignores latency, input-heavy prompts, hidden reasoning tokens, tool charges, and quality differences inside your own acceptance tests.
How does cached input pricing change LLM cost?
Cached input pricing discounts prompt tokens that a provider can reuse, such as a repeated system prompt or stable document context. The calculator applies the discount only when the exact model has a published cache-hit rate. A blank cache field does not mean caching is free or unavailable: when the rate is missing, the calculator charges every input token at the standard rate and labels the row “Standard input rate”.
Does the LLM pricing calculator include retries?
No. The estimate counts the requests you enter at the tokens you enter. A retried call is a second paid request, so add it to requests per month, or raise the per-request token counts if the retry resends the prompt. Reasoning and tool charges are also excluded unless the provider folds them into the output count.
Is the cheapest LLM API also the cheapest completed task?
Not necessarily. A low input rate loses when the model writes long answers, misses the cache or needs more retries to pass the same acceptance test. Compare cost per acceptable result on your own requests; the worked example on this page shows a lower input rate producing the higher monthly bill.
Are open-weight LLMs free to run?
Open-weight models can be free to download, but inference still requires GPUs, electricity, hosting, and engineering time. Hosted APIs for open models also charge per token. Compare the API table with the self-host estimate, then use the dedicated break-even calculator for your hardware, utilization, and traffic assumptions.
Get the weekly price-change file
Material API price changes, newly priced models, and deprecations that could change a production shortlist.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.