Skip to main content
BenchLM
Data
Provider pricing hub

DeepSeek API Pricing (October 2026)

The DeepSeek API costs $0.30/$1.20 per million input/output tokens at peak for DeepSeek V4.1 Flash, served as deepseek-flash since September 10, 2026, and half that off-peak ($0.15/$0.60). Cache-hit input is $0.006 at peak. DeepSeek V4 Pro costs $1.32/$3.96 at peak and $0.66/$1.98 off-peak, rates in force since August 16. DeepSeek announced that deepseek-v4-pro would route to V4.1 Flash from September 14, then said on its pricing page that V4 Pro service continues with billing unchanged. V4 Flash is retired.

Last synced
36 confirmed releases in the last 30 daysFollow DeepSeek price changes with RadarSources attached

DeepSeek is the price floor of the frontier-adjacent API market, and on September 10, 2026 it replaced its volume tier. DeepSeek V4.1 Flash ($0.30/$1.20 at peak, $0.15/$0.60 off-peak) is the first model in a new architecture family: 552B parameters with 8B active for input and 16B for output, native image input, MIT-licensed weights, and a KV cache DeepSeek says is a quarter the size of V4 Flash's. The flagship V4 Pro ($1.32/$3.96 at peak, $0.66/$1.98 off-peak) is still served. DeepSeek said on September 10 that it was phasing V4 Pro out because V4.1 Flash beats it on performance, cost, speed, and total time, then kept it in service. Automatic context caching still makes repeated input cheaper.

The model IDs changed with the launch. deepseek-flash serves V4.1 Flash. The retired deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs temporarily route to V4.1 Flash at the Flash rate. DeepSeek's September 10 announcement said deepseek-v4-pro requests would route there too from 04:00 UTC on September 14, 2026, billed at the V4.1 Flash price, until V4.1 Pro launches. That did not happen as announced. A footnote on DeepSeek's pricing page, read September 17, 2026, says V4 Pro API service continues after September 14 with the billing method unchanged, and that DeepSeek will give notice if that changes. The announcement had not been edited when we read it the same day, so the two pages disagree. Treat the pricing page as current.

Cheap only matters if the quality clears your bar: see the Chinese model rankings, the price-vs-performance view, and the Chinese LLM deep dive. For the complete current and historical roster, use the DeepSeek model directory.

Citable stat2 current DeepSeek V4 API tiers span $0.3–$1.32 per million input tokens and $1.2–$3.96 per million output tokens as of 2026-09-25.

Operator receipt

I checked DeepSeek's live pricing, release, cache, model-ID, and concurrency documentation on August 12, 2026. At the calculator's default 10M input and 2M output tokens with no cache hits, V4 Pro cost $6.09 and V4 Flash cost $1.96 at the rates published then. DeepSeek had signaled a future overall increase but had not published the replacement rates or an effective date.

The V4.1 Flash row was synced from DeepSeek's September 10, 2026 pricing table on launch day. At the same default volume it costs $5.40 at peak and $2.70 off-peak. The V4 Pro row was refreshed on September 25, 2026 to the peak/off-peak rates DeepSeek has charged since 16:00 UTC on August 16: at the same default volume it costs $21.12 at peak and $10.56 off-peak.

Honest limit. The estimate excludes taxes, retries, reasoning-token expansion, any fallback provider, and any future price change. Cache savings use the hit share you enter; DeepSeek describes caching as best effort and does not guarantee a hit rate.

Reviewed by Glevd on August 12, 2026. V4.1 Flash row added September 10, 2026. V4 Pro row refreshed September 25, 2026.

Every DeepSeek API price per 1M tokens

DeepSeek API prices per one million tokens, by model.
ModelAPI model IDInput $/MCached input $/MOutput $/MContextEffectivePrice source
DeepSeek V4 Pro 0813deepseek-v4-proCurrent$1.32$0.044$3.961MAug 16, 2026Official current pricing
DeepSeek V4.1 Flashdeepseek-flashCurrent$0.3$0.006$1.21MSep 10, 2026Official current pricing
DeepSeek R1deepseek-reasonerHistorical mapping$0.55$0.14$2.19128KJan 20, 2025Official R1 release
DeepSeek V3.2 (Thinking)deepseek-reasonerRetired alias$0.55$0.14$2.19128KDec 1, 2025Official legacy pricing
DeepSeek V3.2deepseek-chatRetired alias$0.28$0.028$0.42128KDec 1, 2025Official legacy pricing
DeepSeek V3deepseek-chatHistorical mapping$0.27$0.07$1.1128KFeb 8, 2025Official V3 release
DeepSeek V4 Flash 0731deepseek-v4-flashRetired$0.14$0.0028$0.281MJul 31, 2026Official changelog (retired Sep 10, 2026)

V4.1 Flash and V4 Pro rates are DeepSeek's peak schedule (01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays); off-peak hours bill at half. DeepSeek lists no Batch API discount; High/Max configurations inherit their family's rate and are collapsed into one row.

Estimate your monthly DeepSeek API bill

Estimated monthly DeepSeek API cost for the workload above, by model.
ModelEst. monthly cost
DeepSeek V4 Pro 0813$21.12
DeepSeek V4.1 Flash$5.4

Estimates use DeepSeek's standard per-token rates from the table above. For token-level presets and cross-provider comparison, use the full LLM API pricing calculator.

DeepSeek API quickstart: one authenticated request

DeepSeek exposes an OpenAI-compatible base URL at https://api.deepseek.com. Create a key in the DeepSeek platform, keep it on the server, and send it as a Bearer token. This request uses the current Flash model name rather than the retired deepseek-v4-flash or deepseek-chat aliases:

curl https://api.deepseek.com/chat/completions \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $DEEPSEEK_API_KEY" \
  -d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Reply with OK"}]}'

A successful response includes token usage. Do not place the API key in browser code, logs, or a public repository. The LLM API pricing calculator can then turn measured input, output, and cache-hit tokens into a monthly estimate.

The published rates remain the billing baseline

DeepSeek's September 10, 2026 pricing table lists peak and off-peak rates for cache-hit input, cache-miss input, and output on deepseek-flash and deepseek-v4-pro. Off-peak is half of peak; peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays. The V4.1 Flash and V4 Pro rows on this page store the peak rate so an estimate never understates the bill.

DeepSeek moved V4 Pro to the peak/off-peak schedule at 16:00 UTC on August 16, 2026. Its pricing page, read September 25, 2026, lists deepseek-v4-pro at $1.32 cache-miss input, $0.044 cache-hit input, and $3.96 output at peak, and $0.66, $0.022, and $1.98 off-peak. The table, structured price data, and calculator use those peak rates. The calculator does not apply the off-peak discount. Halve the result for work that runs entirely off-peak.

Cache-hit pricing: DeepSeek’s biggest discount is automatic

DeepSeek prices input tokens in two lanes: cache miss (the standard rate) and cache hit, billed when the service can reuse a prompt prefix. A V4.1 Flash cache hit costs $0.006 per million tokens at peak and $0.003 off-peak. V4 Pro's cache-hit input is $0.044 at peak and $0.022 off-peak. Those are explicit provider rates, not a percentage inferred from the standard input price. DeepSeek says V4.1 Flash's smaller KV cache is what let it cut prices, and cache-hit charges are often the largest share of an agent bill.

Agents, RAG pipelines, and chat applications can benefit when they resend the same system prompt or document context. But the table does not promise a hit rate: a workload with mostly new prefixes pays the cache-miss rate. Enter your measured cached share in the estimator above. DeepSeek lists no separate Batch API tier.

The response exposes prompt_cache_hit_tokens and prompt_cache_miss_tokens in its usage object. Log those two fields and calculate the real hit share before budgeting. DeepSeek says unused cache entries are usually cleared after hours to days, so a one-time warm-up does not establish a permanent discount.

Concurrency and 429 handling

DeepSeek currently lists account-level concurrency limits of 500 for V4 Pro and 2,500 for V4.1 Flash. A request occupies one slot until its response completes. Requests beyond the account limit receive HTTP 429, regardless of how many API keys the account uses.

Production clients should cap concurrency, retry 429 responses with jittered backoff, and keep a fallback path for availability-sensitive work. DeepSeek also documents empty-line or SSE keep-alive traffic while a request waits; custom HTTP parsers must ignore those keep-alives without treating them as the JSON result.

V4.1 Flash vs V4 Pro: which DeepSeek tier to use

V4.1 Flash ($0.30/$1.20 peak) is now the default for new work. DeepSeek's launch table puts it ahead of V4 Pro on every agentic coding, cyber, and tool-use row it reports, it reads images natively, and it keeps the 1M-token context window. V4 Pro ($1.32/$3.96 peak) is the previous-generation flagship with separate reasoning-effort rows in the current Chinese model rankings. DeepSeek cancelled the September 14 reroute and still serves it, with no end date published, so keep it only for integrations that depend on it and plan the move to V4.1 Flash.

Older V3.2 and R1 rows remain in the registry for historical comparison. DeepSeek retired the legacy deepseek-chat and deepseek-reasoner aliases after July 24, 2026 at 15:59 UTC, and retired deepseek-v4-flash on September 10, 2026. New requests must use deepseek-flash or deepseek-v4-pro. DeepSeek publishes the V4.1 Flash and V4 Pro 0813 weights under the MIT license, so both models can be self-hosted.

DeepSeek API price and model-ID history

The current prices and the legacy rows answer different questions. Use the V4.1 Flash row for a new API integration; use the older rows only to reconstruct a historical bill or migration. DeepSeek changes the model behind an alias, so a familiar model ID does not prove that the underlying version stayed fixed.

DateChangeInput / output $/MFirst-party evidence
Read Sep 17, 2026Pricing-page footnote: deepseek-v4-pro service continues after Sep 14 with billing unchanged. The reroute announced on Sep 10 was cancelled; DeepSeek did not date the footnote.deepseek-v4-pro listed at $1.32 input, $3.96 output at peak; half that off-peakOfficial current pricing
Sep 10, 2026V4.1 Flash launched as deepseek-flash; V4 Flash retired and its ID routed to V4.1 Flash. DeepSeek announced that V4 Pro would route there from Sep 14.$0.30/$1.20 peak, $0.15/$0.60 off-peakV4.1 Flash release
Aug 16, 2026Peak/off-peak billing began at 16:00 UTC for V4 Pro and V4 Flash, with off-peak at half of peak.V4 Pro $1.32/$3.96 peak, $0.66/$1.98 off-peakDeepSeek changelog
Jul 31, 2026The existing Flash API ID moved to DeepSeek-V4-Flash-0731 public beta.Published rates unchangedV4 Flash 0731 update
Jul 24, 2026deepseek-chat and deepseek-reasoner reached their retirement deadline.No new rate; migrate model IDsV4 release notice
Apr 24, 2026V4 Flash and V4 Pro launched with dedicated API model IDs.$0.14/$0.28 and $0.435/$0.87V4 release notice
Dec 1, 2025The two legacy aliases moved to V3.2 non-thinking and thinking modes.$0.28/$0.42 and $0.55/$2.19DeepSeek changelog

The price table above links every row to DeepSeek's current or archived first-party evidence and gives the effective version date. BenchLM does not treat historical alias pricing as a currently callable SKU.

DeepSeek vs OpenAI and Claude pricing

The gap is not subtle. On output, V4 Pro at $1.32/$3.96 peak costs about 3x less than GPT-5.6 Terra ($2/$12), about 5x less than GPT-5.6 Sol ($20) and about 6x less than Claude Opus 4.8 ($25). V4.1 Flash is priced against GPT-5.6 Luna ($0.20/$1.20) instead: at its $0.30/$1.20 peak rate it costs about 50% more than Luna on input and the same as Luna on output. Off-peak, at half the peak rate, it costs about 25% less than Luna on input and half as much as Luna on output. The honest caveat: the Western flagships still score higher on BenchLM's overall leaderboard, so the decision is quality-per-dollar, not price alone.

Compare the full tables: OpenAI API pricing and Claude API pricing, or use the current product-level DeepSeek vs Claude comparison. The Opus 5 vs V4 Pro model page keeps the benchmark-by-benchmark evidence separate.

DeepSeek API pricing FAQ

How much does the DeepSeek API cost?

DeepSeek V4.1 Flash costs $0.30 input / $1.20 output per million tokens at peak and $0.15/$0.60 off-peak, with cache-hit input at $0.006 peak. DeepSeek V4 Pro costs $1.32 input / $3.96 output per million tokens at peak and $0.66/$1.98 off-peak, with cache-hit input at $0.044 peak, the rates in force since August 16, 2026. A reroute of V4 Pro to V4.1 Flash announced for September 14 was cancelled. Use deepseek-flash; the deepseek-v4-flash ID was retired on September 10, 2026, and the old deepseek-chat and deepseek-reasoner aliases reached their retirement deadline on July 24, 2026. DeepSeek lists no separate batch tier.

Is the DeepSeek API free?

No — the hosted API is pay-per-token, though at prices low enough that light usage costs cents per day. DeepSeek publishes open weights under the MIT license for V4.1 Flash, the April V4 Preview checkpoints and the V4 Pro 0813 checkpoint the API serves. The genuinely free path is self-hosting an available open checkpoint or quantized variant on your own hardware; see the open-source model rankings.

Does the DeepSeek API have a subscription?

No. The DeepSeek API is prepaid: you top up a balance and each request deducts tokens at the rates above, with off-peak hours billed at half the peak rate. There is no monthly plan, no seat fee and no published free tier; any granted balance shows only inside the account console.

What is DeepSeek cache-hit pricing?

When your request reuses a recently seen prompt prefix, DeepSeek bills that portion at the cache-hit rate instead of the standard input rate — $0.006/M at peak and $0.003/M off-peak on V4.1 Flash, and $0.044/M at peak and $0.022/M off-peak on V4 Pro. The service applies caching automatically, but the hit share depends on repeated prompt prefixes. Use measured usage rather than assuming every agent or RAG workload reaches the same discount.

Which DeepSeek API model should I use?

Start with V4.1 Flash for everything new: extraction, classification, summarization, high-volume requests, and the agentic coding work that used to justify V4 Pro. DeepSeek says V4.1 Flash beats V4 Pro on performance, cost, speed, and total time. It announced a September 14, 2026 reroute of deepseek-v4-pro to V4.1 Flash, then kept V4 Pro in service with billing unchanged. Use the exact current ID deepseek-flash, then compare against your own acceptance tests before routing production traffic.

How does DeepSeek API pricing compare to OpenAI?

DeepSeek V4 Pro's output tokens cost about 5x less than GPT-5.6 Sol's at peak ($3.96 vs $20 per million). V4.1 Flash output costs $1.20 per million at peak, the same as GPT-5.6 Luna's $1.20, and $0.60 off-peak. OpenAI's flagships still lead on BenchLM's overall scores, so the right comparison is quality-per-dollar for your specific workload — see the OpenAI pricing hub.

Keep comparing