# DeepSeek scheduled a reroute, then cancelled it in a footnote

Sent September 17, 2026

On September 10 DeepSeek shipped V4.1 Flash, retired V4 Flash the same day, and said V4 Pro was next. From 04:00 UTC on September 14, every deepseek-v4-pro request would be answered by V4.1 Flash and billed at Flash rates until a V4.1 Pro exists. We re-read DeepSeek’s pricing page on September 17. A footnote there now says the opposite: V4 Pro stays in service after September 14, billed as before, and DeepSeek will give notice if that changes. The September 10 announcement has not been edited and still describes the reroute. Two DeepSeek pages, two answers. The gap between them is on the invoice. At peak rates V4 Pro lists $1.32 in and $3.96 out per million tokens. V4.1 Flash lists $0.30 and $1.20. That is 4.4× on input and 3.3× on output for one model ID, depending on which page describes the day you call it. Read the model field your responses come back with, and do not budget from an announcement.

[DeepSeek V4.1 Flash: prices, context and launch scores →](/models/deepseek-v4-1-flash)

## Three meters moved this week

- [Our new explainer argues that inference is the part of AI you pay for. This week three providers changed the unit it is sold in.](https://benchlm.ai/blog/posts/what-is-inference-in-ai): 
- [DeepSeek bills by the hour of day.](https://benchlm.ai/llm-pricing): Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Every other hour is half price, which puts V4.1 Flash at $0.15 in and $0.60 out. A batch job that can wait costs half.
- [OpenAI bills GPT-Live-1 by the minute.](https://benchlm.ai/llm-pricing): $0.05 per minute, metered per second, for the voice layer alone. The model doing the reasoning and tool calls behind it is billed separately, so the per-minute figure is a floor.
- [Google bills Gemini 3.8 Live by the audio token.](https://benchlm.ai/llm-pricing): $3.00 per million audio tokens in and $12.00 out, which Google puts at $0.005 and $0.018 a minute. Text is $0.75 and $4.50, and on Extended Thinking the thinking tokens bill as output.

## The limit

DeepSeek and OpenAI pages read September 17, 2026. Google’s pricing page read September 15. All figures are list prices. Voice models are display-only on our site and sit outside the weighted rankings.

## Also launched

- Cognition SWE-2, September 10.: Post-trained from Kimi K3 and served only inside Devin. There is no API, no price and no weights, so every score is Cognition’s own run.
- Atria Dawn Preview, September 14.: Shanghai AI Laboratory’s 744B mixture-of-experts model built on GLM-5.2, with MIT weights and a 256K context. Every result is provider-reported, so it stays unranked until someone else runs it.
- Cohere North Small Translate, September 10.: 218B total parameters, 25B active, 16K in and out. Free on Cohere’s API until rate limits. The weights are CC BY-NC, which rules out commercial self-hosting.
- Gemini 3.8 Live Extended Thinking, September 15.: Google reports 82.6% on its speech-to-speech index and 68.6% task success on τ-Voice at high effort. The default Live model scores 76.0% and 30.1% at the same price.

## Analysis worth opening

- [Inference is the part of AI you pay for](https://benchlm.ai/blog/posts/what-is-inference-in-ai): Prefill and decode, why output costs more than input, and why reasoning models bill tokens you never see. In our September 14 run, identical messages were billed as about 782 prompt tokens on Claude Sonnet 5 and 303 on GPT-5.6 Terra.
- [API credits expire, cap you and differ by provider](https://benchlm.ai/blog/posts/api-credits-explained): OpenAI, Anthropic and Google in one table. Prepaid credits lapse after a year at all three, refunds are close to nonexistent, and the first paid tier caps a month at $100, $500 and $250 however much you have loaded.
- [Outside the benchmarks](https://techcrunch.com/2026/09/15/openai-anthropic-google-have-been-in-talks-on-ai-safety-for-weeks/): Dario Amodei published an essay on September 12 asking the frontier labs to slow down together. By September 15, OpenAI’s policy chief told reporters the company had been in talks with Anthropic and Google DeepMind for weeks about an industry standards body, and Sam Altman said OpenAI would join Anthropic in embedding third-party evaluators, TechCrunch reports. President Trump dismissed the fears as “a hoax.” None of it changes a price or a model ID yet. If outside evaluators do get inside the labs, their results are the part that would reach the tables we keep.

## This issue in numbers

- V4 Pro over V4.1 Flash, peak input price, one model ID: 4.4×
- GPT-Live-1 voice layer, backend model billed on top: $0.05 / min
- Prompt tokens billed for identical messages, Sonnet 5 and Terra: 782 vs 303
- Until prepaid credits expire at OpenAI, Anthropic and Google: 1 year

## New on BenchLM

- [Every confirmed release, with its source →](https://benchlm.ai/model-updates): 
- [What Radar does with a week like this one](https://benchlm.ai/radar): A provider changes something on a page nobody bookmarked. This week it was a footnote under DeepSeek’s price table. Radar is the tool we built for that problem: it reads 20 providers’ own pages, records each change with its source and its date, and matches it to the models and routes you declare. Four things it gives you that a changelog does not.
- [The date for the route you call.](https://benchlm.ai/radar): claude-3-haiku-20240307 retired on April 20 on Anthropic’s API, August 23 on Vertex AI and September 10 on Amazon Bedrock. Three dates, three source pages. Radar shows the one for your route, how precise it is, and the replacement the provider named, then alerts you at 90, 30 and 7 days.
- [Price changes with the receipt.](https://benchlm.ai/radar): GPT-5.6 Sol input went from $5 to $4 per million tokens on September 11, and output from $30 to $20. Each change is kept with the provider’s page and date, beside a cost scenario built from your own usage assumptions.
- [Corrections stay on the record.](https://benchlm.ai/radar): When a provider reverses itself, the way DeepSeek did this week, the correction is kept next to the original and does not overwrite it. You can see what was announced, what replaced it, and when each was read.
- [Your code stays on your machine.](https://benchlm.ai/radar): The Stack exporter finds literal model calls in JavaScript, TypeScript and Python locally, and you import only the metadata you pick. No code upload, no provider keys, no prompts retained.
- [134 Published retirement dates landing in the next 90 days; 31 Confirmed model releases in the last 30 days](https://benchlm.ai/radar): Counts from Radar’s September 4 snapshot of provider notices. Radar records published facts only and leaves unknowns unknown. It never edits code, moves traffic or runs your evaluation, and a documented change tells you what to check, not that your app is broken.
- [Declare five models free on Radar →](https://benchlm.ai/radar): Free covers retirements for five declared models, with those alerts and a morning read. No card. Pro, at $19.99 a month, records all five kinds of change across your whole stack (prices, retirements, API changes, incidents and releases) and delivers them to email, Slack, Discord or a signed webhook.

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Canonical page: https://benchlm.ai/newsletter/issues/2026-09-17
