Skip to main content
BenchLM
Weekly brief and archive

Weekly LLM benchmark digest

DeepSeek scheduled a reroute, then cancelled it in a footnote

On September 10 DeepSeek shipped V4.1 Flash, retired V4 Flash the same day, and said V4 Pro was next. From 04:00 UTC on September 14, every deepseek-v4-pro request would be answered by V4.1 Flash and billed at Flash rates until a V4.1 Pro exists. We re-read DeepSeek’s pricing page on September 17. A footnote there now says the opposite: V4 Pro stays in service after September 14, billed as before, and DeepSeek will give notice if that changes. The September 10 announcement has not been edited and still describes the reroute. Two DeepSeek pages, two answers. The gap between them is on the invoice. At peak rates V4 Pro lists $1.32 in and $3.96 out per million tokens. V4.1 Flash lists $0.30 and $1.20. That is 4.4× on input and 3.3× on output for one model ID, depending on which page describes the day you call it. Read the model field your responses come back with, and do not budget from an announcement.

DeepSeek V4.1 Flash: prices, context and launch scores →

Three meters moved this week

The limit

DeepSeek and OpenAI pages read September 17, 2026. Google’s pricing page read September 15. All figures are list prices. Voice models are display-only on our site and sit outside the weighted rankings.

Also launched

Cognition SWE-2, September 10.

Post-trained from Kimi K3 and served only inside Devin. There is no API, no price and no weights, so every score is Cognition’s own run.

Atria Dawn Preview, September 14.

Shanghai AI Laboratory’s 744B mixture-of-experts model built on GLM-5.2, with MIT weights and a 256K context. Every result is provider-reported, so it stays unranked until someone else runs it.

Cohere North Small Translate, September 10.

218B total parameters, 25B active, 16K in and out. Free on Cohere’s API until rate limits. The weights are CC BY-NC, which rules out commercial self-hosting.

Gemini 3.8 Live Extended Thinking, September 15.

Google reports 82.6% on its speech-to-speech index and 68.6% task success on τ-Voice at high effort. The default Live model scores 76.0% and 30.1% at the same price.

Analysis worth opening

This issue in numbers

4.4×
V4 Pro over V4.1 Flash, peak input price, one model ID
$0.05 / min
GPT-Live-1 voice layer, backend model billed on top
782 vs 303
Prompt tokens billed for identical messages, Sonnet 5 and Terra
1 year
Until prepaid credits expire at OpenAI, Anthropic and Google

New on BenchLM

Every confirmed release, with its source →What Radar does with a week like this oneA provider changes something on a page nobody bookmarked. This week it was a footnote under DeepSeek’s price table. Radar is the tool we built for that problem: it reads 20 providers’ own pages, records each change with its source and its date, and matches it to the models and routes you declare. Four things it gives you that a changelog does not.The date for the route you call.claude-3-haiku-20240307 retired on April 20 on Anthropic’s API, August 23 on Vertex AI and September 10 on Amazon Bedrock. Three dates, three source pages. Radar shows the one for your route, how precise it is, and the replacement the provider named, then alerts you at 90, 30 and 7 days.Price changes with the receipt.GPT-5.6 Sol input went from $5 to $4 per million tokens on September 11, and output from $30 to $20. Each change is kept with the provider’s page and date, beside a cost scenario built from your own usage assumptions.Corrections stay on the record.When a provider reverses itself, the way DeepSeek did this week, the correction is kept next to the original and does not overwrite it. You can see what was announced, what replaced it, and when each was read.Your code stays on your machine.The Stack exporter finds literal model calls in JavaScript, TypeScript and Python locally, and you import only the metadata you pick. No code upload, no provider keys, no prompts retained.134 Published retirement dates landing in the next 90 days; 31 Confirmed model releases in the last 30 daysCounts from Radar’s September 4 snapshot of provider notices. Radar records published facts only and leaves unknowns unknown. It never edits code, moves traffic or runs your evaluation, and a documented change tells you what to check, not that your app is broken.Declare five models free on Radar →Free covers retirements for five declared models, with those alerts and a morning read. No card. Pro, at $19.99 a month, records all five kinds of change across your whole stack (prices, retirements, API changes, incidents and releases) and delivers them to email, Slack, Discord or a signed webhook.

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Want the next issue?

The signup form and another real sample are on the weekly brief page.

Subscribe to the weekly brief