Weekly LLM benchmark digest
DeepSeek scheduled a reroute, then cancelled it in a footnote
On September 10 DeepSeek shipped V4.1 Flash, retired V4 Flash the same day, and said V4 Pro was next. From 04:00 UTC on September 14, every deepseek-v4-pro request would be answered by V4.1 Flash and billed at Flash rates until a V4.1 Pro exists. We re-read DeepSeek’s pricing page on September 17. A footnote there now says the opposite: V4 Pro stays in service after September 14, billed as before, and DeepSeek will give notice if that changes. The September 10 announcement has not been edited and still describes the reroute. Two DeepSeek pages, two answers. The gap between them is on the invoice. At peak rates V4 Pro lists $1.32 in and $3.96 out per million tokens. V4.1 Flash lists $0.30 and $1.20. That is 4.4× on input and 3.3× on output for one model ID, depending on which page describes the day you call it. Read the model field your responses come back with, and do not budget from an announcement.
DeepSeek V4.1 Flash: prices, context and launch scores →Three meters moved this week
The limit
DeepSeek and OpenAI pages read September 17, 2026. Google’s pricing page read September 15. All figures are list prices. Voice models are display-only on our site and sit outside the weighted rankings.
Also launched
Cognition SWE-2, September 10.
Post-trained from Kimi K3 and served only inside Devin. There is no API, no price and no weights, so every score is Cognition’s own run.
Atria Dawn Preview, September 14.
Shanghai AI Laboratory’s 744B mixture-of-experts model built on GLM-5.2, with MIT weights and a 256K context. Every result is provider-reported, so it stays unranked until someone else runs it.
Cohere North Small Translate, September 10.
218B total parameters, 25B active, 16K in and out. Free on Cohere’s API until rate limits. The weights are CC BY-NC, which rules out commercial self-hosting.
Gemini 3.8 Live Extended Thinking, September 15.
Google reports 82.6% on its speech-to-speech index and 68.6% task success on τ-Voice at high effort. The default Live model scores 76.0% and 30.1% at the same price.
Analysis worth opening
This issue in numbers
- 4.4×
- V4 Pro over V4.1 Flash, peak input price, one model ID
- $0.05 / min
- GPT-Live-1 voice layer, backend model billed on top
- 782 vs 303
- Prompt tokens billed for identical messages, Sonnet 5 and Terra
- 1 year
- Until prepaid credits expire at OpenAI, Anthropic and Google