Weekly LLM benchmark digest
Open weight is eating the frontier
Sponsored by Celeris. The performance figures in the sent issue came from Celeris’s provider-run tests.
DeepSeek published V4 Flash 0731 on July 31 with MIT-licensed weights and a $0.14 input / $0.28 output API. Alibaba put Qwen3.8-Max into hosted preview and said weights would follow. Moonshot had already published Kimi K3’s full checkpoint on July 27. Three releases in three weeks pointed the same way, but one of them was still a promise. For teams that needed self-hosting, data residency, or control over serving, the choice was no longer only which model. It was which lab, license, and deployment path.
What changed in seven days
The limit
Qwen3.8-Max and DeepSeek V4 Flash shared no public benchmark result in the same comparison when this issue was sent. DeepSeek’s fresh agent scores were provider-run with a harness it had not released. We did not name a quality winner from that evidence.
What to expect next
Separate the promise from the artifact
A weight announcement matters. A license, checkpoint, configuration, and runnable deployment path matter more.
Cheap API tokens move the decision
Once generation costs less than a retry, task completion, hosting overhead, and operational fit become the useful measures.
Procurement now includes the lab and license
That month’s open-weight moves came from Chinese labs. That was a sourcing fact, not a security verdict; each checkpoint still needed the same legal, governance, and production review.
Analysis worth opening
Partner links
The links below lead to Celeris and are part of the sponsored edition.