Weekly LLM benchmark digest
The frontier is outrunning the scorecards
GPT-5.6 became generally available on July 9. Kimi K3 arrived seven days later. In between, another frontier lab released a 975-billion-parameter open-weight model. The release cycle is compressing, and the next model can now arrive before the last one has enough independent evidence to compare cleanly.
What changed in seven days
The limit
Kimi K3 was still unranked when this issue was sent. Its launch table mixed model and harness effects, and the promised weights were due by July 27. Official results were useful launch evidence, not a clean substitute for independent runs in a comparable setup.
What to expect next
Families, not single flagships
Labs can refresh capability tiers on separate schedules, which makes a model name less durable than its exact version and effort setting.
More agent-dependent scores
The host model, harness, reasoning budget, and parallelism increasingly determine the result together.
Shorter buying windows
Keep a small production fixture and compare cost per completed task. A general leaderboard can narrow the field; it cannot reproduce your stack.