Skip to main content
Weekly brief and archive

Weekly LLM benchmark digest

The frontier is outrunning the scorecards

GPT-5.6 became generally available on July 9. Kimi K3 arrived seven days later. In between, another frontier lab released a 975-billion-parameter open-weight model. The release cycle is compressing, and the next model can now arrive before the last one has enough independent evidence to compare cleanly.

What changed in seven days

The limit

Kimi K3 was still unranked when this issue was sent. Its launch table mixed model and harness effects, and the promised weights were due by July 27. Official results were useful launch evidence, not a clean substitute for independent runs in a comparable setup.

What to expect next

Families, not single flagships

Labs can refresh capability tiers on separate schedules, which makes a model name less durable than its exact version and effort setting.

More agent-dependent scores

The host model, harness, reasoning budget, and parallelism increasingly determine the result together.

Shorter buying windows

Keep a small production fixture and compare cost per completed task. A general leaderboard can narrow the field; it cannot reproduce your stack.

Analysis worth opening

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Want the next issue?

The signup form and another real sample are on the weekly brief page.

Subscribe to the weekly brief