Weekly LLM benchmark digest
The frontier moved again, twice
Two frontier model announcements in one week used to be unusual. Claude Opus 4.7 reached general availability at 94/100 on BenchLM, behind Claude Mythos Preview at 99. Kimi 2.7 from Moonshot AI entered preview, and 14 models had released in April by the send date.
Provider standings
| Rank | Provider | Models | Score | Change |
|---|---|---|---|---|
| 1 | Anthropic | 13 | 95 | +2 |
| 2 | OpenAI | 25 | 91.3 | -1 |
| 3 | 11 | 88 | -1 | |
| 4 | Z.AI | 6 | 81.7 | No change |
| 5 | xAI | 6 | 76.3 | No change |
Cumulative average top-three score through April 2026, as published in the issue.
37 months in numbers
- +405
- Total Elo gain
- 21
- Crown changes
- 5 months
- Longest reign
- +375
- Open-source Elo gain
- 4 points
- Closest gap recorded
Analysis worth opening
Claude API Pricing: Haiku 4.5, Sonnet 4.6, and Opus 4.7Prompt caching rates, batch discounts, and the production cost changes in the updated Claude tier stack.Gemini API Pricing: Current Flash, Flash-Lite, and Pro RatesA breakdown of the Flash-Lite, Flash, Pro, Batch, and Flex prices available at the time.How to Choose an LLM in 2026A model selection framework based on use case, budget, and deployment constraints.Best Open Source LLM in 2026GLM-5, Qwen 3.5, Gemma 4, Kimi K2.5, and Llama ranked using the public benchmark data available then.
This issue in numbers
- 194
- Models ranked
- 14
- New that month
- $23.01
- Average output price per 1M tokens
- 99
- Top overall score