Weekly LLM benchmark digest
I paid seven platforms to beat OpenRouter. OpenRouter won.
Every page ranking for “OpenRouter alternatives” is written by a company that sells one, and each puts itself first. None of them publishes an invoice. So we loaded credits on seven platforms, ran one identical workload through each, and scored them from the billing pages instead of the rate cards: 5,400 requests, $13.07 invoiced. OpenRouter kept the highest score, 97 out of 100 — the fastest median latency on both reference models and the lowest invoiced price on DeepSeek Flash. The alternatives that beat it do so on one axis each, and the gaps on that axis are large.
What the receipts say
The limit
Scores are frozen from the 2026-09-01 run: 900 requests on each of the five scored platforms and 450 on Groq and Cerebras, which host neither reference model and therefore sit in an unscored latency appendix. Cost is invoiced dollars divided by API-reported tokens. The Together AI billing discrepancy (2,928,000 GLM-5.2 input tokens billed against 1,667,041 reported) is stated as observed, not explained; a correction has been invited. BenchLM sells monitoring, not routing, and has no affiliate relationship with any platform scored.
What every other comparison skips
We reconciled every bill against logged usage
Six of seven platforms matched within noise. Together AI billed 76% more GLM-5.2 input tokens than its own API reported, while DeepSeek Flash reconciled to the token in the same session.
Automatic caching sets the real price
We never asked for it. Two platforms applied it anyway, one barely did, and identical requests landed up to 5× apart on the invoice.
“Same model” often isn’t
Three platforms served the current DeepSeek Flash snapshot, one served March’s, and one serves an ID you cannot pin to a version at all.
Serving stacks change quality, not just speed
Tool-call success on Flash ranged from 1.00 down to 0.80 across platforms running the same weights.
Analysis worth opening
This issue in numbers
- 5,400
- Paid requests, one identical workload
- $13.07
- Invoiced across seven billing pages
- 97 / 100
- OpenRouter, the platform we set out to replace
- +76%
- Together’s billed GLM-5.2 input vs its own API count