Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Weekly brief and archive

Weekly LLM benchmark digest

I paid seven platforms to beat OpenRouter. OpenRouter won.

Every page ranking for “OpenRouter alternatives” is written by a company that sells one, and each puts itself first. None of them publishes an invoice. So we loaded credits on seven platforms, ran one identical workload through each, and scored them from the billing pages instead of the rate cards: 5,400 requests, $13.07 invoiced. OpenRouter kept the highest score, 97 out of 100 — the fastest median latency on both reference models and the lowest invoiced price on DeepSeek Flash. The alternatives that beat it do so on one axis each, and the gaps on that axis are large.

What the receipts say

The limit

Scores are frozen from the 2026-09-01 run: 900 requests on each of the five scored platforms and 450 on Groq and Cerebras, which host neither reference model and therefore sit in an unscored latency appendix. Cost is invoiced dollars divided by API-reported tokens. The Together AI billing discrepancy (2,928,000 GLM-5.2 input tokens billed against 1,667,041 reported) is stated as observed, not explained; a correction has been invited. BenchLM sells monitoring, not routing, and has no affiliate relationship with any platform scored.

What every other comparison skips

We reconciled every bill against logged usage

Six of seven platforms matched within noise. Together AI billed 76% more GLM-5.2 input tokens than its own API reported, while DeepSeek Flash reconciled to the token in the same session.

Automatic caching sets the real price

We never asked for it. Two platforms applied it anyway, one barely did, and identical requests landed up to 5× apart on the invoice.

“Same model” often isn’t

Three platforms served the current DeepSeek Flash snapshot, one served March’s, and one serves an ID you cannot pin to a version at all.

Serving stacks change quality, not just speed

Tool-call success on Flash ranged from 1.00 down to 0.80 across platforms running the same weights.

Analysis worth opening

This issue in numbers

5,400
Paid requests, one identical workload
$13.07
Invoiced across seven billing pages
97 / 100
OpenRouter, the platform we set out to replace
+76%
Together’s billed GLM-5.2 input vs its own API count

New on BenchLM

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Want the next issue?

The signup form and another real sample are on the weekly brief page.

Subscribe to the weekly brief