Paid audit · every number from an invoice
Best OpenRouter Alternatives in 2026: I Ran 5,400 Paid Requests Across 7 Platforms
We put real credits on seven platforms and ran an identical 900-request workload (short chat, ~20K-token long context, tool calls) through each scored platform — and its 450-request single-model half through the two appendix platforms — 5,400 requests, $13.07 invoiced. Scored from the invoices: DeepInfra is the best OpenRouter alternative on cost ($0.38/1M on GLM-5.2, 19% under OpenRouter), Cerebras on raw speed (0.35s median) — and OpenRouter itself kept the highest overall score, 97/100, which no page selling you its own gateway will tell you.
Every ranking currently on page one for this search is written by a vendor that places itself first. This page is the opposite bet: BenchLM does not sell a gateway, we paid list price on every platform out of our own pocket, and every score below divides an invoice by logged tokens. The raw per-request logs and the cost ledger are committed alongside this site, so each number traces to a run you can inspect.
Conflict-of-interest box, up front
What BenchLM sells: model monitoring (Radar) and benchmark data. What it does not sell: a gateway, a router, or inference — so nothing of ours appears in the scoring. As of 2026-09-01, BenchLM has no affiliate or referral relationship with any platform on this page; if that ever changes, this box changes first, and affiliate status will never affect inclusion, order, or verdicts (full policy). Test spend: $13.07, ours.
Short on time? The scores move when platforms change.
Radar watches model additions, price edits, and deprecations across providers daily — the update-cadence criterion below is measured with it.
Why this page exists: the current results are vendors ranking themselves
Here is page one for “openrouter alternatives” on the day this audit shipped (2026-09-01, US desktop), and who each page puts first. Eight of the nine organic results are written by a company that sells the thing being ranked; the AI Overview above them is stitched from those same vendor pages. One page — AIMultiple — actually measured anything.
| # | Ranking page | Dated | Who it puts first |
|---|---|---|---|
| 1 | DigitalOcean — "9 OpenRouter Alternatives for Multi-Model AI" | Aug 17, 2026 | DigitalOcean (itself, first in its own list) |
| 2 | Reddit r/RooCode — "Other OpenRouter-like API providers?" | ~1 year old | No product — the one non-vendor result |
| 3 | haimaker.ai — "8 Unified AI APIs Compared" | Undated | haimaker.ai (itself) |
| 4 | TrueFoundry — "Best OpenRouter Alternatives for Production" | Aug 7, 2026 | "TrueFoundry is the leading … alternative" (itself) |
| 5 | AIMLAPI — "The Best OpenRouter Alternatives in 2026" | May 6, 2026 | AI/ML API (itself) |
| 6 | GitHub — FreeRouter README | Undated | FreeRouter (its own repo) |
| 7 | AIMultiple — "AI Gateways for OpenAI" | May 13, 2026 | Ran a real latency benchmark — credit where due; no invoice-based pricing |
| 8 | Merge.dev — "4 OpenRouter alternatives" | Undated | Merge Gateway (itself, named first) |
| 9 | Morph — "11 Providers Compared" | Undated | Morph (itself); does document OpenRouter’s 5.5% credit fee |
None of those pages publishes an invoice, a per-request log, or a reproducible workload. That is the gap this page fills.
The rubric, weights first
Five criteria, weights fixed before the run, formulas published so you can re-derive every score. The workload: 300 short-chat, 100 long-context (~20K tokens, each with a needle question that verifies the model really received the context), and 50 tool-call requests — per reference model, per platform.
| Criterion | Weight | Measured by |
|---|---|---|
| Effective cost | 30 | Invoiced dollars ÷ logged tokens, per reference model — the price after markup, caching, and billing quirks, not the rate card. Cheapest platform per model sets the bar; others score proportionally. |
| Latency | 25 | p50 (70%) and p95 (30%) full-completion latency across the run, fastest platform per model as the bar. No streaming, identical shape everywhere — comparable, though not TTFT. |
| Model coverage | 20 | Half for serving both reference models; half for catalog breadth from the platform’s own /models endpoint on run day (400+ IDs = full marks). |
| Reliability & limits | 15 | Success rate (below 95% scores zero on that half), needle-hit rate on the 20K-token contexts, and tool-call success — response integrity on identical weights, not just HTTP 200s. |
| Update cadence | 10 | Which DeepSeek Flash snapshot the platform served on run day (0731 = current / unpinned = half credit / stale = none) and whether GLM-5.2 was live. Radar tracks this continuously between audits. |
Reference models: DeepSeek V4 Flash (the cheap-open-weight workhorse) and GLM-5.2 — chosen because it is the one frontier-class open model all five scored platforms actually host. Groq and Cerebras host neither, so they run as an unscored latency appendix rather than getting a misleading /100.
Master table: every platform, scored from the run
| Platform | Type | Score /100 | Cost /30 | Latency /25 | Coverage /20 | Reliability /15 | Cadence /10 | One-line verdict |
|---|---|---|---|---|---|---|---|---|
| 1. OpenRouterWinner | Managed router | 97 | 27.2 | 25 | 20 | 14.6 | 10 | Fastest p50 on both reference models and the lowest invoiced $/1M on Flash — the incumbent won our test. |
| 2. DeepInfra | Direct provider | 71 | 26.5 | 7.1 | 14.7 | 14.9 | 7.5 | Cheapest GLM-5.2 tokens of the run via aggressive automatic caching; the trade is 2.4–5.3s medians. |
| 3. Together AI | Direct provider | 61 | 9.4 | 10.2 | 17 | 14.5 | 10 | Clean 100% ok-rate and the best GLM p95 after OpenRouter — but its GLM bill charged 76% more input tokens than its own API reported. |
| 4. Requesty | Managed router | 57 | 8.1 | 9.4 | 20 | 14.7 | 5 | Biggest catalog on paper (692 IDs, region-duplicated) but served a stale Flash snapshot at 3.5× OpenRouter’s cost. |
| 5. Fireworks AI | Direct provider | 48 | 7.1 | 7 | 10.6 | 13.1 | 10 | Priciest Flash of the run (5.3× OpenRouter), weakest auto-caching, and the run’s lowest long-context accuracy. |
| Groq | Direct provider (LPU) | Not scored — hosts neither reference model. Measured on qwen/qwen3.8-27b: 0.43s p50 / 2.00s p95, $0.89/1M blended, 99.1% ok. | Fast and rate-limited; 27B tokens priced like a frontier model. | |||||
| Cerebras | Direct provider (wafer-scale) | Not scored — hosts neither reference model. Measured on gpt-oss-120b: 0.35s p50 / 1.10s p95, $0.37/1M blended, 100% ok. | Fastest platform in the run; 2-model catalog. | |||||
Yes: the incumbent won. We went looking for OpenRouter alternatives, paid for the test, and OpenRouter kept the highest score on this workload. That result cost us $13.07 to earn and it is exactly what a page trying to sell you a gateway cannot say. The alternatives that beat it do so on a single axis — DeepInfra on cost, Cerebras on speed — and the table above shows the trade each one makes.
What one million tokens actually cost, per invoice
The table no vendor publishes: the same models, the same requests, and the invoiced price per million blended tokens — what each platform’s billing page charged, divided by the tokens its API reported back to us. Rate cards tell you the sticker; this is the till receipt, after markups, caching discounts, and billing quirks.
| Platform | DeepSeek Flash — invoiced | Flash $/1M blended | GLM-5.2 — invoiced | GLM $/1M blended | Whole run |
|---|---|---|---|---|---|
| DeepInfra | $0.10 / 1.73M tok | $0.058 | $0.68 / 1.80M tok | $0.38 | $0.78 |
| OpenRouter | $0.08 / 1.82M tok | $0.044 | $0.83 / 1.80M tok | $0.46 | $0.92 |
| Fireworks AI | $0.42 / 1.80M tok | $0.23 | $2.39 / 1.79M tok | $1.34 | $2.81 |
| Requesty | $0.25 / 1.73M tok | $0.14 | $2.95 / 1.81M tok | $1.63 | $3.20 |
| Together AI | $0.20 / 1.80M tok | $0.11 | $3.00 / 1.80M tok | $1.67 | $3.20 |
| Groq | Own model (qwen/qwen3.8-27b): $1.51 / 1.70M tok → $0.89/1M | $1.51 | |||
| Cerebras | Own model (gpt-oss-120b): $0.65 / 1.77M tok → $0.37/1M — fully covered by sign-up promo credits ($0.00 out of pocket) | $0.65 | |||
Reading it: OpenRouter’s Flash tokens cost 5.3× less than Fireworks’ for the identical requests. DeepInfra’s GLM-5.2 lands at $0.38/1M — against a $0.49/1M input rate card — because its automatic caching billed half our input at ~80% off. Together’s GLM figure is inflated by the billing discrepancy documented below. Output shares differed a few percent between snapshots, which nudges blends slightly but reorders nothing. Every cell traces to the committed ledger; totals sum to the $13.07 across all seven billing pages.
Platform by platform
1. OpenRouter — managed router — $0.46/1M on GLM-5.2 — Score 97/100
The platform this audit set out to replace posted the fastest p50 latency on both reference models, the lowest invoiced $/1M on DeepSeek Flash, and near-list pricing on GLM-5.2. On this workload, leaving OpenRouter costs you money.
Where it loses points: Cost on GLM-5.2 — DeepInfra’s automatic caching undercut it by 19% per blended token — and the credit top-up fee (about 5.5% at verification, documented rather than measured here) sits outside the per-token figures.
Take it if you want one account, the widest live catalog we measured (418 model IDs), and invoiced pricing that tracks provider list rates.
Skip it if your workload is one large open-weight model at steady volume — DeepInfra billed 19% less per GLM-5.2 token.
2. DeepInfra — direct provider — $0.38/1M on GLM-5.2 — Score 71/100
Cheapest GLM-5.2 tokens in the run — $0.38 per million blended — because DeepInfra auto-cached 53% of our Flash input and roughly half our GLM input at about 80% off, without being asked. The trade is latency: 2.4–5.3s medians, the slowest of the five scored platforms on GLM-5.2. Zero failed requests across 900 calls.
Where it loses points: Latency (7.1/25 — five-second GLM medians) and the unpinned Flash snapshot.
Take it if cost per token on large open-weight models is the whole game and multi-second medians are acceptable.
Skip it if you need sub-second responses or a pinned model snapshot.
3. Together AI — direct provider — $1.67/1M on GLM-5.2 — Score 61/100
A clean 100% ok-rate on all 900 requests and the best GLM-5.2 p95 after OpenRouter. But its bill charged us for 2,928,000 GLM-5.2 input tokens where its own API responses reported 1,667,041 — 76% more — while the identical methodology reconciled to the token on DeepSeek Flash. Observed, not explained; details in the findings section, and we invite correction.
Where it loses points: Effective cost (9.4/30) — the billed-token inflation on GLM-5.2 lands directly in the invoiced $/1M.
Take it if you want a large direct catalog with solid GLM latency and you reconcile invoices against logged usage monthly.
Skip it if billing you cannot verify is a dealbreaker — until the GLM discrepancy is explained, budget well above rate-card math.
4. Requesty — managed router — $1.63/1M on GLM-5.2 — Score 57/100
The biggest catalog on paper — 692 model IDs, inflated by per-region duplicates — and perfect needle recall on Flash. But it served us a March Flash snapshot (0424) a full month after 0731 shipped, and its GLM-5.2 p95 was 21.5 seconds, the worst tail of the run, at 3.5× OpenRouter’s invoiced cost for the same workload.
Where it loses points: Cost (8.1/30), update cadence (5/10 — stale Flash snapshot), and GLM tail latency.
Take it if you need its routing/governance features or a hosted model nobody else lists.
Skip it if you assume the same model name means the same model — snapshot pinning is on you here.
5. Fireworks AI — direct provider — $1.34/1M on GLM-5.2 — Score 48/100
The most expensive DeepSeek Flash of the run — $0.23 per million blended, 5.3× OpenRouter — with the weakest automatic caching of the platforms that itemize it, and the run’s lowest long-context accuracy: 0.83 needle-hit on GLM-5.2 where every other platform scored 0.91 or better on identical weights.
Where it loses points: Cost (7.1/30), coverage (a 23-model serverless catalog), and reliability (needle 0.83 on GLM, tool-call 0.80 on Flash).
Take it if you are buying its dedicated-deployment and enterprise serving story rather than serverless list price.
Skip it if long-context accuracy matters and you will not verify it yourself first — our 100-request needle test found a real gap.
Groq — direct provider (lpu) — $0.89/1M on qwen/qwen3.8-27b — latency appendix
Hosts neither reference model, so no score out of 100. On its own qwen3.8-27b: 0.43s p50 — but $0.89 per million tokens blended for a 27B model, more per token than DeepInfra charges for GLM-5.2, and the only platform where we hit a rate-limit ceiling mid-run (250K tokens/minute on the on-demand tier). The speed is real; so is its price.
Cerebras — direct provider (wafer-scale) — $0.37/1M on gpt-oss-120b — latency appendix
The fastest platform in the run — 0.35s p50, 1.1s p95, zero failed requests on gpt-oss-120b — with a two-model catalog. If your model is one of the two, nothing else came close; if it is not, there is nothing to route to.
Managed routers vs direct providers vs BYOK gateways
“OpenRouter alternative” hides three different architectures. The second cut of the same data:
| Category | Tested here | When it wins | What the run showed |
|---|---|---|---|
| Managed routers | OpenRouter (97), Requesty (57) | Many models, one account, no gateway to operate. | The category’s spread is the story: same architecture, 3.5× apart on invoiced cost and a two-month gap in model snapshots. |
| Direct providers | DeepInfra (71), Together (61), Fireworks (48); Groq + Cerebras in the appendix | One or two models at steady volume; you want the provider’s own caching, limits, and contracts. | Cheapest GLM tokens (DeepInfra) and fastest tokens (Cerebras) both live here — so does the run’s only billing discrepancy. |
| BYOK / self-hosted gateways | Not yet — LiteLLM, Portkey, Vercel AI Gateway, Cloudflare AI Gateway, Helicone (all UNVERIFIED) | Your own provider contracts must stay in place; the gateway adds routing, budgets, and logs on top. | Round 2 of this audit. They route your keys, so “cost” is your providers’ cost — the test that matters is overhead and reliability, and we won’t publish numbers we haven’t run. |
- LiteLLM (Self-hosted gateway (open source)): Routes your own provider keys; zero markup by design. Round-2 candidate.
- Portkey (BYOK gateway): Governance/observability layer over your keys. Round-2 candidate.
- Vercel AI Gateway (BYOK gateway): Platform-tied gateway. Round-2 candidate.
- Cloudflare AI Gateway (BYOK gateway): Edge caching/limits over your keys. Round-2 candidate.
- Helicone Gateway (BYOK gateway): Observability-first proxy. Round-2 candidate.
Where BenchLM sits in this market (unscored, with the weaknesses stated)
BenchLM’s Radar monitors what these platforms serve — new models, price edits, deprecations — which is how the update-cadence criterion stays measured between audits. Its honest limits: it will not route a single request for you, it does not measure your workload’s latency, and it covers model providers more deeply than gateway catalogs today. If you need routing, you need one of the platforms above; Radar just tells you when the ground moves under it.
What every other comparison skips
1. Nobody reconciles the bill against logged usage — we did, and found +76%
For every request we logged the token counts the platform’s own API returned, then compared them with its billing page. Six of seven platforms reconciled within noise. Together AI billed us for 2,928,000 GLM-5.2 input tokens where its API had reported 1,667,041 — 76% more (output: +24%), while the identical harness reconciled to the token on DeepSeek Flash in the same session. That pattern fits a backend counting difference on this model rather than a rate-card markup; we flag it as observed-not-explained and will update this section if Together clarifies or corrects it.
2. Automatic caching quietly decides the real price
We never requested prompt caching; two platforms applied it anyway. DeepInfra billed 53% of our Flash input at roughly 80% off and half our GLM input likewise — that is the entire reason it wins on cost. Together applied cached tiers too ($0.03/M Flash, $0.26/M GLM). Fireworks cached almost nothing ($0.06 of $1.79 GLM input spend). Identical requests, up to 5× apart on the invoice — rate cards tell you none of this.
3. “Same model” often isn’t: snapshot skew
On run day, OpenRouter, Together, and Fireworks served the current DeepSeek Flash snapshot (0731). Requesty served 0424 — two months older, a full month after 0731 shipped. DeepInfra serves an unpinned ID you cannot version at all. If your evals assume a fixed model, the router chooses otherwise. Model IDs and aliases tracks these mappings.
4. Rate-limit ceilings are real mid-run events
Groq’s on-demand tier capped us at 250K tokens/minute and returned 429s mid-run — the only platform where we hit a documented ceiling with two-way concurrency and 250ms gaps. Fine print until it interrupts production traffic.
5. Serving stacks change model quality, not just speed
The same GLM-5.2 weights scored 0.97 on our 100-question needle test on DeepInfra and 0.83 on Fireworks; tool-call success on Flash ranged from 1.00 (DeepInfra, Requesty) down to 0.80 (Together, Fireworks). Quantization, context handling, and template differences are invisible on a pricing page and very visible in output.
6. Model-launch lag compounds all of it
A platform that trails releases by a month (see Requesty’s snapshot above) silently serves you last quarter’s model at this quarter’s price. That drift is continuous, which is why this audit’s cadence criterion is fed by monitoring rather than a one-day check.
Drift is the part you can automate.
Radar records model additions, removals, and price edits with receipts — the evidence trail behind criteria 3 and 6.
Pick one in 60 seconds
| Your situation | Take | Because (from the run) |
|---|---|---|
| Many models, one account, lowest measured markup | Stay on OpenRouter | 97/100 — fastest p50 both models, $0.044/1M Flash, 418-model catalog. |
| One big open-weight model at volume, cost is everything | DeepInfra | $0.38/1M GLM-5.2 via auto-caching; accept multi-second medians. |
| Hard latency budget, model flexibility negotiable | Cerebras (then Groq) | 0.35s / 0.43s p50 — if your model is in their 2- and 14-model catalogs, and you price the premium. |
| Your own provider contracts must stay in the loop | LiteLLM / Portkey / Cloudflare (BYOK) | Untested here (round 2) — architecture fits, but we publish no numbers we haven’t bought. |
| Deciding which model to route in the first place | BenchLM’s data, then any route | Provider list prices, price-vs-performance, and Radar for when any of it changes. |
OpenRouter alternatives FAQ
Is there anything better than OpenRouter?
On a single axis, yes — measured: DeepInfra billed 19% less per GLM-5.2 token, and Cerebras answered in a third of a second. Overall, no platform we paid beat it: OpenRouter scored 97/100 on cost, latency, coverage, reliability, and cadence combined. “Better” depends on which axis your workload actually lives on — the pick-one table above maps them.
Is OpenRouter an LLM gateway?
Functionally it plays the gateway role — one API, provider routing, fallbacks — but it is a hosted marketplace with its own billing, not software you operate. A gateway you run (LiteLLM, a BYOK layer like Cloudflare’s) keeps your provider contracts and data path; OpenRouter replaces them with one account. That architectural difference, not features, is usually the real reason to switch.
Is OpenRouter free?
The account is free and some models carry free variants with tight rate limits, useful for prototyping. Paid usage tracks provider list rates — our invoiced $0.044/1M on DeepSeek Flash matches the model’s list price — with a fee added when you top up credits (documented at about 5.5%; we measured per-token cost, not the fee).
What is the cheapest OpenRouter alternative?
From the invoices: DeepInfra, at $0.38/1M blended on GLM-5.2 — the only platform that beat OpenRouter’s price on either reference model, thanks to automatic caching. On DeepSeek Flash nothing beat OpenRouter’s $0.044/1M. For direct provider list rates model by model, see LLM pricing.
OpenRouter vs Requesty — which is better?
Same architecture, very different receipts: the identical 900-request workload cost $0.92 on OpenRouter and $3.20 on Requesty (3.5×), and Requesty served a DeepSeek Flash snapshot two months older with a 21.5s GLM p95 tail. Requesty’s case is its routing/governance features and a huge nominal catalog (692 IDs) — if you need those, pin your snapshots explicitly.
What is the best free or open-source OpenRouter alternative?
LiteLLM is the standard answer — an open-source, self-hosted proxy with no markup, routing your own provider keys. We have not run the paid workload through it yet (round 2), so it carries an UNVERIFIED label here rather than a score. Note “free gateway” still means paying providers underneath; free routing is not free tokens.
Can I use these alternatives with Janitor AI or on iOS apps?
Generally yes: anything that accepts a custom OpenAI-compatible base URL and API key — which covers Janitor AI’s proxy setting and most iOS chat clients — can point at any platform tested here. Check the platform allows your use case in its terms, and mind the per-minute rate limits (Groq’s 250K tokens/minute ceiling is the kind that bites busy roleplay traffic).
About this audit
Written by Glevd, who runs BenchLM — the model-benchmark and pricing tracker this site is built on. The test was designed, paid for, and run by us on 2026-09-01; no platform was notified in advance, none supplied credits (Cerebras’ standard sign-up promo happened to cover its $0.65), and none saw this page before publication. Corrections are welcome and will be applied in place with a note.
Methodology, verifiable
Harness: plain OpenAI-compatible chat/completions calls, concurrency 2, 250ms gaps, one retry on 429, 120s timeout, no streaming — p50/p95 are full-completion latencies, identical in shape for every platform (comparable, not TTFT). Workload per model per platform: 300 short-chat + 100 long-context (~20K tokens with a needle question verifying context receipt) + 50 tool-call requests. All seven platforms ran from the same machine, the same day (2026-09-01, 14:53–18:24 UTC), on freshly loaded accounts.
Cost is invoiced dollars from each billing page divided by API-reported tokens — never rate-card transcription. Anything not measured is labeled UNVERIFIED. The per-request CSVs, per-platform summaries, the invoice ledger, and the frozen prompt set are committed in the BenchLM repository (docs/seo/openrouter-pilot-data/ and scripts/seo/openrouter-pilot/), so every number on this page can be traced to a logged request.
This page refreshes in place: the next paid run re-scores everything and bumps the verified date, and Radar tracks catalog and price movement between runs. The winner badge, when a platform requests one, follows the audit result, is never sold, and changes when the result changes.
Related pages
- LLM PricingDirect provider list rates per 1M tokens, synced daily.
- Price vs PerformanceCost-adjusted rankings and current value leaders.
- Model IDs & API AliasesWhich snapshot each provider ID actually points at.
- Alternative FinderSet a reference model and constraints, get ranked replacements.
- LLM SpeedThroughput and latency rankings across models.
Get the next audit
Quarterly re-runs of this test, plus price and model-catalog changes as Radar catches them.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.