Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

Paid audit · every number from an invoice

Best OpenRouter Alternatives in 2026: I Ran 5,400 Paid Requests Across 7 Platforms

By Glevd, running BenchLM since 2025Published Sep 1, 2026Test run + prices verified

We put real credits on seven platforms and ran an identical 900-request workload (short chat, ~20K-token long context, tool calls) through each scored platform — and its 450-request single-model half through the two appendix platforms — 5,400 requests, $13.07 invoiced. Scored from the invoices: DeepInfra is the best OpenRouter alternative on cost ($0.38/1M on GLM-5.2, 19% under OpenRouter), Cerebras on raw speed (0.35s median) — and OpenRouter itself kept the highest overall score, 97/100, which no page selling you its own gateway will tell you.

Every ranking currently on page one for this search is written by a vendor that places itself first. This page is the opposite bet: BenchLM does not sell a gateway, we paid list price on every platform out of our own pocket, and every score below divides an invoice by logged tokens. The raw per-request logs and the cost ledger are committed alongside this site, so each number traces to a run you can inspect.

Conflict-of-interest box, up front

What BenchLM sells: model monitoring (Radar) and benchmark data. What it does not sell: a gateway, a router, or inference — so nothing of ours appears in the scoring. As of 2026-09-01, BenchLM has no affiliate or referral relationship with any platform on this page; if that ever changes, this box changes first, and affiliate status will never affect inclusion, order, or verdicts (full policy). Test spend: $13.07, ours.

Short on time? The scores move when platforms change.

Radar watches model additions, price edits, and deprecations across providers daily — the update-cadence criterion below is measured with it.

Open Radar

Why this page exists: the current results are vendors ranking themselves

Here is page one for “openrouter alternatives” on the day this audit shipped (2026-09-01, US desktop), and who each page puts first. Eight of the nine organic results are written by a company that sells the thing being ranked; the AI Overview above them is stitched from those same vendor pages. One page — AIMultiple — actually measured anything.

#Ranking pageDatedWho it puts first
1DigitalOcean — "9 OpenRouter Alternatives for Multi-Model AI"Aug 17, 2026DigitalOcean (itself, first in its own list)
2Reddit r/RooCode — "Other OpenRouter-like API providers?"~1 year oldNo product — the one non-vendor result
3haimaker.ai — "8 Unified AI APIs Compared"Undatedhaimaker.ai (itself)
4TrueFoundry — "Best OpenRouter Alternatives for Production"Aug 7, 2026"TrueFoundry is the leading … alternative" (itself)
5AIMLAPI — "The Best OpenRouter Alternatives in 2026"May 6, 2026AI/ML API (itself)
6GitHub — FreeRouter READMEUndatedFreeRouter (its own repo)
7AIMultiple — "AI Gateways for OpenAI"May 13, 2026Ran a real latency benchmark — credit where due; no invoice-based pricing
8Merge.dev — "4 OpenRouter alternatives"UndatedMerge Gateway (itself, named first)
9Morph — "11 Providers Compared"UndatedMorph (itself); does document OpenRouter’s 5.5% credit fee

None of those pages publishes an invoice, a per-request log, or a reproducible workload. That is the gap this page fills.

The rubric, weights first

Five criteria, weights fixed before the run, formulas published so you can re-derive every score. The workload: 300 short-chat, 100 long-context (~20K tokens, each with a needle question that verifies the model really received the context), and 50 tool-call requests — per reference model, per platform.

CriterionWeightMeasured by
Effective cost30Invoiced dollars ÷ logged tokens, per reference model — the price after markup, caching, and billing quirks, not the rate card. Cheapest platform per model sets the bar; others score proportionally.
Latency25p50 (70%) and p95 (30%) full-completion latency across the run, fastest platform per model as the bar. No streaming, identical shape everywhere — comparable, though not TTFT.
Model coverage20Half for serving both reference models; half for catalog breadth from the platform’s own /models endpoint on run day (400+ IDs = full marks).
Reliability & limits15Success rate (below 95% scores zero on that half), needle-hit rate on the 20K-token contexts, and tool-call success — response integrity on identical weights, not just HTTP 200s.
Update cadence10Which DeepSeek Flash snapshot the platform served on run day (0731 = current / unpinned = half credit / stale = none) and whether GLM-5.2 was live. Radar tracks this continuously between audits.

Reference models: DeepSeek V4 Flash (the cheap-open-weight workhorse) and GLM-5.2 — chosen because it is the one frontier-class open model all five scored platforms actually host. Groq and Cerebras host neither, so they run as an unscored latency appendix rather than getting a misleading /100.

Master table: every platform, scored from the run

PlatformTypeScore /100Cost /30Latency /25Coverage /20Reliability /15Cadence /10One-line verdict
1. OpenRouterWinnerManaged router9727.2252014.610Fastest p50 on both reference models and the lowest invoiced $/1M on Flash — the incumbent won our test.
2. DeepInfraDirect provider7126.57.114.714.97.5Cheapest GLM-5.2 tokens of the run via aggressive automatic caching; the trade is 2.4–5.3s medians.
3. Together AIDirect provider619.410.21714.510Clean 100% ok-rate and the best GLM p95 after OpenRouter — but its GLM bill charged 76% more input tokens than its own API reported.
4. RequestyManaged router578.19.42014.75Biggest catalog on paper (692 IDs, region-duplicated) but served a stale Flash snapshot at 3.5× OpenRouter’s cost.
5. Fireworks AIDirect provider487.1710.613.110Priciest Flash of the run (5.3× OpenRouter), weakest auto-caching, and the run’s lowest long-context accuracy.
GroqDirect provider (LPU)Not scored — hosts neither reference model. Measured on qwen/qwen3.8-27b: 0.43s p50 / 2.00s p95, $0.89/1M blended, 99.1% ok.Fast and rate-limited; 27B tokens priced like a frontier model.
CerebrasDirect provider (wafer-scale)Not scored — hosts neither reference model. Measured on gpt-oss-120b: 0.35s p50 / 1.10s p95, $0.37/1M blended, 100% ok.Fastest platform in the run; 2-model catalog.

Yes: the incumbent won. We went looking for OpenRouter alternatives, paid for the test, and OpenRouter kept the highest score on this workload. That result cost us $13.07 to earn and it is exactly what a page trying to sell you a gateway cannot say. The alternatives that beat it do so on a single axis — DeepInfra on cost, Cerebras on speed — and the table above shows the trade each one makes.

What one million tokens actually cost, per invoice

The table no vendor publishes: the same models, the same requests, and the invoiced price per million blended tokens — what each platform’s billing page charged, divided by the tokens its API reported back to us. Rate cards tell you the sticker; this is the till receipt, after markups, caching discounts, and billing quirks.

PlatformDeepSeek Flash — invoicedFlash $/1M blendedGLM-5.2 — invoicedGLM $/1M blendedWhole run
DeepInfra$0.10 / 1.73M tok$0.058$0.68 / 1.80M tok$0.38$0.78
OpenRouter$0.08 / 1.82M tok$0.044$0.83 / 1.80M tok$0.46$0.92
Fireworks AI$0.42 / 1.80M tok$0.23$2.39 / 1.79M tok$1.34$2.81
Requesty$0.25 / 1.73M tok$0.14$2.95 / 1.81M tok$1.63$3.20
Together AI$0.20 / 1.80M tok$0.11$3.00 / 1.80M tok$1.67$3.20
GroqOwn model (qwen/qwen3.8-27b): $1.51 / 1.70M tok → $0.89/1M$1.51
CerebrasOwn model (gpt-oss-120b): $0.65 / 1.77M tok → $0.37/1M — fully covered by sign-up promo credits ($0.00 out of pocket)$0.65

Reading it: OpenRouter’s Flash tokens cost 5.3× less than Fireworks’ for the identical requests. DeepInfra’s GLM-5.2 lands at $0.38/1M — against a $0.49/1M input rate card — because its automatic caching billed half our input at ~80% off. Together’s GLM figure is inflated by the billing discrepancy documented below. Output shares differed a few percent between snapshots, which nudges blends slightly but reorders nothing. Every cell traces to the committed ledger; totals sum to the $13.07 across all seven billing pages.

Platform by platform

1. OpenRoutermanaged router$0.46/1M on GLM-5.2 — Score 97/100

The platform this audit set out to replace posted the fastest p50 latency on both reference models, the lowest invoiced $/1M on DeepSeek Flash, and near-list pricing on GLM-5.2. On this workload, leaving OpenRouter costs you money.

Flash: $0.044/1M · p50 0.76s · p95 2.59sGLM-5.2: $0.46/1M · p50 1.20s · p95 3.69sNeedle-hit: 0.94 / 0.93 · Tool calls: 0.98 / 1.00Catalog: 418 model IDs · Served deepseek-v4-flash-0731 — the current snapshot — and GLM-5.2 day one.

Where it loses points: Cost on GLM-5.2 — DeepInfra’s automatic caching undercut it by 19% per blended token — and the credit top-up fee (about 5.5% at verification, documented rather than measured here) sits outside the per-token figures.

Take it if you want one account, the widest live catalog we measured (418 model IDs), and invoiced pricing that tracks provider list rates.

Skip it if your workload is one large open-weight model at steady volume — DeepInfra billed 19% less per GLM-5.2 token.

2. DeepInfradirect provider$0.38/1M on GLM-5.2 — Score 71/100

Cheapest GLM-5.2 tokens in the run — $0.38 per million blended — because DeepInfra auto-cached 53% of our Flash input and roughly half our GLM input at about 80% off, without being asked. The trade is latency: 2.4–5.3s medians, the slowest of the five scored platforms on GLM-5.2. Zero failed requests across 900 calls.

Flash: $0.058/1M · p50 2.42s · p95 6.93sGLM-5.2: $0.38/1M · p50 5.29s · p95 13.82sNeedle-hit: 1.00 / 0.97 · Tool calls: 1.00 / 1.00Catalog: 189 model IDs · Serves an unpinned "DeepSeek-V4-Flash" ID — current weights, but you cannot pin a snapshot.

Where it loses points: Latency (7.1/25 — five-second GLM medians) and the unpinned Flash snapshot.

Take it if cost per token on large open-weight models is the whole game and multi-second medians are acceptable.

Skip it if you need sub-second responses or a pinned model snapshot.

3. Together AIdirect provider$1.67/1M on GLM-5.2 — Score 61/100

A clean 100% ok-rate on all 900 requests and the best GLM-5.2 p95 after OpenRouter. But its bill charged us for 2,928,000 GLM-5.2 input tokens where its own API responses reported 1,667,041 — 76% more — while the identical methodology reconciled to the token on DeepSeek Flash. Observed, not explained; details in the findings section, and we invite correction.

Flash: $0.11/1M · p50 2.55s · p95 10.09sGLM-5.2: $1.67/1M · p50 2.53s · p95 5.66sNeedle-hit: 0.99 / 0.93 · Tool calls: 0.80 / 1.00Catalog: 278 model IDs · Served DeepSeek-V4-Flash-0731 (current) and GLM-5.2; GLM-5.3 already listed in its catalog on run day.

Where it loses points: Effective cost (9.4/30) — the billed-token inflation on GLM-5.2 lands directly in the invoiced $/1M.

Take it if you want a large direct catalog with solid GLM latency and you reconcile invoices against logged usage monthly.

Skip it if billing you cannot verify is a dealbreaker — until the GLM discrepancy is explained, budget well above rate-card math.

4. Requestymanaged router$1.63/1M on GLM-5.2 — Score 57/100

The biggest catalog on paper — 692 model IDs, inflated by per-region duplicates — and perfect needle recall on Flash. But it served us a March Flash snapshot (0424) a full month after 0731 shipped, and its GLM-5.2 p95 was 21.5 seconds, the worst tail of the run, at 3.5× OpenRouter’s invoiced cost for the same workload.

Flash: $0.14/1M · p50 1.49s · p95 4.66sGLM-5.2: $1.63/1M · p50 4.76s · p95 21.51sNeedle-hit: 1.00 / 0.91 · Tool calls: 1.00 / 1.00Catalog: 692 model IDs · Served deepseek-v4-flash-0424 — a snapshot two months behind 0731, which had been out for a month.

Where it loses points: Cost (8.1/30), update cadence (5/10 — stale Flash snapshot), and GLM tail latency.

Take it if you need its routing/governance features or a hosted model nobody else lists.

Skip it if you assume the same model name means the same model — snapshot pinning is on you here.

5. Fireworks AIdirect provider$1.34/1M on GLM-5.2 — Score 48/100

The most expensive DeepSeek Flash of the run — $0.23 per million blended, 5.3× OpenRouter — with the weakest automatic caching of the platforms that itemize it, and the run’s lowest long-context accuracy: 0.83 needle-hit on GLM-5.2 where every other platform scored 0.91 or better on identical weights.

Flash: $0.23/1M · p50 2.72s · p95 8.28sGLM-5.2: $1.34/1M · p50 5.94s · p95 8.47sNeedle-hit: 1.00 / 0.83 · Tool calls: 0.80 / 1.00Catalog: 23 model IDs · Served deepseek-v4-flash-0731 (current) and glm-5p2.

Where it loses points: Cost (7.1/30), coverage (a 23-model serverless catalog), and reliability (needle 0.83 on GLM, tool-call 0.80 on Flash).

Take it if you are buying its dedicated-deployment and enterprise serving story rather than serverless list price.

Skip it if long-context accuracy matters and you will not verify it yourself first — our 100-request needle test found a real gap.

Groqdirect provider (lpu)$0.89/1M on qwen/qwen3.8-27b — latency appendix

Hosts neither reference model, so no score out of 100. On its own qwen3.8-27b: 0.43s p50 — but $0.89 per million tokens blended for a 27B model, more per token than DeepInfra charges for GLM-5.2, and the only platform where we hit a rate-limit ceiling mid-run (250K tokens/minute on the on-demand tier). The speed is real; so is its price.

p50 0.43s · p95 2.00s · ok 99.1%Needle-hit 0.96 · tool calls 0.92 · catalog 14 models

Cerebrasdirect provider (wafer-scale)$0.37/1M on gpt-oss-120b — latency appendix

The fastest platform in the run — 0.35s p50, 1.1s p95, zero failed requests on gpt-oss-120b — with a two-model catalog. If your model is one of the two, nothing else came close; if it is not, there is nothing to route to.

p50 0.35s · p95 1.10s · ok 100%Needle-hit 0.98 · tool calls 1.00 · catalog 2 models

Managed routers vs direct providers vs BYOK gateways

“OpenRouter alternative” hides three different architectures. The second cut of the same data:

CategoryTested hereWhen it winsWhat the run showed
Managed routersOpenRouter (97), Requesty (57)Many models, one account, no gateway to operate.The category’s spread is the story: same architecture, 3.5× apart on invoiced cost and a two-month gap in model snapshots.
Direct providersDeepInfra (71), Together (61), Fireworks (48); Groq + Cerebras in the appendixOne or two models at steady volume; you want the provider’s own caching, limits, and contracts.Cheapest GLM tokens (DeepInfra) and fastest tokens (Cerebras) both live here — so does the run’s only billing discrepancy.
BYOK / self-hosted gatewaysNot yet — LiteLLM, Portkey, Vercel AI Gateway, Cloudflare AI Gateway, Helicone (all UNVERIFIED)Your own provider contracts must stay in place; the gateway adds routing, budgets, and logs on top.Round 2 of this audit. They route your keys, so “cost” is your providers’ cost — the test that matters is overhead and reliability, and we won’t publish numbers we haven’t run.
  • LiteLLM (Self-hosted gateway (open source)): Routes your own provider keys; zero markup by design. Round-2 candidate.
  • Portkey (BYOK gateway): Governance/observability layer over your keys. Round-2 candidate.
  • Vercel AI Gateway (BYOK gateway): Platform-tied gateway. Round-2 candidate.
  • Cloudflare AI Gateway (BYOK gateway): Edge caching/limits over your keys. Round-2 candidate.
  • Helicone Gateway (BYOK gateway): Observability-first proxy. Round-2 candidate.

Where BenchLM sits in this market (unscored, with the weaknesses stated)

BenchLM’s Radar monitors what these platforms serve — new models, price edits, deprecations — which is how the update-cadence criterion stays measured between audits. Its honest limits: it will not route a single request for you, it does not measure your workload’s latency, and it covers model providers more deeply than gateway catalogs today. If you need routing, you need one of the platforms above; Radar just tells you when the ground moves under it.

What every other comparison skips

1. Nobody reconciles the bill against logged usage — we did, and found +76%

For every request we logged the token counts the platform’s own API returned, then compared them with its billing page. Six of seven platforms reconciled within noise. Together AI billed us for 2,928,000 GLM-5.2 input tokens where its API had reported 1,667,041 — 76% more (output: +24%), while the identical harness reconciled to the token on DeepSeek Flash in the same session. That pattern fits a backend counting difference on this model rather than a rate-card markup; we flag it as observed-not-explained and will update this section if Together clarifies or corrects it.

2. Automatic caching quietly decides the real price

We never requested prompt caching; two platforms applied it anyway. DeepInfra billed 53% of our Flash input at roughly 80% off and half our GLM input likewise — that is the entire reason it wins on cost. Together applied cached tiers too ($0.03/M Flash, $0.26/M GLM). Fireworks cached almost nothing ($0.06 of $1.79 GLM input spend). Identical requests, up to 5× apart on the invoice — rate cards tell you none of this.

3. “Same model” often isn’t: snapshot skew

On run day, OpenRouter, Together, and Fireworks served the current DeepSeek Flash snapshot (0731). Requesty served 0424 — two months older, a full month after 0731 shipped. DeepInfra serves an unpinned ID you cannot version at all. If your evals assume a fixed model, the router chooses otherwise. Model IDs and aliases tracks these mappings.

4. Rate-limit ceilings are real mid-run events

Groq’s on-demand tier capped us at 250K tokens/minute and returned 429s mid-run — the only platform where we hit a documented ceiling with two-way concurrency and 250ms gaps. Fine print until it interrupts production traffic.

5. Serving stacks change model quality, not just speed

The same GLM-5.2 weights scored 0.97 on our 100-question needle test on DeepInfra and 0.83 on Fireworks; tool-call success on Flash ranged from 1.00 (DeepInfra, Requesty) down to 0.80 (Together, Fireworks). Quantization, context handling, and template differences are invisible on a pricing page and very visible in output.

6. Model-launch lag compounds all of it

A platform that trails releases by a month (see Requesty’s snapshot above) silently serves you last quarter’s model at this quarter’s price. That drift is continuous, which is why this audit’s cadence criterion is fed by monitoring rather than a one-day check.

Drift is the part you can automate.

Radar records model additions, removals, and price edits with receipts — the evidence trail behind criteria 3 and 6.

Open Radar

Pick one in 60 seconds

Your situationTakeBecause (from the run)
Many models, one account, lowest measured markupStay on OpenRouter97/100 — fastest p50 both models, $0.044/1M Flash, 418-model catalog.
One big open-weight model at volume, cost is everythingDeepInfra$0.38/1M GLM-5.2 via auto-caching; accept multi-second medians.
Hard latency budget, model flexibility negotiableCerebras (then Groq)0.35s / 0.43s p50 — if your model is in their 2- and 14-model catalogs, and you price the premium.
Your own provider contracts must stay in the loopLiteLLM / Portkey / Cloudflare (BYOK)Untested here (round 2) — architecture fits, but we publish no numbers we haven’t bought.
Deciding which model to route in the first placeBenchLM’s data, then any routeProvider list prices, price-vs-performance, and Radar for when any of it changes.

OpenRouter alternatives FAQ

Is there anything better than OpenRouter?

On a single axis, yes — measured: DeepInfra billed 19% less per GLM-5.2 token, and Cerebras answered in a third of a second. Overall, no platform we paid beat it: OpenRouter scored 97/100 on cost, latency, coverage, reliability, and cadence combined. “Better” depends on which axis your workload actually lives on — the pick-one table above maps them.

Is OpenRouter an LLM gateway?

Functionally it plays the gateway role — one API, provider routing, fallbacks — but it is a hosted marketplace with its own billing, not software you operate. A gateway you run (LiteLLM, a BYOK layer like Cloudflare’s) keeps your provider contracts and data path; OpenRouter replaces them with one account. That architectural difference, not features, is usually the real reason to switch.

Is OpenRouter free?

The account is free and some models carry free variants with tight rate limits, useful for prototyping. Paid usage tracks provider list rates — our invoiced $0.044/1M on DeepSeek Flash matches the model’s list price — with a fee added when you top up credits (documented at about 5.5%; we measured per-token cost, not the fee).

What is the cheapest OpenRouter alternative?

From the invoices: DeepInfra, at $0.38/1M blended on GLM-5.2 — the only platform that beat OpenRouter’s price on either reference model, thanks to automatic caching. On DeepSeek Flash nothing beat OpenRouter’s $0.044/1M. For direct provider list rates model by model, see LLM pricing.

OpenRouter vs Requesty — which is better?

Same architecture, very different receipts: the identical 900-request workload cost $0.92 on OpenRouter and $3.20 on Requesty (3.5×), and Requesty served a DeepSeek Flash snapshot two months older with a 21.5s GLM p95 tail. Requesty’s case is its routing/governance features and a huge nominal catalog (692 IDs) — if you need those, pin your snapshots explicitly.

What is the best free or open-source OpenRouter alternative?

LiteLLM is the standard answer — an open-source, self-hosted proxy with no markup, routing your own provider keys. We have not run the paid workload through it yet (round 2), so it carries an UNVERIFIED label here rather than a score. Note “free gateway” still means paying providers underneath; free routing is not free tokens.

Can I use these alternatives with Janitor AI or on iOS apps?

Generally yes: anything that accepts a custom OpenAI-compatible base URL and API key — which covers Janitor AI’s proxy setting and most iOS chat clients — can point at any platform tested here. Check the platform allows your use case in its terms, and mind the per-minute rate limits (Groq’s 250K tokens/minute ceiling is the kind that bites busy roleplay traffic).

About this audit

Written by Glevd, who runs BenchLM — the model-benchmark and pricing tracker this site is built on. The test was designed, paid for, and run by us on 2026-09-01; no platform was notified in advance, none supplied credits (Cerebras’ standard sign-up promo happened to cover its $0.65), and none saw this page before publication. Corrections are welcome and will be applied in place with a note.

Methodology, verifiable

Harness: plain OpenAI-compatible chat/completions calls, concurrency 2, 250ms gaps, one retry on 429, 120s timeout, no streaming — p50/p95 are full-completion latencies, identical in shape for every platform (comparable, not TTFT). Workload per model per platform: 300 short-chat + 100 long-context (~20K tokens with a needle question verifying context receipt) + 50 tool-call requests. All seven platforms ran from the same machine, the same day (2026-09-01, 14:53–18:24 UTC), on freshly loaded accounts.

Cost is invoiced dollars from each billing page divided by API-reported tokens — never rate-card transcription. Anything not measured is labeled UNVERIFIED. The per-request CSVs, per-platform summaries, the invoice ledger, and the frozen prompt set are committed in the BenchLM repository (docs/seo/openrouter-pilot-data/ and scripts/seo/openrouter-pilot/), so every number on this page can be traced to a logged request.

This page refreshes in place: the next paid run re-scores everything and bumps the verified date, and Radar tracks catalog and price movement between runs. The winner badge, when a platform requests one, follows the audit result, is never sold, and changes when the result changes.

Get the next audit

Quarterly re-runs of this test, plus price and model-catalog changes as Radar catches them.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.