Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Platform auditRun 2026-09-01

OpenRouter alternatives in 2026: seven platforms, one paid workload, scored from the invoices

By GlevdPublished September 1, 2026Run and prices verified

We loaded credits on seven platforms and sent each the same workload: 300 short chat requests, 100 long-context requests of about 20,000 tokens with a needle question that verifies the model received the context, and 50 tool-call requests, per reference model. Cost is the invoiced amount divided by the tokens each platform’s API reported; latency, success rate, needle recall, and tool-call success come from the per-request logs. The rubric and its weights were fixed before the run. The logs, the invoice ledger, and the prompt set are committed alongside this site.

Scored from the invoices, DeepInfra is the lowest-cost alternative ($0.38 per million blended tokens on GLM-5.2, 19% below OpenRouter) and Cerebras the fastest (0.35-second median). Across all five criteria, OpenRouter scored highest, 97 out of 100: the fastest median on both reference models and the lowest invoiced cost on DeepSeek Flash. No alternative beat it on more than one criterion.

Independence

BenchLM sells model monitoring (Radar) and benchmark data, not a gateway, a router, or inference, so nothing of ours is scored. As of 2026-09-01 BenchLM has no affiliate or referral relationship with any platform on this page. If that changes, this note changes first, and affiliate status will not affect inclusion, order, or verdicts (policy). The $13.07 of test spend was ours; no platform was notified in advance or saw this page before publication.

Results

OpenRouter scored highest on this workload; the alternatives each win on one criterion

Five platforms host both reference models and receive a score out of 100. Groq and Cerebras host neither, so their own-model results are reported without a score rather than compared on unequal terms.

PlatformTypeScore /100Cost /30Latency /25Coverage /20Reliability /15Cadence /10
1OpenRouterHighestManaged router9727.2252014.610
2DeepInfraDirect provider7126.57.114.714.97.5
3Together AIDirect provider619.410.21714.510
4RequestyManaged router578.19.42014.75
5Fireworks AIDirect provider487.1710.613.110
AppendixGroqDirect provider (LPU)Not scored: hosts neither reference model. On qwen/qwen3.8-27b: 0.43s median, 2.00s p95, $0.89 per million blended, 99.1% of requests succeeded.
AppendixCerebrasDirect provider (wafer-scale)Not scored: hosts neither reference model. On gpt-oss-120b: 0.35s median, 1.10s p95, $0.37 per million blended, 100% of requests succeeded.

The result is not the one the search term implies. OpenRouter kept the highest score on the criteria that carry the most weight, cost and latency, and no alternative was both cheaper and faster on either model. DeepInfra leads on cost and pays for it in latency; Cerebras leads on latency and offers two models. The component columns show each trade.

How the score is computed
CriterionWeightMeasured by
Effective cost30Invoiced dollars divided by API-reported tokens, per reference model. The cheapest platform on each model sets the bar; the others score in proportion.
Latency25Median (70%) and 95th-percentile (30%) full-completion latency, fastest platform per model as the bar. No streaming, so the figures are comparable across platforms but are not time to first token.
Model coverage20Half for serving both reference models; half for catalog breadth from the platform’s own model list on the run date, with 400 or more IDs scoring full marks.
Reliability15Success rate (below 95% scores zero on that half), needle recall on the 20,000-token contexts, and tool-call success.
Update cadence10Which DeepSeek Flash snapshot the platform served on the run date (current, unpinned, or stale) and whether GLM-5.2 was available.

Reference models: DeepSeek V4 Flash, the low-cost open-weight workhorse, and GLM-5.2, the one frontier-class open model all five scored platforms host. Weights sum to 100 and were set before any request was sent.

Cost

Invoiced cost per million tokens

Rate cards state list prices. These figures are what each billing page charged for the run, divided by the tokens the platform’s API reported back, so they include markups, caching discounts, and any counting differences. Sorted by GLM-5.2 cost.

PlatformDeepSeek Flash, invoicedFlash $ per 1MGLM-5.2, invoicedGLM $ per 1MWhole run
DeepInfra$0.10 / 1.73M tok$0.058$0.68 / 1.80M tok$0.38$0.78
OpenRouter$0.08 / 1.82M tok$0.044$0.83 / 1.80M tok$0.46$0.92
Fireworks AI$0.42 / 1.80M tok$0.23$2.39 / 1.79M tok$1.34$2.81
Requesty$0.25 / 1.73M tok$0.14$2.95 / 1.81M tok$1.63$3.20
Together AI$0.20 / 1.80M tok$0.11$3.00 / 1.80M tok$1.67$3.20
GroqOwn model (qwen/qwen3.8-27b): $1.51 / 1.70M tokens, $0.89 per million.$1.51
CerebrasOwn model (gpt-oss-120b): $0.65 / 1.77M tokens, $0.37 per million; covered in full by sign-up promotional credit, $0.00 out of pocket.$0.65

OpenRouter’s Flash tokens cost 5.3 times less than Fireworks’ for the identical requests. DeepInfra’s GLM-5.2 figure, $0.38 against a $0.49 list price for input, reflects automatic caching that billed half our input at about 80% off. Together’s GLM figure includes the billing discrepancy described under findings. Output shares differed by a few percentage points between snapshots, which moves blended figures slightly and reorders nothing. The platform totals sum to the $13.07 invoiced across all seven billing pages.

Platform notes

Platform by platform

1

OpenRouter

Managed routerScore 97/100
Flash, per 1M
$0.044 · p50 0.76s · p95 2.59s
GLM-5.2, per 1M
$0.46 · p50 1.20s · p95 3.69s
Needle · tool calls
0.94 / 0.93 · 0.98 / 1.00
Catalog
418 model IDs

Fastest median on both reference models and the lowest invoiced cost on Flash; no alternative beat it on both.

OpenRouter posted the fastest median latency on both reference models, the lowest invoiced cost per million tokens on DeepSeek Flash, and GLM-5.2 within 23% of the cheapest platform. No alternative in this run was both cheaper and faster on either model.

Served deepseek-v4-flash-0731, the current snapshot, with GLM-5.2 available on the run date.

Loses points on
GLM-5.2 cost: DeepInfra’s automatic caching undercut it by 19% per blended token. The credit top-up fee (documented at about 5.5%; not measured here) sits outside the per-token figures.
Suits
Workloads where one account, the widest live catalog we measured (418 model IDs), and invoiced pricing that tracked provider list rates matter most.
Does not suit
Cases where the workload is one large open-weight model at steady volume; DeepInfra billed 19% less per GLM-5.2 token.
2

DeepInfra

Direct providerScore 71/100
Flash, per 1M
$0.058 · p50 2.42s · p95 6.93s
GLM-5.2, per 1M
$0.38 · p50 5.29s · p95 13.82s
Needle · tool calls
1.00 / 0.97 · 1.00 / 1.00
Catalog
189 model IDs

Cheapest GLM-5.2 tokens in the run through automatic caching; the slowest medians of the five scored platforms.

DeepInfra billed the lowest cost per GLM-5.2 token in the run, $0.38 per million blended, because it cached 53% of our Flash input and about half of our GLM input at roughly 80% off with no caching parameter in the request. Median latency was 2.4 to 5.3 seconds, the slowest of the five scored platforms on GLM-5.2. All 900 requests succeeded.

Serves an unpinned DeepSeek-V4-Flash ID: current weights, with no snapshot to select.

Loses points on
Latency (7.1 of 25, with five-second GLM medians) and an unpinned Flash model ID.
Suits
Workloads where cost per token on large open-weight models dominates and multi-second medians are acceptable.
Does not suit
Cases where sub-second responses or a pinned model snapshot are required.
3

Together AI

Direct providerScore 61/100
Flash, per 1M
$0.11 · p50 2.55s · p95 10.09s
GLM-5.2, per 1M
$1.67 · p50 2.53s · p95 5.66s
Needle · tool calls
0.99 / 0.93 · 0.80 / 1.00
Catalog
278 model IDs

Every request succeeded and its GLM p95 was second only to OpenRouter; its GLM bill charged 76% more input tokens than its API reported.

All 900 requests succeeded, and Together’s GLM-5.2 p95 was second only to OpenRouter. Its invoice charged 2,928,000 GLM-5.2 input tokens against 1,667,041 reported by its own API responses, 76% more, while the same method reconciled to the token on DeepSeek Flash. We record this as observed and unexplained; the findings section has the figures, and we have invited a correction.

Served DeepSeek-V4-Flash-0731 (current) and GLM-5.2; GLM-5.3 was already listed in its catalog on the run date.

Loses points on
Effective cost (9.4 of 30): the billed-token inflation on GLM-5.2 lands directly in the invoiced cost per million.
Suits
Workloads where a large direct catalog with solid GLM latency is wanted and invoices are reconciled against logged usage.
Does not suit
Cases where billing cannot be verified; until the GLM discrepancy is explained, budgets should sit above rate-card arithmetic.
4

Requesty

Managed routerScore 57/100
Flash, per 1M
$0.14 · p50 1.49s · p95 4.66s
GLM-5.2, per 1M
$1.63 · p50 4.76s · p95 21.51s
Needle · tool calls
1.00 / 0.91 · 1.00 / 1.00
Catalog
692 model IDs

The largest catalog by ID count; a two-month-old Flash snapshot and the run’s longest GLM tail at 3.5× OpenRouter’s cost.

Requesty lists 692 model IDs, inflated by per-region duplicates, and returned every needle on Flash. It served deepseek-v4-flash-0424, a snapshot two months old, a month after 0731 shipped. Its GLM-5.2 p95 was 21.5 seconds, the longest tail of the run, and the same workload cost 3.5 times what OpenRouter invoiced.

Served deepseek-v4-flash-0424, two months behind 0731, which had been available for a month.

Loses points on
Cost (8.1 of 30), update cadence (5 of 10, for the stale Flash snapshot), and GLM tail latency.
Suits
Workloads where its routing or governance features are needed, or it lists a hosted model no other platform does.
Does not suit
Cases where a model name is assumed to identify a snapshot; versions must be pinned explicitly here.
5

Fireworks AI

Direct providerScore 48/100
Flash, per 1M
$0.23 · p50 2.72s · p95 8.28s
GLM-5.2, per 1M
$1.34 · p50 5.94s · p95 8.47s
Needle · tool calls
1.00 / 0.83 · 0.80 / 1.00
Catalog
23 model IDs

The highest invoiced Flash cost of the run at 5.3× OpenRouter, the least caching, and the lowest long-context accuracy.

Fireworks invoiced the highest cost per DeepSeek Flash token in the run, $0.23 per million blended, 5.3 times OpenRouter. Among the platforms that itemise caching it cached the least. Its GLM-5.2 needle-hit rate was 0.83; every other platform scored 0.91 or better on the same weights.

Served deepseek-v4-flash-0731 (current) and glm-5p2.

Loses points on
Cost (7.1 of 30), coverage (a 23-model serverless catalog), and reliability (needle 0.83 on GLM, tool-call 0.80 on Flash).
Suits
Workloads where dedicated deployments and enterprise serving are the purchase, rather than serverless list price.
Does not suit
Cases where long-context accuracy matters and will not be verified first; our 100-request needle test found a measurable gap.
Appendix

Groq

Direct provider (LPU)Not scored
Model tested
qwen/qwen3.8-27b
Per 1M · latency
$0.89 · p50 0.43s · p95 2.00s
Success · needle · tool
99.1% · 0.96 · 0.92
Catalog
14 models

Groq hosts neither reference model, so it receives no score. On qwen3.8-27b it returned a 0.43-second median at $0.89 per million tokens blended, more per token than DeepInfra charged for GLM-5.2, and it was the only platform to return rate-limit errors mid-run (250,000 tokens per minute on the on-demand tier).

Appendix

Cerebras

Direct provider (wafer-scale)Not scored
Model tested
gpt-oss-120b
Per 1M · latency
$0.37 · p50 0.35s · p95 1.10s
Success · needle · tool
100% · 0.98 · 1.00
Catalog
2 models

Cerebras hosts neither reference model, so it receives no score. On gpt-oss-120b it was the fastest platform in the run: a 0.35-second median, 1.1-second p95, and no failed requests, on a two-model catalog.

Architectures

Three architectures behind one search term

“OpenRouter alternative” covers managed routers, direct providers, and gateways that route keys you already hold. The same results, cut by architecture:

ArchitectureTested hereWhere it fitsWhat the run showed
Managed routersOpenRouter (97), Requesty (57)Many models behind one account, with no gateway to operate.The same architecture spanned a 3.5× difference in invoiced cost and a two-month gap in model snapshots.
Direct providersDeepInfra (71), Together (61), Fireworks (48); Groq and Cerebras in the appendixOne or two models at steady volume, using the provider’s own caching, limits, and contracts.The cheapest GLM-5.2 tokens (DeepInfra) and the fastest responses (Cerebras) are both here; so is the run’s one billing discrepancy.
Gateways over your own keysNot yet tested: LiteLLM, Portkey, Vercel AI Gateway, Cloudflare AI Gateway, Helicone (unverified)Provider contracts stay in place; the gateway adds routing, budgets, and logs.Planned for the second round. Because they route your keys, cost is your providers’ cost; the measurements that matter are overhead and reliability, and we do not publish figures we have not run.
  • LiteLLM (Self-hosted gateway (open source)): Routes your own provider keys; zero markup by design. Round-2 candidate.
  • Portkey (BYOK gateway): Governance/observability layer over your keys. Round-2 candidate.
  • Vercel AI Gateway (BYOK gateway): Platform-tied gateway. Round-2 candidate.
  • Cloudflare AI Gateway (BYOK gateway): Edge caching/limits over your keys. Round-2 candidate.
  • Helicone Gateway (BYOK gateway): Observability-first proxy. Round-2 candidate.
Findings

Six things the invoices showed that rate cards do not

1

Billed tokens did not always match the tokens the API reported

For every request we logged the token counts each platform returned, then compared the sum with its billing page. Six of seven platforms reconciled within noise. Together AI billed 2,928,000 GLM-5.2 input tokens where its API responses summed to 1,667,041, 76% more (output: 24% more), while the same harness reconciled to the token on DeepSeek Flash in the same session. The pattern is consistent with a counting difference on this model rather than a rate-card markup. We record it as observed and unexplained, and will update this section if Together clarifies or corrects it.

2

Automatic caching, not the rate card, set the effective price

No request asked for prompt caching. DeepInfra applied it anyway, billing 53% of our Flash input and about half of our GLM input at roughly 80% off, which is the whole reason it leads on cost. Together applied cached tiers as well ($0.03 per million on Flash, $0.26 on GLM). Fireworks cached almost nothing: $0.06 of $1.79 in GLM input spend. Identical requests landed up to five times apart on the invoice.

3

The same model name resolved to different snapshots

On the run date OpenRouter, Together, and Fireworks served the current DeepSeek Flash snapshot (0731). Requesty served 0424, two months older, a month after 0731 shipped. DeepInfra serves an unpinned ID with no version to select. An evaluation that assumes a fixed model is at the mercy of the platform’s alias table.

4

One platform hit its documented rate limit during the run

Groq’s on-demand tier returned 429 responses once the run crossed 250,000 tokens per minute, with two concurrent requests and 250 ms gaps. No other platform returned a rate-limit error at this load.

5

Serving stacks changed output quality on identical weights

GLM-5.2 answered the 100-question needle test at 0.97 on DeepInfra and 0.83 on Fireworks. Tool-call success on DeepSeek Flash ranged from 1.00 (DeepInfra, Requesty) to 0.80 (Together, Fireworks). Quantisation, context handling, and prompt templates differ between platforms and do not appear on a pricing page.

6

Model-launch lag compounds every other difference

A platform that trails releases by a month serves last quarter’s model at this quarter’s price. The update-cadence criterion is scored from what each platform served on the run date; between runs, catalog and price changes are recorded by monitoring rather than by another one-day check.

BenchLM Radar

Catalog and price changes between runs are recorded, not re-checked by hand.

Radar records model additions, removals, and price edits across providers with a dated source for each. It is the evidence behind the update-cadence criterion between audits.

Open Radar
Choosing

Which platform, by situation

SituationPlatformBasis in the run
Many models, one account, lowest measured markupOpenRouter97 of 100; fastest median on both models; $0.044 per million on Flash; 418 model IDs.
One large open-weight model at volume, cost firstDeepInfra$0.38 per million on GLM-5.2 through automatic caching; medians of several seconds.
A hard latency budget, flexibility on the modelCerebras, then Groq0.35 s and 0.43 s medians, on catalogs of two and fourteen models, at a per-token premium.
Provider contracts that must stay in the loopLiteLLM, Portkey, or Cloudflare AI GatewayArchitecturally the fit; not yet measured here (second round).
Deciding which model to route in the first placeBenchLM’s data, then any routeProvider list prices, price against performance, and Radar for changes.
Prior work

Existing comparisons of these platforms

Nine pages ranked for “openrouter alternatives” on Google (United States) on the run date. Seven are published by companies that sell a gateway or inference service, and each of those lists its own product first or names itself the best alternative. One, AIMultiple’s, ran a latency benchmark. None publishes a paid workload, per-request logs, or invoiced prices, which is the gap this audit is designed to fill. They are listed so a reader can weigh each source knowing who wrote it.

SourceDatedMethod and placement
DigitalOcean, 9 OpenRouter Alternatives for Multi-Model AI2026-08-17Feature comparison; lists its own platform first.
Reddit, r/RooCode, Other OpenRouter-like API providers?2025User discussion; no product to place.
haimaker.ai, 8 Unified AI APIs ComparedUndatedFeature comparison; names itself the best alternative.
TrueFoundry, Best OpenRouter Alternatives for Production2026-08-07Feature comparison; names itself the leading alternative.
AIMLAPI, The Best OpenRouter Alternatives in 20262026-05-06Feature comparison; lists its own API.
FreeRouter (GitHub), Free, self-hosted AI model routerUndatedProject README.
AIMultiple, AI Gateways for OpenAI2026-05-13Latency benchmark across five platforms; no invoiced pricing.
Merge.dev, 4 OpenRouter alternativesUndatedFeature comparison; lists its own gateway first.
Morph, 11 Providers ComparedUndatedFeature comparison; lists itself. Documents OpenRouter’s 5.5% credit fee.
Questions

Questions people ask about OpenRouter alternatives

Is there anything better than OpenRouter?

On a single criterion, yes, and by a measurable margin: DeepInfra billed 19% less per GLM-5.2 token, and Cerebras answered in a third of a second. Across the five criteria combined, no platform we paid scored higher than OpenRouter’s 97 out of 100. The decision table above maps each situation to the platform that wins on the criterion that matters for it.

Is OpenRouter an LLM gateway?

It performs the gateway role, one API with provider routing and fallbacks, but it is a hosted marketplace with its own billing rather than software you operate. A gateway you run, such as LiteLLM or a bring-your-own-key layer, keeps your provider contracts and data path in place; OpenRouter replaces them with one account. That architectural difference is usually the real reason to switch, not a feature gap.

Is OpenRouter free?

The account is free, and some models carry free variants with tight rate limits that suit prototyping. Paid usage tracked provider list rates in our run: the invoiced $0.044 per million tokens on DeepSeek Flash matches the model’s list price. A fee is added when credits are purchased (documented at about 5.5%); we measured per-token cost, not the fee.

What is the cheapest OpenRouter alternative?

From the invoices, DeepInfra: $0.38 per million blended tokens on GLM-5.2, the only platform that billed less than OpenRouter on either reference model, through automatic caching. On DeepSeek Flash nothing beat OpenRouter’s $0.044. Direct provider list rates by model are on the LLM pricing page.

OpenRouter vs Requesty: which is better?

Same architecture, different receipts. The identical 900-request workload cost $0.92 on OpenRouter and $3.20 on Requesty, and Requesty served a DeepSeek Flash snapshot two months older with a 21.5-second GLM p95. Requesty’s case rests on its routing and governance features and a large nominal catalog (692 IDs); if those matter, pin model snapshots explicitly.

What is the best free or open-source OpenRouter alternative?

LiteLLM is the usual answer: an open-source, self-hosted proxy with no markup that routes your own provider keys. We have not run the paid workload through it, so it carries an unverified label here rather than a score. A free gateway still bills you for the providers underneath it; free routing is not free tokens.

Can these alternatives be used with Janitor AI or on iOS apps?

Generally, yes. Any client that accepts a custom OpenAI-compatible base URL and API key, which covers Janitor AI’s proxy setting and most iOS chat clients, can point at any platform tested here. Check that the platform’s terms allow the use case, and note per-minute rate limits: Groq’s 250,000 tokens-per-minute ceiling is the kind that interrupts busy roleplay traffic.

Method

Method and limits

Every platform was called through a plain OpenAI-compatible chat-completions endpoint from the same machine on the same day, 2026-09-01, between 14:53 and 18:24 UTC, on freshly loaded accounts. The workload per reference model was 300 short-chat requests, 100 long-context requests of about 20,000 tokens carrying a needle question, and 50 tool-call requests; the five scored platforms ran it on both reference models (900 requests each), and Groq and Cerebras ran it once on their own flagship (450 each), for 5,400 in total.

Latency is full-completion time without streaming, so it is comparable across platforms but is not time to first token. Cost is invoiced dollars from each billing page divided by API-reported tokens; nothing is transcribed from a rate card. The needle test scores whether the answer contained the passphrase planted in the context. Anything not measured is labelled unverified.

Limits: one day, one region, one account tier per platform, two reference models, and no streaming. Groq and Cerebras are reported without a score because a score would compare different models. The Together AI discrepancy is a comparison of two numbers the platform itself produced, not an audit of its metering. The per-request logs, per-platform summaries, invoice ledger, and frozen prompt set are committed in the BenchLM repository under docs/seo/openrouter-pilot-data/ and scripts/seo/openrouter-pilot/.

The page is refreshed in place: the next paid run re-scores every platform and advances the verified date, and Radar records catalog and price changes between runs. A platform that wants to cite its result may request a badge; the badge follows the result, is not sold, and changes when the result changes.

The next run, when it happens

Quarterly re-runs of this audit, and the catalog and price changes Radar records in between.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.