Cerebras’s public catalog puts GPT-OSS 120B near 3,000 tokens per second. Groq’s production table puts the same weight near 500. On the one public text model both sell, the wafer is roughly six times the LPU.
That is the first paragraph. It is not the article.
The article is which weights you can actually run at that speed, and whether your system is waiting on tokens at all. We did not send a bake-off prompt to either API. Every speed and price below is the vendor’s public table, retrieved 18 August 2026, except where we name Artificial Analysis as the source. Same model, same day, two hosts is the only fair race. Mixed catalogs are not a race.
If you came here from the Groq essay, the identity question is already answered. This piece is the buy.
The overlap set is the only race
Strip the homepages and one shared production SKU remains.
| Host | Model | Listed speed | Input / output per 1M | Context (paid) |
|---|---|---|---|---|
| Cerebras | gpt-oss-120b |
~3,000 tok/s | $0.35 / $0.75 | 131k |
| Groq | openai/gpt-oss-120b |
~500 tok/s | $0.15 / $0.60 | 131k |
| Cerebras | gemma-4-31b |
~1,850 tok/s | $0.99 / $1.49 | 131k |
| Groq | openai/gpt-oss-20b |
~1,000 tok/s | $0.075 / $0.30 | 131k |
Same public weight. Six times the listed speed. A higher sticker. Not a run we performed.
Cerebras is faster and more expensive on the overlap. Groq is the cheaper 20B sprinter. Neither row is a Claude, a Kimi, or a DeepSeek. If your shortlist starts with those names, this table is a detour.
Cerebras’s own September 2025 comparison claimed up to 6× Groq on identical models and cited Artificial Analysis figures in that range. Treat that post as vendor advocacy. The current public tables still rhyme with it on GPT-OSS 120B. They do not let us extend the multiple to Qwen, Llama 3.3 70B, or any other weight that one of them no longer lists.
Two models on the public door
Cerebras’s model catalog is blunt. Public endpoints: gpt-oss-120b and gemma-4-31b. Dedicated endpoints and partners (OpenRouter, Hugging Face, Vercel, AWS Marketplace) carry more families. The shared door does not.
That is the wafer constraint in one sentence. Speed is real. The menu is the product.
Groq at least still shows a preview Qwen and an enterprise MiniMax. Cerebras’s public list is two open-weight IDs and a dedicated-endpoint footnote. If you need a model that is only on dedicated hardware, you are in a sales conversation, not a developer-tier price table.
We are not going to invent the rest of the menu from old Llama 3.3 blog charts. Those posts are still on the internet. They are not today’s catalog.
The wafer just got a frontier guest
On 13 August 2026 OpenAI previewed Ultrafast, a Cerebras-powered service tier for GPT-5.6 Sol. OpenAI and Cerebras say up to 750 output tokens per second and up to 14× Standard processing. It is waitlist-gated. There is no public Ultrafast price and no generally available model ID you can paste into a billing forecast.
This is the first time the wafer story includes a closed frontier weight people actually shortlist. It is also still a preview. Sol’s standard API remains $5 / $30. Whether Ultrafast is a premium on that sticker, a limited free taste, or a capacity experiment is unpublished.
So the honest status is: Cerebras can now be the hidden host behind an OpenAI SKU, and you cannot yet buy that SKU as a normal line item. Do not write it into a production cost model. Do write it into the next-quarter watch list. A generally available Sol-on-Cerebras price would change the buy. A quiet preview that stays gated would not.
Tokens per second is the wrong unit for most agents
A 400-token reply at 3,000 tokens per second is a blink. The same agent with three tool calls and a reasoner that thinks before it speaks is a different machine.
On our 17 August Artificial Analysis-derived runtime row, Claude Fable 5 shows a 139.55-second time to first answer. Faster decoding does nothing to that wait before the first chunk. The latency-tax post is the mechanism: the critical path is the longest dependency chain, not the fastest isolated completion.
Buy Cerebras when two things are true at once. The model you need is on the wafer, or on a dedicated endpoint you have actually contracted. And generation, not tools or silent reasoning, is what the user is waiting on. Voice, live coding autocomplete, tight classification, a UI that must stream. Those jobs feel the 3,000-token-per-second number.
Skip Cerebras when the weight is missing, or when a GPU host at a few hundred tokens per second is waiting on search, a repo tool, or a human approval. You do not need a racecar for a traffic jam.
Price the overlap honestly. On GPT-OSS 120B, Groq is $0.15 / $0.60 and slower. Cerebras is $0.35 / $0.75 and faster. If the quality is the same weight, you are buying time. Run the cost calculator on your actual input/output mix before you assume the cheaper sticker wins.
What would change the buy
Two rows would rewrite this page.
A generally available frontier closed model on Cerebras, with a public price. Sol Ultrafast is the preview of that row. It is not the row.
A collapse of the tokens-per-second gap on ordinary GPU hosts. Blackwell already ate part of Groq’s original demo. If GPU shops print 1,000-plus tokens per second on the same 120B at Groq’s price, the wafer becomes a specialty, not a default.
Until then, treat Cerebras as the fast, expensive door with two public keys and a waitlist for the interesting one. The speed crown is real on the overlap. It is worthless unless the model is on the wafer.
Reader questions
Frequently asked questions
01Is Cerebras faster than Groq?
On the one public text model both list, yes, by a wide vendor-claimed margin. Cerebras lists GPT-OSS 120B at about 3,000 tokens per second. Groq lists about 500. Those are their tables, not a same-day run we performed. A model only one of them serves is not a race.
02Which models does Cerebras run?
The public catalog on 18 August 2026 lists two endpoints: gpt-oss-120b and gemma-4-31b. Dedicated endpoints and partners carry more families. OpenAI previewed GPT-5.6 Sol Ultrafast on Cerebras hardware on 13 August, waitlist-gated, with no public Ultrafast price.
03How much does Cerebras inference cost?
Cerebras’s developer table lists GPT-OSS 120B at $0.35 input and $0.75 output per million tokens, and Gemma 4 31B at $0.99 / $1.49. Groq’s GPT-OSS 120B is $0.15 / $0.60. Cerebras is the faster, more expensive host on that overlap. Ultrafast Sol pricing is unpublished.
04When is GPU inference good enough?
When the model you need is not on the wafer, or when tools, retries, and reasoning time dominate the critical path. A Blackwell host at a few hundred tokens per second is fine if the agent is waiting on a search call. Buy Cerebras when generation is the wait and the weight is actually served.
05Does tokens-per-second matter for agents?
Only when generation sits on the critical path. An agent with three tool calls and a long silent reasoner is not a streaming demo. Fable 5’s first-answer latency on our 17 August Artificial Analysis-derived row is 139.55 seconds. No wafer fixes more than two minutes of hidden thinking.
Source ledger
External sources linked in this article
- 01GPT-OSS 120Binference-docs.cerebras.ai
- 02Ultrafastopenai.com
Continue with live BenchLM data
Share or save
