Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Groq is not a model

After Nvidia licensed the chip and hired the team, what is left is a fast, narrow catalog. Most people searching “groq” still think they are buying a brain.

Hand-inked Groq lightning bolt on bone ledger paper, cobalt blot at the tip
Published
Last reviewed
Data as of
Reading time
11 min
External sources
3
Tags: groq, inference, pricing, speedData and scoring methodology
In this article5 sections

On 16 August 2026 Groq shut down llama-3.1-8b-instant for free and developer-tier traffic. That was the cheap, famous sprinter. The deprecation note tells you to move to openai/gpt-oss-20b. Two days later, the public production text catalog is two GPT-OSS sizes, Whisper, and a Compound system.

Groq is not a model. It never was. People still type the word as if it were one, or as if it were Grok. Both readings are wrong. Groq sells throughput on a short list of someone else’s weights.

We did not run the Groq API for this article. Speed numbers below are Groq’s published production rates, dated 18 August 2026. We will not crown a “fastest inference company” from another site’s model-level medians, including our own speed table, which currently follows Artificial Analysis rather than Groq or Cerebras endpoints.

People are still searching the wrong noun

“Groq” is a 47,000-a-month query. The results mix a hardware cloud, a Grok typo, and a static provider table. The useful sentence is narrower than any of those pages.

Groq is an inference company. It pioneered an LPU, then an LPX story that sits next to Nvidia GPUs. It does not train a closed frontier model you can call groq-5. If you want Grok 4.5, you want SpaceXAI. If you want 500 to 1,000 tokens per second on an open-weight chat model, you want a Groq SKU.

Get that split wrong and you will evaluate the wrong object. You will compare Groq to Claude. You should compare Groq’s GPT-OSS 120B route to Together’s, Fireworks’, Cerebras’s, or your own GPU box.

The catalog after the Llama shutdown

The current supported-models page is the receipt. Production text models are not a zoo.

Table 1
Model ID Groq’s listed speed Price (in / out per 1M) Context Status
openai/gpt-oss-20b ~1,000 tok/s $0.075 / $0.30 131,072 Production
openai/gpt-oss-120b ~500 tok/s $0.15 / $0.60 131,072 Production
groq/compound ~450 tok/s unlisted on the model table 131,072 Production system
whisper-large-v3 n/a $0.111 / hour n/a Production audio
qwen/qwen3.6-27b ~500 tok/s $0.60 / $3.00 131,072 Preview
minimaxai/minimax-m2.7 ~260 tok/s sales-only 196,608 Preview, enterprise
llama-3.1-8b-instant Shut down 16 Aug 2026 for free and developer tiers

Three cards: Llama 3.1 8B Instant marked shut down, GPT-OSS 20B remaining at about 1,000 tokens per second, GPT-OSS 120B remaining at about 500.

Two production text models remain. The rest of the old demo is a deprecation notice.

Empty cells stay empty. Compound has no token price on that table. MiniMax has no public rate. Llama 3.1 8B Instant is gone unless you are an enterprise customer on a committed-spend contract.

That last row is the product lesson. Groq will deprecate a beloved SKU to keep the fleet on the models the chip wants to run. A narrow catalog is not a temporary embarrassment. It is the constraint you are buying.

Qwen 3.6 27B is the price trap on the preview list: $0.60 in and $3.00 out. If you price it off the input rate, long completions will surprise you.

Speed we will not crown

Groq’s own production numbers are already a story. One thousand tokens per second on a 20B open-weight model is a real product. Five hundred on a 120B is still faster than most GPU hosts advertise for that class.

We are not going to turn those rows into “Groq is the fastest.” We did not send the prompts. Groq’s listed rates are vendor claims with a date. Artificial Analysis’s Groq provider page is a third-party snapshot with a different methodology. Our llm-speed surface is currently a model-level AA feed, not an endpoint bake-off.

The only fair race is same model, same prompt, same day, two hosts. The overlap that matters right now is GPT-OSS 120B against Cerebras and the GPU shops. That comparison belongs in the Cerebras piece, because Cerebras is the other specialist and the public overlap set is tiny.

After Nvidia

In late December 2025 Nvidia paid about $20 billion to license Groq’s inference architecture and hire most of the team, including the founder. It was structured as a license and talent deal, not a clean acquisition, which is why Groq still exists as a company. TechCrunch reported in May that the remaining firm was raising to grow the inference cloud. Groq’s homepage now leads with a $350 million Series A and an LPX story that sits alongside Nvidia GPUs.

Read that as a catalog company, not as an independent chip vendor you can bet a five-year architecture on. Nvidia wanted the low-latency design. Groq kept a cloud. Your lock-in is the endpoint and the short model list, not a secret weight.

If your system only works because Groq has a particular 8B Instant SKU, the 16 August shutdown already explained the risk.

When a SKU is the right buy

Buy Groq when generation is the critical path and the model you need is on that table. Voice, routing, classification, tight agent steps, cheap structured extraction. GPT-OSS 20B at $0.075 / $0.30 and about 1,000 tokens per second is a serious volume product for that work.

Do not buy Groq when the job is a frontier closed model, a Qwen family Groq does not produce, or a Kimi / DeepSeek / Claude weight that is not on the page. Speed without the weight is a demo. A 500-token-per-second 120B does not make up for the model you actually needed.

Price the whole task, not the sticker. Use the cost calculator. If the step is inside an agent, read the latency-tax note before you celebrate tokens per second. A reasoner that thinks for a minute before the first token does not care that Groq can stream a 20B like a firehose.

The number to watch next is whether the next catalog add is a frontier weight or another 8B/20B sprinter. The Llama shutdown already told you which way the fleet prefers to move.

Reader questions

Frequently asked questions

01What is Groq?

Groq is an inference cloud, not a foundation-model lab. It serves other people’s weights on its own chips. As of 18 August 2026 the public production text models are OpenAI’s GPT-OSS 20B and 120B, plus Whisper and Groq Compound. It does not sell a Groq-branded frontier brain.

02Is Groq the same as Grok?

No. Groq is the inference company. Grok is SpaceXAI’s model family. The names collide in search and almost nowhere else. If you want Grok 4.5, you are buying an xAI API, not a Groq endpoint. If you want 500 to 1,000 tokens per second on an open-weight model, you are in the Groq catalog.

03Which models does Groq serve?

Production text models on 18 August 2026 are openai/gpt-oss-20b at about 1,000 tokens per second and openai/gpt-oss-120b at about 500, plus whisper-large-v3 and the groq/compound systems. Qwen 3.6 27B and MiniMax M2.7 are preview or enterprise. Llama 3.1 8B Instant shut down for free and developer tiers on 16 August 2026.

04How fast is Groq compared with a GPU API?

Groq’s own production table lists about 1,000 tokens per second on GPT-OSS 20B and about 500 on GPT-OSS 120B. Those are vendor rates, not a run we performed. They are faster than a typical GPU host on the same small open-weight class. They say nothing about a model Groq does not serve.

05Is Groq still independent after Nvidia?

The remaining company still sells GroqCloud. In December 2025 Nvidia licensed Groq’s inference architecture and hired most of the team in a roughly $20 billion transaction that was not a full acquisition. Groq later raised to grow the neocloud. Build against the catalog, not against an independent chip roadmap.

Source ledger

External sources linked in this article

3
  1. 01deprecation note
  2. 02supported-models page
  3. 03TechCrunch reported

Share or save

Share on XShare on LinkedIn

Keep reading

All research

Choose the right model before an expensive mistake. One weekly recommendation: what to choose, what costs less, and what is not worth switching to.

Read a sample issue

Join 2,000+ readers.