Fastest LLMs by Output and First Answer
Celeris-1 (Celeris) is the fastest measured model at 1535 tokens/sec. Among models scoring 70 or higher, Gemini 3.8 Flash leads at 245 tok/s. LFM2-24B-A2B has the lowest first-answer latency at 0.42s.
The fastest output model is not always the fastest model for a short request. Tokens/sec measures sustained generation after output begins. First-answer latency measures the opening wait. Sort both, keep a quality floor, and replay the request shape your application will use.
Median tokens/s, Latency first answer chunk (s). Measurements are screening evidence, not provider service-level guarantees.
Celeris-1
1535 tok/s · Celeris
LFM2-24B-A2B
0.42s to first answer · LiquidAI
Gemini 3.8 Flash
245 tok/s · Score: 73.92
Top 15 — output speed (tok/s)
- Celeris-11535 tok/s, Celeris, 0.56 seconds first-answer latency, score 28.9
- Mercury 2871 tok/s, Inception, 4.54 seconds first-answer latency, score 35.82
- Mercury 2.5661 tok/s, Inception, 3.45 seconds first-answer latency, score 35.18
- Mercury 2.5 Preview661 tok/s, Inception, 3.45 seconds first-answer latency
- Trinity-Large-Thinking342 tok/s, Arcee AI, 7.19 seconds first-answer latency, score 34.86
- Gemini 3.5 Flash-Lite339 tok/s, Google, 9.44 seconds first-answer latency, score 51.53
- LFM2.5-8B-A1B335 tok/s, LiquidAI, 7.62 seconds first-answer latency, score 29.76
- Ling 3.0 Flash325 tok/s, InclusionAI, 9.04 seconds first-answer latency, score 46.38
- GPT-OSS 20B313 tok/s, OpenAI, 0.65 seconds first-answer latency, score 34.27
- Gemini 3.7 Flash296 tok/s, Google, 13.5 seconds first-answer latency, score 67.94
- Nemotron 3 Nano Omni 30B A3B264 tok/s, NVIDIA, 7.98 seconds first-answer latency, score 31.28
- Gemini 3.1 Flash-Lite250 tok/s, Google, 5.8 seconds first-answer latency, score 49.49
- Gemini 3.8 Flash245 tok/s, Google, 21.76 seconds first-answer latency, score 73.92
- DeepSeek V4 Flash 0731219 tok/s, DeepSeek, 10.18 seconds first-answer latency
- Muse Spark 1.1217 tok/s, Meta, 12.12 seconds first-answer latency, score 66.6
Bar length is relative to the fastest model in the current filtered view. Exact median speed is shown at right; open a model for its complete evidence profile.
Average speed by provider
Inception
731 tok/s avg · 3 models
3.8s avg latency
InclusionAI
215 tok/s avg · 2 models
5.1s avg latency
LiquidAI
208 tok/s avg · 3 models
6.7s avg latency
NVIDIA
196 tok/s avg · 4 models
11.6s avg latency
177 tok/s avg · 13 models
17.2s avg latency
Cohere
143 tok/s avg · 2 models
19.4s avg latency
DeepSeek
142 tok/s avg · 5 models
5.5s avg latency
StepFun
139 tok/s avg · 3 models
18.9s avg latency
Meta
133 tok/s avg · 6 models
15.2s avg latency
Mistral
116 tok/s avg · 7 models
2.7s avg latency
OpenAI
113 tok/s avg · 28 models
55.5s avg latency
xAI
96 tok/s avg · 9 models
26.2s avg latency
Alibaba
88 tok/s avg · 14 models
39.4s avg latency
MiniMax
78 tok/s avg · 3 models
32.9s avg latency
Xiaomi
72 tok/s avg · 3 models
39s avg latency
Anthropic
65 tok/s avg · 15 models
125.7s avg latency
Z.AI
63 tok/s avg · 9 models
30.5s avg latency
Moonshot AI
60 tok/s avg · 5 models
27.4s avg latency
Microsoft
29 tok/s avg · 2 models
1.7s avg latency
145 models match these filters.
Runtime rows use the same cross-provider feed. Prompt shape, region, provider load, account tier, cache state, and reasoning settings can change production results.
Two clocks choose different winners
First-answer latency dominates short interactive work: routing, classification, a tool argument, or the opening phrase of a voice turn. Sustained output dominates long reports and code generation after the response has started.
A rough completion estimate is first-answer time plus requested output tokens divided by tokens per second. It will not capture jitter, stop time, reasoning tokens, or retries, but it reveals when a slower starter can overtake a slower writer.
Read the TTFT explainer for timing boundaries, prompt prefill, and a reproducible measurement plan.
Keep speed behind a quality gate
A model that returns the wrong tool arguments quickly is not a fast production system. Build a task suite first, then identify the lowest-latency candidates that clear it. The score column is a broad screening signal, not a substitute for that application test.
Compare at least p50 and p95 latency from the deployment region. Hold prompt, output length, concurrency, API route, and reasoning budget constant. Keep failures and rate limits in the report so the surviving requests do not create a flattering distribution.
For voice, measure end of user speech to first audible output across transcription, model, speech generation, network, and playback. Text first-answer latency is only one segment.
Questions
What does tokens per second mean for LLMs?
Tokens per second measures sustained output after generation begins. A higher rate shortens long answers, but it does not reduce the initial wait by itself. Tokenization also differs across model families, so use the rate to screen candidates and confirm total completion time on the prompts your application sends.
What does the latency column measure?
Latency is the cross-provider runtime feed’s time from request to the first usable answer chunk. Lower is better within that test. For reasoning models it can include work performed before visible output. It is not full voice latency, browser render time, or a service-level guarantee for another region and account.
Which LLM is the fastest?
Currently, Celeris-1 by Celeris is the fastest at 1535 tokens/second. The fastest model scoring above 70 overall is Gemini 3.8 Flash at 245 tok/s.
Why are reasoning models slower?
Some reasoning settings perform additional work before a usable answer appears, which can increase first-answer latency and billed output. The effect varies by model, provider, prompt, and reasoning budget. Measure the exact setting you will deploy, then compare task success with latency and cost rather than assuming every reasoning label behaves alike.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.