Skip to main content
BenchLM

Fastest LLMs by Output and First Answer

Celeris-1 (Celeris) is the fastest measured model at 1535 tokens/sec. Among models scoring 70 or higher, Gemini 3.8 Flash leads at 245 tok/s. LFM2-24B-A2B has the lowest first-answer latency at 0.42s.

Runtime feed verified · Artificial Analysis

The fastest output model is not always the fastest model for a short request. Tokens/sec measures sustained generation after output begins. First-answer latency measures the opening wait. Sort both, keep a quality floor, and replay the request shape your application will use.

Median tokens/s, Latency first answer chunk (s). Measurements are screening evidence, not provider service-level guarantees.

Fastest Output

Celeris-1

1535 tok/s · Celeris

Lowest Latency

LFM2-24B-A2B

0.42s to first answer · LiquidAI

Fastest (Score 70+)

Gemini 3.8 Flash

245 tok/s · Score: 73.92

Top 15 — output speed (tok/s)

Ultra FastFastMediumSlow
  1. Celeris-11535 tok/s, Celeris, 0.56 seconds first-answer latency, score 28.9
  2. Mercury 2871 tok/s, Inception, 4.54 seconds first-answer latency, score 35.82
  3. Mercury 2.5661 tok/s, Inception, 3.45 seconds first-answer latency, score 35.18
  4. Mercury 2.5 Preview661 tok/s, Inception, 3.45 seconds first-answer latency
  5. Trinity-Large-Thinking342 tok/s, Arcee AI, 7.19 seconds first-answer latency, score 34.86
  6. Gemini 3.5 Flash-Lite339 tok/s, Google, 9.44 seconds first-answer latency, score 51.53
  7. LFM2.5-8B-A1B335 tok/s, LiquidAI, 7.62 seconds first-answer latency, score 29.76
  8. Ling 3.0 Flash325 tok/s, InclusionAI, 9.04 seconds first-answer latency, score 46.38
  9. GPT-OSS 20B313 tok/s, OpenAI, 0.65 seconds first-answer latency, score 34.27
  10. Gemini 3.7 Flash296 tok/s, Google, 13.5 seconds first-answer latency, score 67.94
  11. Nemotron 3 Nano Omni 30B A3B264 tok/s, NVIDIA, 7.98 seconds first-answer latency, score 31.28
  12. Gemini 3.1 Flash-Lite250 tok/s, Google, 5.8 seconds first-answer latency, score 49.49
  13. Gemini 3.8 Flash245 tok/s, Google, 21.76 seconds first-answer latency, score 73.92
  14. DeepSeek V4 Flash 0731219 tok/s, DeepSeek, 10.18 seconds first-answer latency
  15. Muse Spark 1.1217 tok/s, Meta, 12.12 seconds first-answer latency, score 66.6

Bar length is relative to the fastest model in the current filtered view. Exact median speed is shown at right; open a model for its complete evidence profile.

Average speed by provider

Inception

731 tok/s avg · 3 models

3.8s avg latency

InclusionAI

215 tok/s avg · 2 models

5.1s avg latency

LiquidAI

208 tok/s avg · 3 models

6.7s avg latency

NVIDIA

196 tok/s avg · 4 models

11.6s avg latency

Google

177 tok/s avg · 13 models

17.2s avg latency

Cohere

143 tok/s avg · 2 models

19.4s avg latency

DeepSeek

142 tok/s avg · 5 models

5.5s avg latency

StepFun

139 tok/s avg · 3 models

18.9s avg latency

Meta

133 tok/s avg · 6 models

15.2s avg latency

Mistral

116 tok/s avg · 7 models

2.7s avg latency

OpenAI

113 tok/s avg · 28 models

55.5s avg latency

xAI

96 tok/s avg · 9 models

26.2s avg latency

Alibaba

88 tok/s avg · 14 models

39.4s avg latency

MiniMax

78 tok/s avg · 3 models

32.9s avg latency

Xiaomi

72 tok/s avg · 3 models

39s avg latency

Anthropic

65 tok/s avg · 15 models

125.7s avg latency

Z.AI

63 tok/s avg · 9 models

30.5s avg latency

Moonshot AI

60 tok/s avg · 5 models

27.4s avg latency

Microsoft

29 tok/s avg · 2 models

1.7s avg latency

145 models match these filters.

RankModelSpeed / latency
1
Celeris-1Celeris · score 28.9
1535 tok/s0.56s latency
2
Mercury 2Inception · score 35.82
871 tok/s4.54s latency
3
Mercury 2.5Inception · score 35.18
661 tok/s3.45s latency
9
GPT-OSS 20BOpenAI · score 34.27
313 tok/s0.65s latency
16
Command A+Cohere · score 36.6
215 tok/s9.83s latency
19
Qwen3.7 MaxAlibaba · score 63.44
203 tok/s14.09s latency
21
LFM2.5-2.6BLiquidAI · score 30.95
197 tok/s12.02s latency
37
InklingThinking Machines Lab · score 54.78
160 tok/s14.36s latency
38
o3-miniOpenAI · score 40.84
160 tok/s7.12s latency
43
Quasar 438BMultiverse Computing · score 49.83
143 tok/s15.16s latency
44
GPT-4.1OpenAI · score 39.28
140 tok/s0.88s latency
49
GPT-5 nanoOpenAI · score 37.08
137 tok/s83.3s latency
53
GPT-4oOpenAI · score 30.04
126 tok/s1.05s latency
56
Grok 4.3xAI · score 55.33
120 tok/s27.88s latency
58
o3OpenAI · score 47.72
115 tok/s6.53s latency
60
Nova ProAmazon · score 22.15
110 tok/s1.07s latency
62
Grok 4.20xAI · score 59.18
105 tok/s20.5s latency
64
GPT-5.1OpenAI · score 58.95
103 tok/s38.42s latency
65
MiniMax M3MiniMax · score 55.16
98 tok/s21.74s latency
66
o1OpenAI · score 40.06
98 tok/s32.29s latency
67
GPT-5.4OpenAI · score 68.89
97 tok/s190.56s latency
70
GPT-5.5OpenAI · score 69.38
93 tok/s52.42s latency
74
Qwen3.5 397BAlibaba · score not scored
89 tok/s38.05s latency
76
Hy3Tencent · score 53.22
88 tok/s25.74s latency
77
GLM-5.2Z.AI · score 63.15
86 tok/s26.39s latency
80
GPT-5 miniOpenAI · score 44.75
86 tok/s65.32s latency
82
Kimi K2.5Moonshot AI · score 53.65
82 tok/s38.77s latency
83
GPT-5.6 SolOpenAI · score 78.91
81 tok/s116.23s latency
85
Grok 4.7xAI · score not scored
81 tok/s88.21s latency
89
Solar Pro 4Upstage · score not scored
79 tok/s27.31s latency
91
GLM-5Z.AI · score 55.04
76 tok/s42.73s latency
92
Qwen3.5-27BAlibaba · score 48.31
75 tok/s32.15s latency
93
GPT-5.4 ProOpenAI · score 70.72
74 tok/s151.79s latency
95
GPT-5 (high)OpenAI · score not scored
71 tok/s1.27s latency
97
GPT-5.2OpenAI · score 61.92
70 tok/s137.63s latency
99
Grok 4.6xAI · score 69.38
69 tok/s35.9s latency
100
GLM-4.7Z.AI · score 49.83
67 tok/s31.13s latency
102
GLM-5.3Z.AI · score 65.66
65 tok/s34.13s latency
108
Grok 4.5xAI · score 65.63
58 tok/s10.66s latency
110
Qwen3 MaxAlibaba · score 41.08
57 tok/s2.19s latency
112
Qwen3.6 PlusAlibaba · score 55.19
56 tok/s101.51s latency
113
Qwen3.6-27BAlibaba · score 49.3
56 tok/s104.41s latency
115
GLM-5.1Z.AI · score 57.74
55 tok/s70.85s latency
116
Grok 4xAI · score 53.37
54 tok/s15.6s latency
117
Kimi K2Moonshot AI · score 29.64
53 tok/s1.57s latency
118
GPT-6 AstraOpenAI · score 88.69
51 tok/s320.21s latency
119
Kimi K2.6Moonshot AI · score 60.17
51 tok/s2.7s latency
120
GLM-4.5Z.AI · score not scored
51 tok/s1.45s latency
122
LongCat-2.0Meituan · score not scored
48 tok/s44.97s latency
127
Qwen3.8-27BAlibaba · score 57.56
44 tok/s49.56s latency
128
MiMo-V2.5Xiaomi · score not scored
44 tok/s64.65s latency
130
Phi-4Microsoft · score 26.37
40 tok/s2.61s latency
131
Qwen3.8 MaxAlibaba · score 72.1
39 tok/s53.87s latency
134
Gemma 4 31BGoogle · score 41.07
36 tok/s49.85s latency
136
Kimi K3Moonshot AI · score 72.12
34 tok/s63.59s latency
137
GLM-4.6Z.AI · score 40.38
33 tok/s4.92s latency
138
GPT-4o miniOpenAI · score 27.57
33 tok/s3.16s latency
139
GPT-4 TurboOpenAI · score 22.31
32 tok/s3.49s latency
140
Gemma 3 27BGoogle · score 29.46
31 tok/s2.04s latency
144
o3-proOpenAI · score 47.88
27 tok/s84.93s latency

Runtime rows use the same cross-provider feed. Prompt shape, region, provider load, account tier, cache state, and reasoning settings can change production results.

Two clocks choose different winners

First-answer latency dominates short interactive work: routing, classification, a tool argument, or the opening phrase of a voice turn. Sustained output dominates long reports and code generation after the response has started.

A rough completion estimate is first-answer time plus requested output tokens divided by tokens per second. It will not capture jitter, stop time, reasoning tokens, or retries, but it reveals when a slower starter can overtake a slower writer.

Read the TTFT explainer for timing boundaries, prompt prefill, and a reproducible measurement plan.

Keep speed behind a quality gate

A model that returns the wrong tool arguments quickly is not a fast production system. Build a task suite first, then identify the lowest-latency candidates that clear it. The score column is a broad screening signal, not a substitute for that application test.

Compare at least p50 and p95 latency from the deployment region. Hold prompt, output length, concurrency, API route, and reasoning budget constant. Keep failures and rate limits in the report so the surviving requests do not create a flattering distribution.

For voice, measure end of user speech to first audible output across transcription, model, speech generation, network, and playback. Text first-answer latency is only one segment.

Questions

What does tokens per second mean for LLMs?

Tokens per second measures sustained output after generation begins. A higher rate shortens long answers, but it does not reduce the initial wait by itself. Tokenization also differs across model families, so use the rate to screen candidates and confirm total completion time on the prompts your application sends.

What does the latency column measure?

Latency is the cross-provider runtime feed’s time from request to the first usable answer chunk. Lower is better within that test. For reasoning models it can include work performed before visible output. It is not full voice latency, browser render time, or a service-level guarantee for another region and account.

Which LLM is the fastest?

Currently, Celeris-1 by Celeris is the fastest at 1535 tokens/second. The fastest model scoring above 70 overall is Gemini 3.8 Flash at 245 tok/s.

Why are reasoning models slower?

Some reasoning settings perform additional work before a usable answer appears, which can increase first-answer latency and billed output. The effect varies by model, provider, prompt, and reasoning budget. Measure the exact setting you will deploy, then compare task success with latency and cost rather than assuming every reasoning label behaves alike.

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.