Gemini 3.5 Transcribe launch evaluation
Google’s August 26, 2026 launch post reports word error rates on an independent transcription leaderboard and FLEURS for pre-recorded and streaming transcription, plus a latency gain over Chirp 3.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- independent-leaderboard WER (non-streaming and streaming), FLEURS multilingual WER, and time-to-final-transcript versus Chirp 3
- Primary metric
- Word error rate (lower is better)
- Owner
- Available evidence
- Exact figures published in the official launch post
What it measures
Available results
- Independent WER leaderboard (pre-recorded)
- 2.6%
- Independent WER leaderboard (streaming)
- 4.0%
- FLEURS WER (pre-recorded)
- 5.04%
- FLEURS WER (streaming)
- 5.50%
- Time to final transcript
- 70% faster than Chirp 3
Average across the independent leaderboard’s transcription set; lower is better.
Real-time streaming mode.
Multilingual FLEURS benchmark.
Multilingual FLEURS benchmark, streaming mode.
Relative latency claim versus Google’s previous transcription model.