Gemini 3.8 Audio model evaluation
Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Live API voice-agent evaluations using ServiceNow EVA-Bench, a source-owned speech-to-speech index and τ-Voice, and Sierra τ³-Banking
- Primary metric
- Task success and speech-to-speech index (higher is better)
- Owner
- Google DeepMind
- Available evidence
- Exact bar-chart labels in Google’s model evaluation PDF; EVA-Bench scatterplot points are not tabulated
What it measures
Available results
- Gemini 3.8 Live speech-to-speech index
- 76.0%
- Gemini 3.8 Live τ-Voice task success
- 30.1%
- Gemini 3.8 Live Extended Thinking speech-to-speech index
- 82.6%
- Gemini 3.8 Live Extended Thinking τ-Voice task success
- 68.6%
- Gemini 3.8 Live Extended Thinking Sierra τ³-Banking
- 35.1%
Source-owned index shown in Google’s September 2026 evaluation PDF.
Google’s comparison chart for the default Live model.
High-effort Extended Thinking run.
High-effort Extended Thinking run; do not attribute this result to the default Live model.
High-effort Live API run on the banking task suite.