# Gemini 3.8 Audio model evaluation: Voice Benchmark Profile

> Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index.

- Scope: Live API voice-agent evaluations using ServiceNow EVA-Bench, a source-owned speech-to-speech index and τ-Voice, and Sierra τ³-Banking
- Measurement lanes: Task completion, Conversation dynamics, Voice experience
- Primary metric: Task success and speech-to-speech index (higher is better)
- Available evidence: Exact bar-chart labels in Google’s model evaluation PDF; EVA-Bench scatterplot points are not tabulated
- Owner: Google DeepMind
- Source snapshots refreshed: 2026-09-15

## Interpretation limit

Google’s PDF combines evaluations run by separate teams using different APIs and effort settings. The Extended Thinking task rows use high effort; the EVA-Bench scatterplot labels are plotted without exact figures. These voice-protocol results do not enter weighted text-model scores, and the plotted audio-hour costs are not canonical API prices.

## Primary sources

- [Dataset](https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_live_model_evaluation.pdf)
- [Owner page](https://deepmind.google/models/model-cards/gemini-3-8-audio/)

Canonical page: https://benchlm.ai/voice-benchmarks/gemini-3-8-audio-evaluation
