Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Official provider results

Gemini 3.8 Audio model evaluation

Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
Live API voice-agent evaluations using ServiceNow EVA-Bench, a source-owned speech-to-speech index and τ-Voice, and Sierra τ³-Banking
Primary metric
Task success and speech-to-speech index (higher is better)
Owner
Google DeepMind
Available evidence
Exact bar-chart labels in Google’s model evaluation PDF; EVA-Bench scatterplot points are not tabulated

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Voice experience
Is the exchange natural, robust, and responsive?

Available results

Gemini 3.8 Live speech-to-speech index
76.0%

Source-owned index shown in Google’s September 2026 evaluation PDF.

Gemini 3.8 Live τ-Voice task success
30.1%

Google’s comparison chart for the default Live model.

Gemini 3.8 Live Extended Thinking speech-to-speech index
82.6%

High-effort Extended Thinking run.

Gemini 3.8 Live Extended Thinking τ-Voice task success
68.6%

High-effort Extended Thinking run; do not attribute this result to the default Live model.

Gemini 3.8 Live Extended Thinking Sierra τ³-Banking
35.1%

High-effort Live API run on the banking task suite.

Interpretation limit

Google’s PDF combines evaluations run by separate teams using different APIs and effort settings. The Extended Thinking task rows use high effort; the EVA-Bench scatterplot labels are plotted without exact figures. These voice-protocol results do not enter weighted text-model scores, and the plotted audio-hour costs are not canonical API prices.