# GPT-Live-1 API launch evaluation: Voice Benchmark Profile

> OpenAI’s API launch post reports spoken task success, full-duplex interactivity, tool calling, response quality, and turn-taking latency for GPT-Live-1 against GPT-Realtime-2.1 and GPT-Realtime-2.

- Scope: Tau3 (Voice) airline, retail, and telecom pass@1; Tau Banking (Voice) knowledge pass@1 over 97 tasks; Artificial Analysis Conversational Dynamics; Full Duplex Bench v1.5 interactivity, v1 turn-taking latency, and v3 tool calling and response quality
- Measurement lanes: Task completion, Conversation dynamics, Latency
- Primary metric: Pass@1 and average scores (higher is better); turn-taking latency in seconds (lower is better)
- Available evidence: Exact figures embedded in the official launch post charts
- Owner: OpenAI
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

These are provider-reported launch results against OpenAI’s own Realtime models, not an independent voice-agent ranking. The Tau3 and Tau Banking rows pair GPT-Live-1 with GPT-6 Astra at medium reasoning effort as the delegated backend, and the Full Duplex Bench v3 rows use the Terra backend at low effort. The Full Duplex Bench versions and runs differ from the v3 paper table BenchLM mirrors separately.

## Primary sources

- [Owner page](https://openai.com/index/introducing-gpt-live-1-in-the-api/)

Canonical page: https://benchlm.ai/voice-benchmarks/gpt-live-1-launch-evaluation
