GPT-Live-1 API launch evaluation
OpenAI’s API launch post reports spoken task success, full-duplex interactivity, tool calling, response quality, and turn-taking latency for GPT-Live-1 against GPT-Realtime-2.1 and GPT-Realtime-2.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Tau3 (Voice) airline, retail, and telecom pass@1; Tau Banking (Voice) knowledge pass@1 over 97 tasks; Artificial Analysis Conversational Dynamics; Full Duplex Bench v1.5 interactivity, v1 turn-taking latency, and v3 tool calling and response quality
- Primary metric
- Pass@1 and average scores (higher is better); turn-taking latency in seconds (lower is better)
- Owner
- OpenAI
- Available evidence
- Exact figures embedded in the official launch post charts
What it measures
Available results
- Tau3 (Voice) intelligence pass@1
- 86.2%
- Tau Banking (Voice) knowledge pass@1
- 32.0%
- Artificial Analysis Conversational Dynamics
- 97.3%
- Full Duplex Bench v1.5 interactivity
- 80.1%
- Full Duplex Bench v1 turn-taking latency
- 0.798 s
- Full Duplex Bench v3 tool calling pass@1
- 87.0%
- Full Duplex Bench v3 response quality
- 90.0%
Spoken airline, retail, and telecom customer-service tasks with each domain weighted equally; GPT-Live-1 delegates to GPT-6 Astra (medium). GPT-Realtime-2.1 scores 45.7% and GPT-Realtime-2 42.4%.
Fraction of 97 banking_knowledge tasks completed with knowledge retrieval and account tools, again with the Astra (medium) backend. GPT-Realtime-2.1 scores 12.4% and GPT-Realtime-2 10.3%.
Average score across pause handling, conversational turn taking, interruptions, and backchannels. GPT-Realtime-2.1 scores 95.7% and GPT-Realtime-2 95.3%.
Reactions to background speech, speech to another person, listener backchannels, and interruptions. GPT-Realtime-2.1 scores 45.4% and GPT-Realtime-2 47.8%.
Time from the end of the user’s turn to the start of the reply; lower is better. GPT-Realtime-2.1 takes 1.41 s and GPT-Realtime-2 1.63 s.
Tool-call sequences from spoken requests containing natural pauses, hesitations, and self-corrections, with the Terra (low) backend. GPT-Realtime-2.1 scores 60.0% and GPT-Realtime-2 58.0%.
How well the spoken answer to tool-using requests matches the reference intent, with the Terra (low) backend. GPT-Realtime-2.1 scores 88.0% and GPT-Realtime-2 81.0%.