Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Official provider results

GPT-Live-1 API launch evaluation

OpenAI’s API launch post reports spoken task success, full-duplex interactivity, tool calling, response quality, and turn-taking latency for GPT-Live-1 against GPT-Realtime-2.1 and GPT-Realtime-2.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
Tau3 (Voice) airline, retail, and telecom pass@1; Tau Banking (Voice) knowledge pass@1 over 97 tasks; Artificial Analysis Conversational Dynamics; Full Duplex Bench v1.5 interactivity, v1 turn-taking latency, and v3 tool calling and response quality
Primary metric
Pass@1 and average scores (higher is better); turn-taking latency in seconds (lower is better)
Owner
OpenAI
Available evidence
Exact figures embedded in the official launch post charts

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Latency
How long do responses, tool calls, and full tasks take?

Available results

Tau3 (Voice) intelligence pass@1
86.2%
Leads OpenAI’s three-model chart

Spoken airline, retail, and telecom customer-service tasks with each domain weighted equally; GPT-Live-1 delegates to GPT-6 Astra (medium). GPT-Realtime-2.1 scores 45.7% and GPT-Realtime-2 42.4%.

Tau Banking (Voice) knowledge pass@1
32.0%
Leads OpenAI’s three-model chart

Fraction of 97 banking_knowledge tasks completed with knowledge retrieval and account tools, again with the Astra (medium) backend. GPT-Realtime-2.1 scores 12.4% and GPT-Realtime-2 10.3%.

Artificial Analysis Conversational Dynamics
97.3%
Leads OpenAI’s three-model chart

Average score across pause handling, conversational turn taking, interruptions, and backchannels. GPT-Realtime-2.1 scores 95.7% and GPT-Realtime-2 95.3%.

Full Duplex Bench v1.5 interactivity
80.1%
Leads OpenAI’s three-model chart

Reactions to background speech, speech to another person, listener backchannels, and interruptions. GPT-Realtime-2.1 scores 45.4% and GPT-Realtime-2 47.8%.

Full Duplex Bench v1 turn-taking latency
0.798 s
Fastest of the three models charted

Time from the end of the user’s turn to the start of the reply; lower is better. GPT-Realtime-2.1 takes 1.41 s and GPT-Realtime-2 1.63 s.

Full Duplex Bench v3 tool calling pass@1
87.0%
Leads OpenAI’s three-model chart

Tool-call sequences from spoken requests containing natural pauses, hesitations, and self-corrections, with the Terra (low) backend. GPT-Realtime-2.1 scores 60.0% and GPT-Realtime-2 58.0%.

Full Duplex Bench v3 response quality
90.0%
Leads OpenAI’s three-model chart

How well the spoken answer to tool-using requests matches the reference intent, with the Terra (low) backend. GPT-Realtime-2.1 scores 88.0% and GPT-Realtime-2 81.0%.

Interpretation limit

These are provider-reported launch results against OpenAI’s own Realtime models, not an independent voice-agent ranking. The Tau3 and Tau Banking rows pair GPT-Live-1 with GPT-6 Astra at medium reasoning effort as the delegated backend, and the Full Duplex Bench v3 rows use the Terra backend at low effort. The Full Duplex Bench versions and runs differ from the v3 paper table BenchLM mirrors separately.