By end of 2027, native full-duplex speech-to-speech will be the default real-time voice API for a majority of top frontier providers, with at least three advertising sub-500ms median response latency
Prediction Statement
By December 31, 2027, native full-duplex speech-to-speech will be the default real-time voice offering — not merely an option — from a majority of the top frontier voice providers, defined as the set of OpenAI, Google, Microsoft/Azure, Amazon, and xAI. Concretely: at least three of these five will offer a generally available (not preview) full-duplex voice API in which a single model ingests and emits audio simultaneously, and at least three will publicly advertise a median conversational response latency at or below 500 milliseconds for that API. A "default" means the full-duplex model is the one the vendor documentation steers new real-time voice developers to first, with the older cascade-oriented realtime endpoint positioned as legacy or special-purpose.
Reasoning and Analysis
The direction is already set. As of July 2026, OpenAI has shipped GPT-Live (full-duplex, in ChatGPT, API stated to follow), Microsoft has GPT-Realtime 1.5 and an Azure-Realtime model in public preview, xAI has a fast Grok Voice Agent, Amazon has Nova 2 Sonic, Google has Gemini 3.1 Flash Live, and NVIDIA has shipped PersonaPlex. Kyutai Moshi has demonstrated that native full-duplex can reach roughly 200-millisecond end-to-end latency by modeling user and system audio as parallel streams. The architecture works, the latency is achievable, and the biggest distribution channel in consumer AI now ships it by default.
The economic and product logic points the same way. Cascades cannot produce overlap, back-channels, or graceful interruption regardless of how fast they get, because they must detect the end of a turn before responding. Once a competitor ships a model that shares the floor like a person, "wait for me to finish" reads as broken, and every vendor is forced to follow. The reflex-plus-reasoning split that GPT-Live uses also makes always-on full-duplex economically viable by metering expensive reasoning instead of paying frontier prices per millisecond of listening, which removes the main cost objection.
The 500-millisecond bar is deliberately set above the ~200ms human floor and below the ~800ms-to-3s range where most 2026 systems sit, so it captures genuine progress without requiring every vendor to hit the theoretical best. Median rather than best-case latency is specified to prevent cherry-picked demo numbers from satisfying the claim.
Confidence Factors
Confidence is 72 — highly likely but not near-certain — and would rise toward the mid-80s if, within the next two quarters, the GPT-Live API reaches general availability and a second major provider promotes a full-duplex endpoint as its documented default. It would fall toward 55 if full-duplex adoption stalls at the preview stage through 2027, if latency claims cluster stubbornly above 500ms in independent testing, or if a serious safety or privacy incident around always-listening voice triggers regulatory friction that slows GA rollouts.
The largest genuine uncertainty is the gap between "shipped" and "default." Vendors may keep cascade endpoints as the documented default for compatibility and enterprise inertia even after full-duplex models exist, which would leave the capability available but not the default — failing the strict version of this claim even as the technology clearly wins.
Key Indicators
Signals to monitor between now and the target date: whether the GPT-Live developer API moves from "coming soon" to generally available; whether Azure promotes GPT-Realtime 1.5 or a successor out of preview and steers developers to it first; whether xAI, Amazon, or Google ship a documented full-duplex default; whether independent latency benchmarks (of the kind published through early 2026) show median response times crossing below 500ms for GA endpoints; and whether a new generation of full-duplex behavioral benchmarks becomes standard in vendor release notes, which would indicate the scoreboard itself has shifted.
Validation Criteria
At the target date, score accuracy as follows. Full credit (90-100%) if at least three of the five named providers offer a GA native full-duplex voice API AND at least three publicly advertise sub-500ms median response latency for it AND the full-duplex model is documented as the default path for at least three. Strong partial credit (70-89%) if the full-duplex-as-default condition is met by at least three providers but the sub-500ms latency claim is met by only two, or vice versa. Partial credit (50-69%) if full-duplex is generally available from three providers but positioned alongside rather than ahead of cascade endpoints. Low credit (30-49%) if only one or two providers reach GA full-duplex. Miss (0-29%) if full-duplex remains preview-only across the field or is abandoned. Latency claims are judged on vendor-published median figures corroborated by at least one independent benchmark where available.
Published: July 13, 2026
Prediction ID: full-duplex-voice-api-default-2027