English

Product LaunchesOpenAI

OpenAI Redesigns WebRTC Stack to Achieve Low-Latency Voice AI

This article is a translation. Read the Japanese original

OpenAI's technical staff team, the division responsible for real-time AI interactions, announced that they have redesigned their WebRTC stack. This move resolves existing constraints that were incompatible with OpenAI's infrastructure.

Under the previous approach, challenges included media termination processes that assigned ports per session, as well as the stable ownership management of stateful ICE and DTLS sessions. Furthermore, it was necessary to keep the latency of the first hop in global routing low.

The new architecture adopts a "split relay transceiver" method, which changes how packets are routed internally within OpenAI while maintaining standard WebRTC behavior for the client.

WebRTC standardizes the difficult aspects of interactive media, such as NAT traversal, encrypted transport, and codec negotiation. Building upon this protocol stack already implemented in browsers and mobile platforms, OpenAI is focusing on constructing infrastructure that connects real-time media with its models.

In AI applications, it is critical that audio arrives as a continuous stream. This enables transcription, inference, and tool calling while the user is still speaking. This mechanism forms a decisive difference from push-to-talk systems.

OpenAI is building upon the WebRTC ecosystem and open-source implementations. Justin Uberti, a former architect of WebRTC, and Sean DuBois, a developer of Pion, are both employed at OpenAI and are supporting the utilization of the media infrastructure.


Source: How OpenAI delivers low-latency voice AI at scale(HN 510pt・146コメント) (HN Search (backfill), 2026-05-05)