
How we built a realtime system for responsive voice AI in six months
OpenAI developed GPT-Live, a full-duplex voice system that allows simultaneous listening and speaking to eliminate the need for turn detectors. The system is designed for low latency to make AI conversations feel more natural.
Why it matters
Users experience sub-second responsiveness in ChatGPT Voice, enabling natural interruptions during conversation. This technology also powers new capabilities like controlling computers and coordinating agents in the desktop app.
The details
- The WARP protocol reduces media startup from six network round trips to one.
- Media frontend and inference logic were rewritten in Go for smoother frame delivery.
- GPT-Live can asynchronously delegate complex reasoning tasks to frontier models like GPT-5.5.
Show entities and relationshipsHide entities and relationships
In this article
Key connections
OpenAI owns ChatGPT Voice
OpenAI operates ChatGPT Voice.
OpenAI owns Realtime API
OpenAI provides the Realtime API.
OpenAI owns Advanced Voice Mode
OpenAI created Advanced Voice Mode.
OpenAI owns GPT-Live API
OpenAI is developing the GPT-Live API.
GPT-Live delegates deep reasoning and tool execution to GPT-5.5.
GPT-Live uses WebRTC for low-latency media transport.
Show 11 more connectionsShow fewer connections
GPT-Live uses WARP to minimize protocol handshake latency.
GPT-Live media frontend and inference logic were implemented in Go.
GPT-Live replaced an earlier Python asyncio implementation.
GPT-Live is related to ChatGPT Voice
GPT-Live powers interaction capabilities in ChatGPT Voice.
GPT-Live competes with Advanced Voice Mode
GPT-Live succeeds Advanced Voice Mode.
GPT-Live is related to Realtime API
GPT-Live extends the voice infrastructure established by Realtime API.
GPT-Live API is related to GPT-Live
GPT-Live API will expose GPT-Live capabilities to developers.
TSVWG is a working group within the IETF.
OpenAI works with standard collaborators within IETF to advance WARP.
WARP proposals are being advanced through the IETF TSVWG working group.
WARP utilizes DTLS 1.3 for faster transport handshakes.
Related events
OpenAI Unveils GPT-Live Realtime Voice Architecture
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.