OpenAI ships gpt-realtime-2.1 for faster voice and multimodal apps
The updated realtime model family cuts p95 latency by at least a quarter while improving recognition, tool use and instruction following for voice agents.

OpenAI has released gpt-realtime-2.1 and a smaller gpt-realtime-2.1-mini variant, aimed squarely at developers building low-latency voice and multimodal experiences.
What improved
- ▸At least 25% lower p95 latency across the Realtime voice model line
- ▸Better recognition and noise handling in real-world audio conditions
- ▸Stronger reasoning, tool use and instruction following mid-conversation
Why it matters
Voice agents are unusually sensitive to latency — a half-second delay breaks the illusion of a live conversation. Shaving p95 latency rather than just average latency targets the worst-case stalls that make voice products feel broken, which is typically the harder engineering problem to solve.
Source: OpenAI


