Live

OpenAI ships gpt-realtime-2.1 for faster voice and multimodal apps

The updated realtime model family cuts p95 latency by at least a quarter while improving recognition, tool use and instruction following for voice agents.

TKTomislav Krajina
Published 7 Jul 2026, 14:29Updated 28 Jul 2026, 05:051 min read
OpenAI ships gpt-realtime-2.1 for faster voice and multimodal apps

OpenAI has released gpt-realtime-2.1 and a smaller gpt-realtime-2.1-mini variant, aimed squarely at developers building low-latency voice and multimodal experiences.

What improved

  • At least 25% lower p95 latency across the Realtime voice model line
  • Better recognition and noise handling in real-world audio conditions
  • Stronger reasoning, tool use and instruction following mid-conversation

Why it matters

Voice agents are unusually sensitive to latency — a half-second delay breaks the illusion of a live conversation. Shaving p95 latency rather than just average latency targets the worst-case stalls that make voice products feel broken, which is typically the harder engineering problem to solve.

Source: OpenAI

Related stories

All →