OpenAI released voice models GPT-Live-1 and GPT-Live-1 mini for ChatGPT
OpenAI updated ChatGPT's voice mode: the new models GPT-Live-1 and the lighter GPT-Live-1 mini run on full-duplex architecture — listening and speaking simultaneously. The main difference: users can interrupt the assistant's response at any moment in the conversation, rather than waiting for the model to finish, as before.
AI-processed from 36Kr (36氪); edited by Hamidun News
OpenAI in early July 2026 released a new generation of voice models for ChatGPT—GPT-Live-1 and a lightweight version GPT-Live-1 mini, fully transitioning the voice mode of the assistant to a full-duplex architecture. Chinese technology digest 36Kr (36氪) reported this citing sources.
What Changed in ChatGPT's Voice Mode
Previous versions of ChatGPT's voice assistant operated on the principle of strict turn-taking: the user speaks a phrase, the model waits for a pause, and only then formulates an answer that cannot be interrupted until the very end. GPT-Live-1 and GPT-Live-1 mini change this scenario completely—both models are built on a "full duplex" architecture that allows the model to simultaneously accept user speech and generate its own, instead of working in turns.
- New models are named GPT-Live-1 (main version) and GPT-Live-1 mini (lightweight, for less resource-intensive scenarios)
- Both are built on a full-duplex architecture—speech input and generation happen in parallel, not sequentially
- The user can interrupt the model's answer at any point in the conversation
- The update affected the entire voice functionality of ChatGPT as a whole
Why Conversation with the Assistant Became More Natural
The main consequence of full duplex is the elimination of the awkward pauses and rigid turn-taking characteristic of voice bots of the previous generation. If a user clarifies a question midway through a phrase, changes the formulation, or simply wants to interject "wait, that's not what I meant," GPT-Live-1 reacts immediately rather than finishing a pre-generated answer to the end, as happened in earlier ChatGPT voice modes.
Technically, full duplex is more complex to implement than the familiar chain "recognize speech → formulate answer → synthesize voice": the model must constantly decide whether to continue speaking, listen, or do both simultaneously without creating an echo effect and without losing the thread of conversation. This is precisely why such systems have remained rare even among major laboratories until recently.
Why There's a Lightweight Version
The existence of two models—GPT-Live-1 and GPT-Live-1 mini—is a typical approach for voice services where response latency is critical to the feeling of "live" dialogue. The lightweight version is designed for scenarios where response speed and lower infrastructure load matter more, while the main model can be used where quality and accuracy of the answer are the priority.
Where This Will Be Useful First
Responsiveness in conversation is especially noticeable where voice is the only way to interact with the assistant: while driving, while working with hands, in dictation scenarios or voice customer support. In such situations, an extra second of waiting or inability to quickly correct the model feels much more acutely than in a text chat where one can simply add a clarification in the next message. Full duplex removes precisely this source of frustration—conversation flows without artificial pauses of "speak after the signal."
What This Means
The launch of GPT-Live-1 shows that competition in voice AI assistants has shifted from pure speech synthesis quality to the naturalness of the dialogue itself—how much conversation with the model feels alive rather than an exchange of lines with an answering machine. Users who increasingly communicate with ChatGPT by voice rather than text get an assistant that behaves closer to a live conversation partner with their interruptions and clarifications, rather than a classic voice menu with rigid scenarios. Given that Google already has its own voice mode Gemini Live, ChatGPT's voice engine update looks like a direct response to growing competition for users who prefer to speak with AI rather than type.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.