OpenAI's Talk-Listen Model and Its Implications artwork

OpenAI's Talk-Listen Model and Its Implications

AI Update

July 8, 2026

In this episode, we explore the implications of OpenAI's new model that engages in both talking and listening. Join us for a fascinating discussion on future applications.
**SPEAKER_1** (0:00)
Are you subscribed to aichatdaily.com? Yeah, I have a newsletter that I send out every single day, and it's essentially all of the same content, all of the same stories that I talk about in the podcast, but I do them in these really easy to break down formats, and I also have links to articles to go more in depth. So it's kind of like show notes for the podcast, but it goes into way more depth, and we have extra stories that we can't get around to covering in the podcast. So if every single day you want to start the day out or end the day by getting a roundup of the latest in AI news and a really easy to scan format, go to aichatdaily.com. There's a big subscribe button in the top corner. I'll probably leave a link in the show notes to the newsletter. And also is just a full news site with tons of articles. I'll probably publish about 10 of them every single day. So if you want to go check that out, aichatdaily.com. Okay, OpenAI is shipping GPT Live once. This is a full duplex voice model. And what duplex voice model means? Well, this is actually kind of crazy because Miriam Murady, who was formerly at OpenAI, she released something recently with her most recent company, Thinking Machine Labs, where it's, I mean, essentially, it looks like OpenAI has just knocked her off. So this is going to be interesting. Also, China has flagged Anthropix Cloud code. They say that this is a security risk because there's some backdoor risks with it, which I don't know. I just think anything with security risk is ironic coming from China. Nvidia's Nemotron 3 Ultra has just hit closed model parity at a 10x lower cost, and this is from some experiments they've been doing on laying chain. I'm really excited when we're able to get the costs down on some of these models. Meta is adding a tamper-proof LED to their AI Meta glasses, and they're also getting more data collection from other parts of the glasses. So there's kind of a mixed bag with what's going on there. An AI, the OpenAI chief futurist, John Etchiam, is leaving. He's been there for nine years. He's one of the OG trust and safety people over OpenAI. But let's kick the podcast off with OpenAI's latest model, GPT Live 1 This is a voice model. It listens and it speaks at the same time, which is something that you've probably noticed. A lot of these voice models have struggled with in the past. If you're talking to what voice mode is on ChatGPT right now, which by the way, this replaces their advanced voice mode. But if you've ever talked to it before, you talk, there's a couple seconds of latency, then it responds back to you. And they've done a lot of work to try to get that latency down as much as possible. But what's really cool about this new Live 1 model is that it's literally listening, and it's able to listen and talk at the same time. Because that's how actual humans talk when we're in a conversation. We kind of interrupt each other, we talk over each other. They're saying something, you're saying something. And I mean, the goal is not to interrupt and talk over each other, but it's just part of communication, right? Like, we don't always wait for there to be a perfect pause at the end of everything that someone says before we jump in. So that's what these models are getting really good at, which is really interesting. And what's to me really crazy about this is that Miriam Marotti released a model just like this. A number of months ago, I remember covering it on the podcast, and it was like the big thing from Thinking Machine Labs, and I was like, man, this is a really cool idea and concept. Well, turns out, you know, a couple months later, anyone that releases any sort of cool new feature, a couple months later, the frontier labs are going to be able to clone it reverse engineer and put it out themselves because, you know, they don't want to fall behind. The reason why I think this is one in particular is an important feature is because there's three separate things that are getting collapsed into one step. Number one, it's cutting the delay, so it's going to let your conversation be way more natural. And they're also going to be rolling this out to 150 million ChatGPT voice users. And finally, the way this works is really interesting to me. So there's something called GPT Live 1 mini, and this is going to be for basically all of the free users. It's kind of the default. But if you're on the paid tier, if you're paying $20 a month for ChatGPT, you get the larger GPT Live one, which is going to send some of the harder questions over to GPT 5.5 mid conversation. So you're going to be talking in the middle of the conversation if you're saying something kind of complex. And it's going to basically get half of your sentence, send it over to GPT 5.5 to try to come up with a response and send it back and put it out before you even finish your sentence, which is really crazy. It can be interrupted in the middle of a sentence. So if it's talking, you can interrupt it. It handles silences while you think. So just because you're pausing to think about something doesn't mean it has to jump in, which is something really annoying. I feel like I can barely keep my thoughts. I can barely think when I'm talking to these things because I got to talk so fast because if I stop talking, it's going to want to jump in and start talking to me. So I really appreciate that. And it's going to show charts or images alongside whatever you're talking to it about, which you know before it was just kind of this like animated strobe on your screen. But now if you're talking, you're like, hey, you know, what are the, you know, what's the weather going to be like? It's actually pulling up weather charts on the screen while it talks to you. So you get visual and audio, which is really cool.

10 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000776031643