Topics: Technology
**Andrew** (0:05)
Today, we're exploring the way engineers are relearning how to build software from scratch with Sahaj Garg, co-founder and CTO of Wispr Flow. And Wispr is a company that's been on everyone's lips lately, and for good reason. They're at the center of a shift that many teams are feeling, moving from voice to text, from never worked to a possible workflow that people now use every day, all day. And this change in the bottleneck going from different parts of the process to actually being your keyboard has led many leaders and senior engineers to discover that they can express their context, taste, and intent faster, and to help their teams move quicker with voice. Because what if the limiting factor in your organization isn't actually like a context window, but how quickly you can get the right context out of someone's head? So today we're going to be talking about shared context as being a valuable currency and all of the different changes that this introduces to how people can communicate with software. And we're going to get really practical about that as well. So Sahaj, welcome to Dev Interrupted.
**Sahaj Garg** (1:10)
Thank you, Andrew. It's fantastic to be here. Really excited for this conversation.
**Andrew** (1:14)
Me too. And I want to start at the top by just mentioning a little bit about using Whisper. I've been using Whisper Flow very recently, and I've totally fallen in love with the type of software that it is. It's a true delight to use, and I can see why it's dismantling the way that the engineers have traditionally approached working with code on their keyboard. I myself have definitely written a whole novel at this point, and Whisper even tells me as much. And so I'm actually so blown away by how much I can trust this technology to express what I'm trying to say. And that's what I really want to explore today, because I know me and our listeners too, we've all been burned in the past by bad transcripts, or that voice detect didn't really capture what I said, or that voice note had a crazy typo in it. And those little tiny burns, they add up over time, and people walk away from the technology. But we're seeing a shift now where you're rebuilding trust in a technology that many had dismissed. So what has that been like for you at Whispr, and how did you approach that challenge?
**Sahaj Garg** (2:15)
Yeah, it's a fantastic question. It's actually one of the reasons why we almost never built the product that we did, because in some ways, I actually thought using voice for communication, for typing, for interaction on your computer was like a fundamentally doomed thing after 20 years of being disappointed by every product there was in the market. And I think the thing that we learned as we built it, and especially the first version was like, oh my God, when this works, it's magical. And that is the thing that we have so consistently heard from people who use the product over time, even now, which is when it works, it's magical. And all the work that we try to do here is to expand the settings and the context and the places in which you did get it right for the first try. Because as a user, what I want is a system that just gets me intuitively, I don't have to explain myself a bunch of times. It should know if I'm talking to Cloud Code and I'm talking about a.env file, what that actually means, right? Not show up in a completely nonsensical way. I'd say the biggest framing for us of this problem is we want to build you a voice interaction where you never have to go back and fix the mistake. For us, we call this zero edit rate. Something where you have to fix no mistakes with what you're doing. That means both getting everything you said right and figuring out what you actually meant to convey, so that we can help you fix that up on your behalf in a way that sounds just like you and that you can actually use downstream.
**Andrew** (3:43)
Really, it's like a two-part equation because you have the traditional layer of like, oh yes, we can turn this voice into the words it is, but then there's also this contextual layer of understanding what are you operating in, what are the words around you, and what have you said recently, what are the things that matter to you as the user? And combining those two is like the formula I think that Whispers is getting right and is what's letting people work very quickly with it. So I'm curious in how when you solve the problem of rebuilding trust and doing it in this two-part way and acknowledging that it's a nuanced engineering problem, and so you're approaching this with your teams and we're now in a green field because we've acknowledged that this is a lot more nuanced than how we solve this. So what assumptions there in that world do you throw away to ultimately arrive at an application or a system that's a delight, that gets that zero edit rate? I'm curious like what becomes the secret levers.
31 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID