**Russ d'Sa** (0:00)
We knew upfront going into it that we were going to take a lot of outages. Because it was novel and there were going to be issues that we just did not have software mitigations for. But we knew that long-term, it was the right decision because you can't just depend on AWS, but it just means you can also depend on GCP, or Azure, or DigitalOcean, or LinodacMI. It's like you have to actually leverage all of them together. And so this was another really important decision where we kind of run an overlay network across all of them. How do you make sure that you don't take an outage if somebody's cloud provider goes down? Well, you have to build across multiple cloud providers. It's just not realistic to expect a system to be perfect forever. When you have to transition to the system that I'm describing, that we built kind of from the very beginning, it's extremely difficult to do while the plane is in flight. Very difficult to do.
**Jerry Li** (0:53)
Hello and welcome to The Engineering Leadership Podcast, brought to you by ELC, The Engineering Leadership Community. I'm Jerry Li, founder of ELC. And I'm Patrick Gallagher, and we're your hosts. Our show shares the most critical perspectives, habits, and examples of great software engineering leaders to help evolve leadership in the tech industry.
Russ D'SA, CEO and co-founder at LiveKit joins us to deconstruct the product paradigm shift happening right now and how this will impact strategies and roadmaps. We are talking about the shift toward voice-driven interfaces and agent-centric UX. Plus, we dive into a bunch of stories behind LiveKit's origin, including some of their high-stakes scaling lessons. From powering OpenAI's voice mode, hitting the next level of scale and navigating real-time bottlenecks with Character AI's voice mode launch, to the architectural necessity of a multi-cloud strategy from day one and the foundations of a co-founder relationship that can effectively blend engineering and business strategy. Let me introduce you to Russ and LiveKit. Russ D'SA is the co-founder and CEO of LiveKit, the voice AI platform behind ChatGPT, Character AI, and 11 Labs. He started his first company in the 2007 batch of YC, and was the second front-end engineer at Twitter. LiveKit began as an open-source project for building live streaming and video conferencing applications using WebRTC. Over time, they evolved into a developer platform for building voice, video, and physical AI agents. What started with just a media server and some SDKs is now a full ecosystem of APIs and tools for multimodal computing. Enjoy our conversation with RUSS D'SA.
RUSS, I just want to say welcome. What's been kind of fun is every time you and I have talked, we have gotten into some sort of paradigm shifting idea. And I think that actually kind of sets the stage for our conversation today. And so, I mean, we have a few things around like product paradigm shifts that are happening. And so I guess as is tradition, since you and I have connected a couple times now, let's get into some paradigm shifts. So can you introduce us a little bit to this product paradigm shift that's going on? We've been kind of talking about this world of voice-driven apps and computer interfaces and some of the major shifts going on there. So I guess bring us in, like, what are some of these fundamental shifts happening in this world right now?
**Russ d'Sa** (2:59)
So there's so many things that are going on right now, 2026, that feel like sci-fi even five years ago. One of those things that is kind of relevant to my work is the way that you interact with computers. There is this paradigm shift in how computers feel to use them, right? The way you interact with computers has changed over time. So it started off with these punch cards, and then you had a keyboard and a mouse, and then you have a touchscreen, and the computers that you wear even now, like you wear an Apple Watch, it has a touchscreen on it though, and that's the primary way that you interact with that thing. But with LLMs and Generative AI, suddenly we have a computer in the abstract sense that can understand human thoughts through a written or audio-based form, right? Like it's read the Internet and it can write. The way that the trajectory of these models is that they're becoming more and more realistic in the way that they're able to understand what you say to it, and then speak back to you in a convincingly human way, right? It has all of this knowledge from the Internet, and then the way it phrases things, like it can do so and write in very human-like ways, and it doesn't just mean that it writes like perfect English all the time, it can also like speak in slang if you instruct it to and all of these kinds of things, different languages. Now, with an AI model like that, you can create software with that AI model at its core that you can talk to. That's I think the magic unlock that we've seen recently in this industry we call Voice AI, is that now you can interact with software applications using just your natural human input and output. I can speak to the computer, it can understand me, maybe it can even pick up on my mood or how I'm feeling that day, and then it can commensurately respond in a way that feels empathetic and either helps me think through something that I'm trying to work through in my own mind, or it can even do work for me or perform tasks on my behalf, and I don't even have to go and sit in front of the computer for a while and perform those tasks myself. Now the computer can kind of do it, almost like a person might if I was working with another person on a particular thing. And so this is just very sci-fi.
42 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773908625