Dialpad's Chief Strategy Officer, Dan O'Connell, on AI-Powered Business Communications artwork

Dialpad's Chief Strategy Officer, Dan O'Connell, on AI-Powered Business Communications

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

January 9, 2024

In this episode, Nathan sits down with Dan O’Connell, Chief Strategy Officer at Dialpad. They discuss building their own language models using 5 billion minutes of business calls, custom speech recognition models for every customer, and the challenges of bringing AI into business.
Speakers: Erik Torenberg, Nathan Labenz, Dan O'Connell
**Erik Torenberg** (0:00)
Turpentine is a network of podcasts, newsletters and more covering tech, business and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors like Moment of Zen and my show Upstream. We're looking for industry-leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erikaturpentine.co. That's E-R-I-K at turpentine.co.

**Nathan Labenz** (0:45)
It's not that I'm racing to replace people or cut costs or whatever, but I always come back to a Bezos style like, what does the customer really want? And the customer wants immediate response 24-7, where I can pause the conversation where I want to at my convenience and be able to come back and pick it up right where I left off and maybe even switch modalities. And Chantubt offers me all these things today. Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host Eric Thornburg. Hello and welcome back to The Cognitive Revolution. As we head into 2024, I've been thinking a lot about where we are in terms of AI's impact on knowledge work. While 2023 certainly brought explosive growth in AI adoption, to be honest, things have moved a little less quickly than I had expected. At retail prices, $1 billion only buys 1 to 2 GPT-4 API calls for each of the world's 8 billion citizens, which really just goes to show what a tiny toehold language models have established in knowledge work globally.
Even if this were to go 100x over the next year, it's still just 1 GPT-4 API call per person per day, still a tiny fraction of the knowledge work that humans are doing. So why this delay relative to my admittedly very high expectations, and what are the likely solutions? First, while I've definitely argued that OpenAI has modes, I have been a bit surprised by how long it has taken for other companies to match the quality of GPT-4.
It's fair to say I think at this point that GPT-35 level models are effectively commoditized. But GPT-4 is different. Only Anthropic, and now Google with Gemini, are really even in the ballpark in the West. Though it's worth noting, and I do hope to have another episode on this soon, that Baidu's Ernie 4 also appears to be a worthy contender.
A second major issue is that language models are weird, and the know-how to successfully implement them into automated workflows remains relatively scarce. Most companies are naturally excited about opportunities to save 90% of their costs by automating tasks that nobody enjoys anyway. But if they don't have or know someone who can execute on such a project, they really have no choice but to wait.
A third issue is that AI agents, which promise to allow people to delegate work to AI on a real-time, more ad hoc basis, have not matured as fast as I thought they might. In part, that's due to ongoing GPU shortages. I've repeatedly said that I think GPT-4 Vision, originally introduced in March, but only now hitting general availability, will boost agent performance. And indeed, we already are starting to see new reports of much better performance at significantly lower cost simply because computer interfaces are designed to be interpreted visually, and GPT-4 Vision now makes that possible. What else is standing in the way? Another challenge is that it turns out that we can boost AI performance significantly by decomposing tasks into subparts. But then, when we try to string long chains of these subparts back together into meaningful work, it becomes very tricky to determine what information to give the AI at each step along the way. Too much information can be unwieldy and in any case makes the agents slower and much more expensive, while too little information leads to bad decisions and overall failure. GPT-4 fine-tuning, which is also still in Early Access Only mode as of now, will probably help here. Fine-tuning in general is very useful for shaping model behavior, and regular listeners will know from the recent Emergency episode on the Mamba architecture that I expect state-space models to deliver more coherent, agentic behavior as well. But even if so, this would pose another challenge. Where are we going to get all the long-episode, context-rich, how-we-work-in-practice sort of data that we're going to need to train these long-context AIs? Where does that data exist today, if at all?

73 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000641003634