**Daniel Faggella** (0:12)
Welcome everyone to the Emerge AI in Business Podcast. Today's guest is Luke Rotta, Director of Site Reliability Engineering at Charles Schwab. Luke joins Emerge's Yolandi de Weerdt on today's episode to examine why modern IT operations struggle to keep pace with incident volume and data flow. He outlines how fragmented tooling, slow contextualization, and human-driven triage create the real drag on response speed. His perspective highlights the operational shift leaders must make to move from reactive firefighting to more intelligent, reliable enterprise operations. Today's episode is sponsored by BigPanda. Just a quick note for our audience that views expressed by Luke Rotta on today's program do not reflect that of Charles Schwab or its leadership. Emerj works with a select group of AI vendors to reach Fortune 500 decision makers through research, media and direct access. If you want to be considered, download our media kit at emerj.com. slash ad one. That's emerj.com. slash ad and the number one. Now the conversation with Luke.
**Yolandi de Weerdt** (1:32)
Luke, it's great to have you in our studio today.
**Luke Rotta** (1:34)
Thank you. I really appreciate you having me. I'm excited.
**Yolandi de Weerdt** (1:38)
We had our chat a while ago in buildup of this conversation, and I know that we're in for a treat today, and our listeners as well. And I'm going to hit the ground running here thinking that there is this assumption that the hardest part of IT operations is the technology, getting the different systems to talk to each other, getting the tooling right. But I feel like we keep coming back to that the harder problem seems to be the pace mismatch, how fast incidents arrive versus how fast teams can actually respond to them.
And you've been inside this environment where cost of a slow response is immediate and very real. So for question one, I want to get into that and I want to ask you what pressures inside modern IT operations make it hardest for teams to stay ahead of incidents?
**Luke Rotta** (2:27)
So I think there's several aspects day to day that make it challenging for any first responder in a operational role, whether they're level one or level two or beyond that. And I think the amount of data that is coming in these days into tools is a challenge. I think the number of tools that are being used and are not integrated together is also a challenge. So tools over the last decade or two have been trying to solve this alert noise problem. I still don't really see it being solved all that well, but I think with things, with AI coming into the mix, perhaps it can actually be solved these days. So certainly the amount of data that's coming in, contextualizing that data, for the most part, when an issue occurs, someone is trying to answer a question and getting to that question quicker allows you to solve the problem quicker. And it really comes down to the data that's available and how much time you need to spend researching. And that's going to shorten or lengthen the resolution period.
**Yolandi de Weerdt** (3:30)
That's interesting that you mentioned that one of the challenges would be the number of tools that are not integrated or talking to each other. And this is something that I feel comes up constantly in this type of conversation.
How much of the problem is a too many tools problem? And how much of it is that the tools are not talking to each other? Or is it both? What have you been seeing?
**Luke Rotta** (3:50)
I think it's both. And I think it's actually becoming a larger issue before it's going to kind of reduce. So I think almost every day, someone sends me a message about some new tool, observability tool, that they are advertising, that solves mean time to resolution reduction, that solves alert noise. The messaging is almost the same, except now there's AI in the mix. And what I don't see is someone solving the data problem because I think that's the real issue here, is that the data is not contextualized quicker and in the form of how humans react or respond to incidents.
**Yolandi de Weerdt** (4:38)
That's interesting and it actually sparks something that I feel like there's a whole other conversation to also have.
We've been talking about the context layer and how important it is on the show, but I don't think we've ever touched on how fast the data gets contextualized. That's definitely something that we need to look into. I want us to also just hit on that mean time to resolution that you just brought up. Where does a slow MTTR hurt us the most? Is it customer impact? Is it cost? Is it team burnout? Where do we see that being the biggest problem?
13 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778216552