Redefining Incident Response: Insights from the Chaos Engineer Behind Jeli.io | Nora Jones artwork

Redefining Incident Response: Insights from the Chaos Engineer Behind Jeli.io | Nora Jones

Dev Interrupted

April 25, 2023

If you think your org doesn’t have any incidents, it’s time to change your definition of an incident. This week we’re joined by Nora Jones, Jeli's founder & CEO, to help us make sense of incident analysis and explain why so many incidents go underreported.
Speakers: Nora Jones

Topics: Technology

**Nora Jones** (0:00)
I tell orgs, you should actually over index on calling incidents. A big warning sign is when I talk to a leader and they're like, we don't have incidents. We actually don't have any incidents.
You'd be surprised, I hear this from very large companies that everyone uses. And that scares, I'm like, I know you have more incidents than anyone else, and incidents are surprises. Incidents are anything that pulls you out of your day that you did not prepare for, right? There are always things to learn from.

**SPEAKER_2** (0:25)
At Dev Interrupted, we work to give engineering leaders actionable ways to improve their teams. That's why we're producing a three-part summer workshop series with Linear B. Each of the three workshops will explore the processes that elite software engineering organizations and executives use to deliver better business outcomes and reduce cycle time by 47% on average in just 120 days. You'll learn how to assess your current performance and benchmark it against industry averages, streamline processes through automation, and improve business outcomes through resource allocation. Learn from the best and take your team to the next level.
Visit our website to learn more and secure your spot today.
We're back on Dev Interrupted, and I'm excited to be joined by Nora Jones. Nora is the founder and CEO of jelly.io, an incident management platform, and she's also the ex-head of chaos engineering at Slack. Welcome to the show.

**Nora Jones** (1:18)
Thank you so much for having me. I'm excited to be here.

**SPEAKER_2** (1:20)
I'm really stoked to have you on here because you have a couple of, I would say almost controversial opinions about how engineering works.
One of which is the talk you gave today about leading from incidents and how you think incidents should drive the decision-making for a company. Can you expand on that?

**Nora Jones** (1:37)
Yeah, absolutely.
So incidents and unexpected situations are happening all the time. We are constantly getting surprise in our organizations. You know, we have this plan, right? And we set out to do it, and we have this feature roadmap, and then all of a sudden, a customer is complaining about something, something doesn't look right. We get this weird, spidey sense feeling, and we all jump on and we fix it. And then we have to go back to our roadmaps afterwards, or back to the things that we had promised to be delivered, and we kind of just leave all that happened there. And there's so much data hiding in that. There's so much investment that we can get out of, you know, this thing we already spent a bunch of money on, because it took a lot of time and energy to spend money from the incident. And there's just, there's a lot of value in understanding how people talked to each other during an incident, who they brought in, what they were looking at. Like all that data can be used to understand how your organization works and also help create this learning culture.
So that was kind of what my talk was about today. And then like you said, I have many controversial engineering opinions.

**SPEAKER_2** (2:45)
I'm excited. Let's dig in a bit more though. I'd love to talk through an example, if you have one that comes to mind about how a company can leverage an incident to really learn and grow out of it.

**Nora Jones** (2:55)
Yeah. So I talked about one pretty extensively in my talk today, but I think the real data is hiding in your seemingly innocuous incidents. You know, the ones where you went, oh, phew, no one noticed that. Like it wasn't a step one. Yeah, we're all good, right?
That's actually where you can get the most learnings because emotions aren't as high afterwards. When you have a very impacting incident, like there was a data dog one recently, right?
And I guarantee you things were emotional and turn internally there afterwards, like 100%. And if they're not practicing every day, you know, the seemingly innocuous events, the ones that take five minutes, the ones that no one noticed, like it's gonna be hard to glean data from those large scale emotional ones. And so an example I would take would be find an incident in your organization that didn't take very long compared to your normative standards, but maybe took a lot of people in your org to help fix.
And I would ask you, like, did you actually take the chance to learn from that incident, to do a review on it, to interview the people that participated in it? That's an example I would use. And that's one I go over in my talk today. It's like a seemingly innocuous search incident. And then, you know, the team did an incident review afterwards and I'll put it in quotes, but they really just kind of went through a checklist and then got back to work. And it just, it wasn't that valuable to the organization. And so in the talk I give, I kind of take an alternate approach to doing that.

34 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID