Topics: Technology
**Dan Lines** (0:00)
Hey, everyone, welcome to Dev Interrupted. I'm Dan Lines, and today's topic is site reliability engineering. As one of the most wanted roles by engineering orgs around the globe, I brought on my friend and SRE manager at G-Research, Brian Murphy, aka Murph, to discuss everything we need to know about SRE.
Murph, thanks for coming on the show.
**Brian Murphy** (0:24)
Yeah, of course, Dan, no problem.
**Dan Lines** (0:27)
Yeah, awesome to have you.
Before we jump right into our SRE discussion, I usually ask a few pieces of context about you and your role and your background and your org. So what do you got going on now?
**Brian Murphy** (0:43)
All right, so let's see. Some background. I was a developer. I did Java. I did Python. I did all that stuff. I felt a desire to do more things operational.
So I really started getting into more of the operation side.
And then there was an opportunity to do an SRE role. So I started out doing SRE for a startup back in the day. You're well aware of Dan.
**Dan Lines** (1:13)
We both worked at that startup.
**Brian Murphy** (1:15)
Yep, exactly. And then from there, it came sort of building out a team, right? Like the SRE impact was being felt. It was very positive. It was very good, build out a team to do more of that.
Yeah, so now I lead another SRE team. Slightly bigger, certainly a lot more experience. I have a bunch of like Zooglers and ex-Googlers and stuff like that. So really pushing observability, monitoring metrics, alerting, all of those types of things.
**Dan Lines** (1:49)
Yeah, that's great.
**Brian Murphy** (1:50)
Resident management, all of those.
**Dan Lines** (1:54)
Can you say anything about what G-Research does?
**Brian Murphy** (1:58)
Yeah, of course. So G-Research is a software firm, right? That's what we do. We build software. We just happen to build software for financial trading, quantitative financial trading.
So we essentially build out a platform such that these, I'm gonna call them math nerds, but they're really much smarter than that. And they're actually really good and decent people for the most part, the ones I know.
They pull together these algorithms and they're like, if popcorn goes up and Sony goes down, buy Cheez Whiz. It makes no sense to me. I literally cannot fathom any of it, but it works. And it works well over time. One of the things about financial trading is that reliability is key, right? Because if you're not in the market, if your software is down and you're not in the market, it means you're not buying and you're not selling.
**Dan Lines** (2:57)
Not making money.
**Brian Murphy** (2:59)
That's not making money.
That's not good. So you always want to be in the market. You always want to be buying and selling.
**Dan Lines** (3:06)
Yeah, actually that's great context. So let's get into some of the SRE stuff. I'll give a little overview that our producer, Nico, wrote for me. So if something sounds wrong, we can blame him. But here's a little history of SRE, and then we'll check with you to see what you think of this. So this type of role has probably been around for decades before that, but called something like production ops, disaster recovery, prod, testing, monitoring type things. SRE started to rise with cloud computing and all of that started to take shape. Engineers really needed to be able to work in production.
The role then became increasingly complex as orgs transitioned from large monolithic infrastructures to distributed microservices.
And maybe that's where we're at today. Does that cover how you think of SRE or how do you think of it?
**Brian Murphy** (4:14)
Yeah, there's, it's funny. Like you can describe it in exactly that fashion, but there was a quote from one of the, from Ben Treanor, and I'm gonna get his name wrong, but Ben Treanor, where he said, SRE is essentially like throwing a software engineer at an operations problem, right? Cause you come from that, like developer mindset, that design and think about all of these things. So think about it as a developer, but apply it to an operational type of problem.
**Dan Lines** (4:48)
Yeah.
**Brian Murphy** (4:48)
And I think it all encapsulates into that.
**Dan Lines** (4:51)
And who is Ben?
**Brian Murphy** (4:53)
He was like the original SRE at Google. He was like the founder, like, I don't know if you want to call him the father of SRE or whatever, but he's certainly the founder of that.
**Dan Lines** (5:05)
The overlord of SRE.
**Brian Murphy** (5:07)
Yes, probably, yes.
**Dan Lines** (5:08)
19 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID