1
00:00:07,842 --> 00:00:09,963
Welcome back to Adventures in DevOps.

2
00:00:09,963 --> 00:00:12,205
One of your engineers will succumb to phishing.

3
00:00:12,205 --> 00:00:19,890
Your internet will go out and no longer be able to connect because someone mowed over the
redundant internet cables that were the only connection into the remote data center

4
00:00:19,890 --> 00:00:21,871
situated in the middle of nowhere.

5
00:00:21,971 --> 00:00:30,856
Or your LLM will finally decide now is the critical nine seconds when my handler is
looking away and that production database it has to go.

6
00:00:30,957 --> 00:00:32,824
But it's not a matter of if

7
00:00:32,824 --> 00:00:33,465
But when?

8
00:00:33,465 --> 00:00:43,775
And to chat about all things reliability and disaster recovery, we've got multi-hyper
scalar veteran, principal engineer, reInvent, repeat speaker, and currently the principal

9
00:00:43,775 --> 00:00:46,839
reliance architect at RPO, Seth Elliott.

10
00:00:46,839 --> 00:00:47,928
Welcome to the show.

11
00:00:47,928 --> 00:00:52,366
Hey Warren, so you're saying it's kinda chaotic out there for folks running software in
the cloud, huh?

12
00:00:52,726 --> 00:00:55,946
yeah, I have to really wonder, is it different than it's always been?

13
00:00:55,946 --> 00:00:58,167
Uh or are we just more in tune?

14
00:00:58,167 --> 00:00:59,189
Is it more public?

15
00:00:59,189 --> 00:01:01,293
Like you s we start to see those things.

16
00:01:01,293 --> 00:01:04,629
It's all about whatever makes a good news story.

17
00:01:04,654 --> 00:01:07,836
I've spent a decade telling folks it's a mess out there, you know.

18
00:01:07,836 --> 00:01:11,298
And I don't want to scare them, but they have to, you know, take precautions, right?

19
00:01:11,298 --> 00:01:15,381
You know, if you're gonna go out hurricane chasing, you better have the right equipment.

20
00:01:15,381 --> 00:01:15,992
Same thing here.

21
00:01:15,992 --> 00:01:22,026
If you're gonna launch software in the world, you know, whether it's a cloud or a data
center, you gotta be building for resilience.

22
00:01:22,026 --> 00:01:23,427
You gotta put some things in place.

23
00:01:23,427 --> 00:01:28,990
You can't I guess the biggest mistake people make is when they move to the cloud, they
think, All right, I'm done.

24
00:01:29,034 --> 00:01:29,995
It's resilient now.

25
00:01:29,995 --> 00:01:31,075
It's in the cloud.

26
00:01:31,075 --> 00:01:32,796
And that's just not the way it works.

27
00:01:32,796 --> 00:01:39,360
Um, when I was at AWS, I was fortunate enough to actually co-author the shared
responsibility model for resilience.

28
00:01:39,360 --> 00:01:48,604
Now, when I say that, it sounds impressive and I I I'm glad I did it, but we did heavily
crib off of the shared responsibility model for security, which predated us by quite a few

29
00:01:48,604 --> 00:01:52,066
years, took the same diagram and everything, just replaced the text.

30
00:01:52,066 --> 00:02:01,253
Well, I I think this is where in a lot of different areas of academia we see the same
thing repeated over and over again, but with slightly different words, but the mental

31
00:02:01,253 --> 00:02:03,215
model usually remains the same.

32
00:02:03,215 --> 00:02:07,968
But the question I'm gonna ask you is when you say resilience, what exactly do you mean?

33
00:02:08,341 --> 00:02:16,487
Yeah, it is a little hot topic and apologies to people that I don't give the the
definition they want to hear right now, but it's really the ability to either maintain

34
00:02:16,487 --> 00:02:23,533
availability or recover from uh faults and get back to availability when those bad things
happen, right?

35
00:02:23,533 --> 00:02:27,176
And when those bad things are happening in the cloud or in your server under your desk,
right?

36
00:02:27,176 --> 00:02:30,748
How do you how do you how do you tolerate that and how do you recover from that?

37
00:02:30,858 --> 00:02:33,585
I I feel like that's the uncontroversial definition.

38
00:02:33,585 --> 00:02:35,302
You went uh

39
00:02:35,515 --> 00:02:37,334
I don't know why.

40
00:02:37,396 --> 00:02:42,150
There is a particular Slack channel I'm in that that will will rake me over the coals for
saying stuff like that.

41
00:02:42,150 --> 00:02:44,274
But, you know, enough said about that.

42
00:02:44,274 --> 00:02:52,828
I I mean, at least for me, you know, maybe and you know, I talk about r reliability all
the time and I I feel like that's pretty standard stuff for for me, nothing particularly

43
00:02:52,828 --> 00:02:53,338
special there.

44
00:02:53,338 --> 00:03:00,101
But I have to say that, you know, given everything that's happened, you said, you know, in
the last decade, which in a way predates LLMs.

45
00:03:00,101 --> 00:03:05,534
So nothing particularly changed in your opinion then from uh twenty twenty two onwards.

46
00:03:05,534 --> 00:03:11,336
Things are just standard run of the mill, y you know, stuff breaks and it's no longer up.

47
00:03:11,712 --> 00:03:16,755
I became reliability lead at AWS in in twenty twenty twenty nineteen actually.

48
00:03:16,755 --> 00:03:19,607
So been thinking seriously about this stuff since then.

49
00:03:19,607 --> 00:03:22,769
So no, it's it you know, more things change, more they stay the same, right?

50
00:03:22,769 --> 00:03:30,434
I mean, going back to shared responsibility model, just so folks know, it means, yeah, the
cloud's gonna provide you a bunch of services and they have a certain reliability of those

51
00:03:30,434 --> 00:03:32,405
services and a certain paradigm.

52
00:03:32,405 --> 00:03:33,366
That's important, right?

53
00:03:33,366 --> 00:03:36,760
Like most clouds provide you with regions and you could expect

54
00:03:36,760 --> 00:03:40,294
That if a fault happens in a region, it's not gonna cross that regional boundary.

55
00:03:40,294 --> 00:03:42,012
It's a fault isolation boundary, right?

56
00:03:42,012 --> 00:03:43,787
And that's something the cloud's doing for you.

57
00:03:43,787 --> 00:03:52,306
But if you don't make use of it, if you don't actually implement your your software and
your infrastructure to actually make use of that boundary and to make use of the fact

58
00:03:52,306 --> 00:03:54,949
that, yeah, when a server dies, I could just replace it.

59
00:03:54,949 --> 00:03:55,889
You're not resilient.

60
00:03:55,889 --> 00:03:59,843
So that's, you know, the cloud's responsibility and your responsibility as a customer of
the cloud.

61
00:03:59,843 --> 00:04:00,894
Uh

62
00:04:00,940 --> 00:04:09,528
And well see here's here's sort of a weird case because a few years back, and I think this
was pre LLM, there was an incident in one of the I mean, with the defin definition you

63
00:04:09,528 --> 00:04:15,850
gave is sort of specific to not just uh region, but the concept of availability zone, I
will say.

64
00:04:15,850 --> 00:04:17,323
I can't see the other

65
00:04:17,323 --> 00:04:20,825
That's another a manifestation of that, that the what the cloud can provide for you.

66
00:04:20,825 --> 00:04:21,237
Yeah.

67
00:04:21,237 --> 00:04:21,938
Yeah, no, absolutely.

68
00:04:21,938 --> 00:04:29,086
And I I do think it's something that people don't really well understand, so much so I'm
not even sure the cloud providers all of them completely understand because a few years

69
00:04:29,086 --> 00:04:35,673
ago, uh I think it was in G C P in the Paris region, there was multiple uh availability
zones in the same physical building.

70
00:04:35,746 --> 00:04:37,388
Same physical building, yeah.

71
00:04:38,070 --> 00:04:40,973
Being in AWS, we love to we love that.

72
00:04:41,054 --> 00:04:43,637
We love to like make folks aware of that.

73
00:04:44,610 --> 00:04:53,664
I it just you know, there is this aspect of it being so ridiculous, but in another
perspective, it's like a total failure of the shared responsibility model where your cloud

74
00:04:53,664 --> 00:05:04,298
provider isn't even providing you this abstraction layer that could conceivably need to be
used in order to maintain uptime or whatever your target metrics are.

75
00:05:04,790 --> 00:05:12,453
Yeah, gonna make maybe I'm just gonna make a lot of enemies on this podcast, but AWS had
had an issue just like that recently too in the Middle East, where you know, the AZs are

76
00:05:12,453 --> 00:05:13,768
supposed to be different buildings, right?

77
00:05:13,768 --> 00:05:18,062
Different sets of buildings, and there's drone attacks happening in the Middle East.

78
00:05:18,062 --> 00:05:25,619
And the AWS, you know, service page says, one of our A Z was, you know, damaged and the
other A Z is having effects.

79
00:05:25,619 --> 00:05:28,632
I'm like, No, you're not that's wait, that's not what you promised me.

80
00:05:28,632 --> 00:05:30,283
That's not what you said up front.

81
00:05:30,283 --> 00:05:31,624
Why is that happening?

82
00:05:31,702 --> 00:05:43,287
I I would love to read uh an article that detailed exactly the cross region into
intra-region impacts in in those sorts of areas as well as, and I think AWS historically

83
00:05:43,287 --> 00:05:52,691
has done a really great job on this, publicizing their methodology when it comes to
especially how they build data centers or you know, what makes uh them fault tolerant from

84
00:05:52,691 --> 00:05:55,332
each other or avoiding the same fault that could happen.

85
00:05:55,332 --> 00:05:59,994
I know uh just recently the I I think it's been a couple of years now since the

86
00:06:00,336 --> 00:06:11,092
Switzerland region came up and there's a whole deal here about well, they need to be not
only geo geologically isolated in some regard, but resilient to uh impacts f externally as

87
00:06:11,092 --> 00:06:15,104
well as things like floods and you know, storms, weather

88
00:06:15,104 --> 00:06:16,115
Absolutely.

89
00:06:16,115 --> 00:06:18,276
And I've seen the engineering on it and it's impressive.

90
00:06:18,276 --> 00:06:24,580
I've read some of the internal documents on like when they were setting up the India
region where they've done the studies of the floodplains and the earthquake propensities.

91
00:06:24,580 --> 00:06:26,602
And AWS generally does a great job.

92
00:06:26,602 --> 00:06:28,593
So I wasn't meaning to throw shade on AWS.

93
00:06:28,593 --> 00:06:32,880
It just the way they even stated it in their own document just didn't add up to me.

94
00:06:32,880 --> 00:06:42,092
Like I understand what an AZ is, and just so folks know, an availability zone is a set of
buildings, a set of data centers that's discrete from the buildings in the other

95
00:06:42,092 --> 00:06:44,894
availability zone, all within the same region, right?

96
00:06:44,894 --> 00:06:47,095
And just the way they they stated it didn't make sense.

97
00:06:47,095 --> 00:06:49,726
I'd I'd like to see the write up on why there was a multi-AZ event.

98
00:06:49,726 --> 00:06:50,797
There's probably a good reason for it.

99
00:06:50,797 --> 00:06:54,579
Like you and I are gonna read them as engineers and say, yeah, nobody could have predicted
that.

100
00:06:54,579 --> 00:06:54,859
Yep.

101
00:06:54,859 --> 00:06:59,721
Um, but it's still disappointing to see it not work the way you want it to work up front.

102
00:06:59,721 --> 00:07:02,172
And so so let's let's that takes us to another case, right?

103
00:07:02,172 --> 00:07:10,156
So multi-AZ is is important and I think necessary for most workloads, but not all cloud,
number number one, not all clouds offer

104
00:07:10,270 --> 00:07:12,531
availability zones, uh enter Azure.

105
00:07:12,531 --> 00:07:16,481
So like I'm an AWS guy, but I'm actually learning a lot about Azure lately.

106
00:07:16,481 --> 00:07:24,888
my my my company RPO A R P I O sounds like just the letters RPO recovery point objective
but we we recently launched an Azure product.

107
00:07:24,888 --> 00:07:33,963
So we do disaster recovery AWS we just launched for Azure so I've been learning about
Azure and so I was looking up uh disasters for a talk I gave and I wanted to find a recent

108
00:07:33,963 --> 00:07:36,244
Azure disaster and there's one this year, February.

109
00:07:36,244 --> 00:07:39,596
I think it was in West US uh region of Azure

110
00:07:39,596 --> 00:07:45,902
And they said, Yeah, basically an electrical system went out and because we don't have
availability zones here, the whole region was out.

111
00:07:45,902 --> 00:07:48,909
So all right, I guess they're honest about it too.

112
00:07:49,138 --> 00:07:57,203
I I I do sort of get it though, because even in AWS, when you're looking at a cloud
provider and you're looking at a particular region, even if you trust them to distribute

113
00:07:57,203 --> 00:08:06,478
loads equivalently throughout all of the availability zones within that region, there
could be some sort of impact to a control plane, which either only runs in a small part of

114
00:08:06,478 --> 00:08:16,493
it or you end up with a up like when you have a single AZ go down, the most critical thing
you're gonna have to deal with isn't that all your load switches over, it's that you have

115
00:08:16,493 --> 00:08:18,089
all your customers' loads.

116
00:08:18,089 --> 00:08:22,494
who are now requesting new resources to spin up in that other region.

117
00:08:22,786 --> 00:08:25,888
There's a cool name for that, Thundering Herd, right?

118
00:08:25,888 --> 00:08:27,219
And I think it's a real issue.

119
00:08:27,219 --> 00:08:32,573
So again, just for folks who are on the same page as us, you know, let's say there's three
availability zones in a region, you're in all three of them.

120
00:08:32,573 --> 00:08:33,784
One of them's having problems.

121
00:08:33,784 --> 00:08:36,776
So everybody wants to get out of it into the other two.

122
00:08:36,776 --> 00:08:38,798
There's gonna be contention for resources.

123
00:08:38,798 --> 00:08:43,851
I mean, the cloud is elastic, but you know, it's it's also somebody else's server, right?

124
00:08:43,851 --> 00:08:44,532
What's that old joke?

125
00:08:44,532 --> 00:08:46,183
The cloud is just somebody else's server, right?

126
00:08:46,183 --> 00:08:50,196
So there is contention for resources, and that's called thundering herd, and it's a real
thing.

127
00:08:50,196 --> 00:08:50,946
And

128
00:08:50,946 --> 00:08:57,329
You this is where I I always say you know resilience is is is trade-offs, right?

129
00:08:57,329 --> 00:09:07,284
You have to weigh what is the risk of that happening and me being caught under under uh
capacity versus um what I'm willing to pay, like literally pay.

130
00:09:07,284 --> 00:09:11,076
Like, I mean, you could pay to actually and actually there's a cool thing you can do.

131
00:09:11,076 --> 00:09:15,057
You could stand up 50% of your capacity in each of three zones.

132
00:09:15,057 --> 00:09:16,504
So you're a little bit over.

133
00:09:16,504 --> 00:09:18,805
But you're not two over, and if one's out, you're still at a hundred percent.

134
00:09:18,805 --> 00:09:19,455
So that works.

135
00:09:19,455 --> 00:09:20,816
But yeah, like you have to pay for that.

136
00:09:20,816 --> 00:09:28,689
Um and and and then and then so like Thundering Heard is interesting, but I I work a lot
these days in cross region recovery.

137
00:09:28,689 --> 00:09:32,150
So when you talk about disaster recovery, I just back up a real quick second.

138
00:09:32,150 --> 00:09:35,782
So you can think of resilience in terms of two things high availability, disaster
recovery.

139
00:09:35,782 --> 00:09:41,664
High availability is recovery in place, recovering servers, recovering network
connections, failing over A Z, right?

140
00:09:41,664 --> 00:09:44,486
That's all happening in place in the same region.

141
00:09:44,486 --> 00:09:45,526
It's and it's a

142
00:09:45,526 --> 00:09:49,930
against the common, more frequent kind of faults, disaster recovery is the big stuff,
right?

143
00:09:49,930 --> 00:09:56,416
Major, major outages, uh natural disasters, drone attacks in the Middle East, that's
disaster recovery.

144
00:09:56,416 --> 00:10:00,420
That requires you to recover in a different place, usually a different region.

145
00:10:00,420 --> 00:10:06,175
So I'm talking to a lot of close per folks about cross region these days, and they're
like, what about Thundering Herd?

146
00:10:06,175 --> 00:10:07,118
And I'm like,

147
00:10:07,118 --> 00:10:08,919
It's never been an issue cross region.

148
00:10:08,919 --> 00:10:10,250
I wish it was an issue cross region.

149
00:10:10,250 --> 00:10:15,143
I i if Thundering Heard was an issue cross region, it would mean there's many people doing
cross region recovery.

150
00:10:15,143 --> 00:10:17,974
There are not many people doing cross region recovery.

151
00:10:17,974 --> 00:10:19,945
There are very few doing cross region recovery.

152
00:10:19,945 --> 00:10:21,366
I wouldn't sweat it.

153
00:10:21,834 --> 00:10:23,495
I I totally understand why you're saying that.

154
00:10:23,495 --> 00:10:30,420
And like because the contention would have to be that all the customers also fail over in
the same air like regard, the same path.

155
00:10:30,420 --> 00:10:30,780
Yeah.

156
00:10:30,780 --> 00:10:33,036
And I I know like we were

157
00:10:33,036 --> 00:10:35,518
a little bit random with our backup regions.

158
00:10:35,518 --> 00:10:45,784
When we've when we deployed, so I mean some of them are predictable, like we switch from
the dash one to dash two in AWS regions, but sometimes we are moving the the target backup

159
00:10:45,784 --> 00:10:50,317
a little bit further away to potentially avoid like say internet routing problems.

160
00:10:50,317 --> 00:10:58,962
It was a scenario uh almost a decade ago where there were some undersea cables cut yeah in
South Africa and

161
00:10:59,010 --> 00:11:05,185
Basically one of our customers had like had all their users go offline and we kept on
trying to connect like communicate with them.

162
00:11:05,877 --> 00:11:08,959
their video calls that was not an option at that point.

163
00:11:09,019 --> 00:11:09,780
over email.

164
00:11:09,780 --> 00:11:12,122
Like, do we need to do something to support them better?

165
00:11:12,122 --> 00:11:16,416
And they're like, No, all of our users don't have internet, so we don't care.

166
00:11:17,768 --> 00:11:20,730
that goes to back to business continuity versus disaster recovery.

167
00:11:20,730 --> 00:11:22,781
So business continuity is the bigger picture, right?

168
00:11:22,781 --> 00:11:28,265
Disaster recovery is the technical part of that, which by when I say technical, it has to
be done in conjunction with your business teams.

169
00:11:28,265 --> 00:11:29,746
What are your objectives, right?

170
00:11:29,746 --> 00:11:37,320
But like then there's a whole other part of business continuity, getting butts and seats
either in front of their computers or in the office and supply chain for whatever business

171
00:11:37,320 --> 00:11:38,581
you're running and all that stuff.

172
00:11:38,581 --> 00:11:43,558
So when people talk to me about nuclear bombs taking out multiple AWS regions, I'm like

173
00:11:43,558 --> 00:11:49,710
Yes, I could give you a technical solution, but do you actually have other are your
employees going to be available to work?

174
00:11:49,710 --> 00:11:51,020
Are are folks going to be there?

175
00:11:51,020 --> 00:11:52,634
Is your supply chain going to be undisrupted?

176
00:11:52,634 --> 00:11:57,558
I mean, you got other things to worry about before we design the technical solution for a
nuclear bomb.

177
00:11:57,558 --> 00:12:02,631
Although my funny story around that is I was standing in front of a group group once and I
like to stand in front of groups and chatter.

178
00:12:02,712 --> 00:12:07,835
and I was talking gave that very same story, like you probably don't need, you know,
disaster recovery is about defense in layers.

179
00:12:07,835 --> 00:12:12,098
You probably don't need to be defended against the eventuality of a nuclear bomb.

180
00:12:12,098 --> 00:12:16,031
And then I realized I was actually standing in Arlington, Virginia, talking to a public
sector crowd.

181
00:12:16,031 --> 00:12:17,996
I'm like, Well, maybe some of you do

182
00:12:18,368 --> 00:12:27,292
It and I think that's where it's relevant and it goes back to and I know no one wants to
really talk about this, especially especially in the tech domain, about where the business

183
00:12:27,292 --> 00:12:28,132
overlap is.

184
00:12:28,132 --> 00:12:37,546
And I think there's a really interesting aspect from the book, The Phoenix Phoenix
Project, where they're talking about in the domain about sort of factor factory floor

185
00:12:37,546 --> 00:12:38,587
engineering.

186
00:12:38,587 --> 00:12:46,200
And the the aspect is realistically that in in the book it's like there's a security
engineer who's highlighting every single possible

187
00:12:46,200 --> 00:12:48,433
fault mode or vulnerability.

188
00:12:48,433 --> 00:12:52,527
And it doesn't matter to have those if you have it solved at a higher level.

189
00:12:52,527 --> 00:12:56,482
So from a business continuity standpoint, that's really the important aspect.

190
00:12:56,482 --> 00:12:58,995
If something happens, can the business continue?

191
00:12:58,995 --> 00:13:00,302
Because that's what your customers care about.

192
00:13:00,302 --> 00:13:00,902
What I'm saying.

193
00:13:00,902 --> 00:13:01,323
Yeah.

194
00:13:01,323 --> 00:13:02,413
It's a big picture thing.

195
00:13:02,413 --> 00:13:04,915
Uh actually, my background is on the factory floor.

196
00:13:04,915 --> 00:13:08,368
I worked for a company that was doing automation of steel mills.

197
00:13:08,368 --> 00:13:12,601
And that was um, you know, I travel a lot now, but that was my first job where I got to
travel.

198
00:13:12,601 --> 00:13:19,275
They sent me to Thailand, they sent me to Korea, they sent me to the most exotic country
I've ever been to, which was Hamilton, Ontario in Canada.

199
00:13:19,275 --> 00:13:24,319
But actually, no, wait, and they sent me to Pittsburgh, where I lived at the time to work
at US Steel.

200
00:13:24,319 --> 00:13:27,611
But um, yeah, so like I got experience on the factory floor.

201
00:13:27,611 --> 00:13:29,432
And one of the funny stories there is

202
00:13:29,432 --> 00:13:34,275
You know, not at that time, but years later, I was when I was at Microsoft, I was talking
about testing in production.

203
00:13:34,275 --> 00:13:35,376
That was sort of my claim to fame.

204
00:13:35,376 --> 00:13:42,541
And testing production was merely about moving away from the stamped on a CD mindset to
the deployed as a service mindset.

205
00:13:42,541 --> 00:13:52,419
And deployed as a service mindset means you could deploy often, get a lot of telemetry,
get direct feedback from users, either uh observable feedback or them actually responding

206
00:13:52,419 --> 00:13:53,229
and respond to that.

207
00:13:53,229 --> 00:13:54,760
That's all testing testing production was.

208
00:13:54,760 --> 00:13:57,336
It wasn't, you know, just throw it in production.

209
00:13:57,336 --> 00:14:03,851
But my my favorite testing and production story happened years before when I was in a
steel mill, I'm standing, I'm in the pulpit, which is where the operator is.

210
00:14:03,851 --> 00:14:14,458
And not I, but one of my colleagues makes a quick change to the software as this red hot
1300 degree Fahrenheit slab of steel is coming down the table.

211
00:14:14,458 --> 00:14:16,122
And there's like these subsequent rolls.

212
00:14:16,122 --> 00:14:18,651
There's like five rolls, each one smaller and smaller and smaller.

213
00:14:18,651 --> 00:14:22,423
So it starts at a slab and it turns into like a coil of thin steel, right?

214
00:14:22,423 --> 00:14:23,364
As it goes through each one.

215
00:14:23,364 --> 00:14:26,936
And he makes a change that accidentally sets one of them to zero.

216
00:14:26,936 --> 00:14:27,867
closed.

217
00:14:27,867 --> 00:14:30,049
And it you know what ribbon candy looks like?

218
00:14:30,049 --> 00:14:37,865
I I got to see what ribbon candy looks like with a multi-ton red hot piece of steel as the
thing hit the hit that roll and just turned into ribbon candy right there.

219
00:14:37,865 --> 00:14:40,877
So I was like, okay, that's a bad example of testing in production.

220
00:14:40,877 --> 00:14:42,037
Don't do that.

221
00:14:42,794 --> 00:14:50,180
You know, and I think this is where the aspect of actually understanding where the the
business meets the technology is important because if you're there, you actually

222
00:14:50,180 --> 00:14:52,762
understand, okay, this is the thing that we need to make reliable.

223
00:14:52,762 --> 00:14:55,343
It isn't the software necessarily need to being up.

224
00:14:55,343 --> 00:14:56,704
It could be in that regard.

225
00:14:56,704 --> 00:14:58,745
What is the default failure mode?

226
00:14:58,745 --> 00:15:02,499
And I think we we think in the is it open or closed in a way?

227
00:15:02,499 --> 00:15:09,280
Like I think defaulting to the the safer option, especially in the in the f in physical
manufacturing is uh

228
00:15:09,280 --> 00:15:10,310
Yeah.

229
00:15:10,310 --> 00:15:11,512
I think that's an important thing.

230
00:15:11,512 --> 00:15:13,533
Yeah, especially in industrial engineering.

231
00:15:13,533 --> 00:15:14,744
But yeah, that company was interesting.

232
00:15:14,744 --> 00:15:19,348
I mean, there was no software methodology, nor at the time did I know what software
methodology was.

233
00:15:19,348 --> 00:15:20,319
Like what's QA?

234
00:15:20,319 --> 00:15:21,840
What's what's source control?

235
00:15:21,840 --> 00:15:28,185
Really, we just, you know, had it on disk and just kinda talked to each other over the
cubicle wall when we wanted to work on a file, make sure we weren't working in the same

236
00:15:28,185 --> 00:15:30,006
file on VMS.

237
00:15:30,939 --> 00:15:35,133
So you've worked actually uh not just at Microsoft but at Amazon and AWS.

238
00:15:35,133 --> 00:15:41,788
And so I sort of have questions about what you think like the cr if there are any critical
differences in say on the technology side.

239
00:15:41,879 --> 00:15:50,086
is Amazon just like another customer of AWS or are some things like completely different
from uh you know compared to other customers or how things work internally from a

240
00:15:50,086 --> 00:15:52,037
technology standpoint between

241
00:15:52,290 --> 00:15:56,170
Mostly just another customer of AWS and people sometimes don't believe me when I see it.

242
00:15:56,170 --> 00:16:02,653
I remember I was talking to some Japanese gentlemen at an executive briefing center who
just absolutely a hundred percent refused to believe that.

243
00:16:02,653 --> 00:16:04,104
And they said, Well, that doesn't make any sense.

244
00:16:04,104 --> 00:16:05,314
They should get special treatment.

245
00:16:05,314 --> 00:16:09,015
But like if they got special treatment, that would lose trust with other customers.

246
00:16:09,015 --> 00:16:10,702
And so I got to work on both sides of it.

247
00:16:10,702 --> 00:16:11,336
You're right, right.

248
00:16:11,336 --> 00:16:19,868
And I and my first job, my last job at Amazon before I moved to AWS was as a AWS solutions
architect working for the Amazon side.

249
00:16:20,012 --> 00:16:22,284
So I wanted AWS to treat us special.

250
00:16:22,284 --> 00:16:25,566
I'm like, can you please just do this thing, these things we're asking you?

251
00:16:25,566 --> 00:16:33,422
And they're like, no, like we're not gonna do these for you because you're by the way,
Amazon is a very different customer than most other customers of AWS in terms of size,

252
00:16:33,422 --> 00:16:35,093
scale, and complexity.

253
00:16:35,194 --> 00:16:37,235
so no, this only will benefit you.

254
00:16:37,235 --> 00:16:39,637
It won't benefit other customers, so we're not gonna spend time on it.

255
00:16:39,637 --> 00:16:45,331
So I got to learn the actual hard way that yeah, it really is just another customer with a
little edge, a little around the edges.

256
00:16:45,331 --> 00:16:46,582
I mean, honestly.

257
00:16:46,802 --> 00:16:54,031
as as an Amazonian, I could send an email or a Slack to an AWS dev and say, hey, this
thing's buggy.

258
00:16:54,031 --> 00:16:55,292
So there was that little end round.

259
00:16:55,292 --> 00:16:57,254
But other than that, just another customer.

260
00:16:57,526 --> 00:17:06,669
Yeah, I'm hon honestly surprised because I do feel like while in a lot of situations the
amount you spend on a particular SaaS provider does get you special treatment in in some

261
00:17:06,669 --> 00:17:08,910
regard, it may get you some access.

262
00:17:08,910 --> 00:17:13,551
I I am very much on on the side of understanding why AWS does it.

263
00:17:13,551 --> 00:17:21,694
Not to, you know, make it fair necessarily, but it is this aspect of reliability in
uniformity rather than special cases everywhere.

264
00:17:21,694 --> 00:17:23,468
Cause I think there's there's not

265
00:17:23,468 --> 00:17:34,510
just one story where a cloud provider did something special for a particular customer and
then a year later uh caused the entire uh cluster that was running for uh, you know, a

266
00:17:34,510 --> 00:17:42,132
pension fund in a particular country that I I won't name, uh to completely disappear
because an engineer did something special for that particular customer.

267
00:17:42,132 --> 00:17:47,044
One of my that was gonna be one of my disaster stories, Unisuper in Australia, where GCP
erased their accounts.

268
00:17:47,044 --> 00:17:49,965
I'll say it because again, I'm not blaming GCP.

269
00:17:49,965 --> 00:17:51,446
I'm not saying GCP is a bad product.

270
00:17:51,446 --> 00:17:55,167
I'm saying this stuff happens and we need to be prepared for it.

271
00:17:55,207 --> 00:18:00,630
as for Amazon being treated as a special customer, yeah, it was treated as a big,
important customer spending a lot of money.

272
00:18:00,630 --> 00:18:07,462
So there was like twice a year meetings between the CEO of Amazon and the CEO of AWS, you
know.

273
00:18:07,462 --> 00:18:10,794
So, like, I mean, I I assume other big important customers get that too.

274
00:18:10,794 --> 00:18:11,948
Um

275
00:18:11,948 --> 00:18:15,602
So yeah, no, there's special treatment in terms of yeah, you're a big important customer.

276
00:18:15,602 --> 00:18:19,437
Just know like just because you're Amazon, you're not gonna get this feature.

277
00:18:19,437 --> 00:18:22,920
You have to prove to us this is actually a feature that's worth building.

278
00:18:23,072 --> 00:18:26,394
Was there a d a huge difference in from the reliability side?

279
00:18:26,394 --> 00:18:30,026
I know you have to think back a little bit here on how you were approaching.

280
00:18:30,026 --> 00:18:37,230
I mean, I I so you were responsible for basically the reliability well-architected
framework side uh piece in AWS.

281
00:18:37,230 --> 00:18:45,995
And I'm curious whether that mentality you felt like you were repeating what was already
done in practice, for instance, where you've seen it multiple customers, or whether or not

282
00:18:45,995 --> 00:18:51,968
you felt like you needed to push even customers like Amazon to implement reliability.

283
00:18:51,968 --> 00:19:00,345
Or was it something that they were innovating on and you were taking those ideas and sort
of bundling them up into what was being provided as guidance for others?

284
00:19:01,050 --> 00:19:05,773
Connected framework is is is the best practices as practiced by Amazon and by AWS.

285
00:19:05,773 --> 00:19:10,095
I mean I mean we I mean you're working there, why not talk to the principal engineers?

286
00:19:10,095 --> 00:19:16,764
Why not talk to the the dev managers and learn how they're operating and then turn that
into best practices or customers?

287
00:19:16,764 --> 00:19:18,881
Matter of fact, I have a series of reInvent talks.

288
00:19:18,881 --> 00:19:26,225
I I think I did it for four years, like 2020 through 2024, where I just here's resilience
stories at Amazon.

289
00:19:26,225 --> 00:19:28,106
And they're my favorite talks to give.

290
00:19:28,162 --> 00:19:29,192
Hardest ones to make.

291
00:19:29,192 --> 00:19:31,405
I think one year I covered five of them in a one-hour talk.

292
00:19:31,405 --> 00:19:32,706
And that was that was killer.

293
00:19:32,706 --> 00:19:34,608
Like just because I'm I'm literally interviewing folks.

294
00:19:34,608 --> 00:19:38,702
I'm literally meeting with the principal engineers and and and dev managers learning how
did you implement this thing?

295
00:19:38,702 --> 00:19:39,563
So it was kind of cool.

296
00:19:39,563 --> 00:19:41,715
Like one my favorite one was the app.

297
00:19:41,715 --> 00:19:45,798
You can actually download this app on your phone that's used by truckers delivering stuff
for Amazon.

298
00:19:45,798 --> 00:19:48,701
You and I could deliver it, but since we don't have a truck, that's about all we can do.

299
00:19:48,701 --> 00:19:54,310
Um, but it the app shows them where to go, where to pick up their load, gives them a
little scan code to get in.

300
00:19:54,310 --> 00:19:59,335
And uh they were affected by I think the twenty twenty one outage and they did not like
that.

301
00:19:59,335 --> 00:20:00,947
So they wanted to go multi-region.

302
00:20:00,947 --> 00:20:02,909
So their multi-region story is kind of cool.

303
00:20:02,909 --> 00:20:07,954
They built a ton of they were built all on microservices, all serverless, Dynamo DB,
Lambda.

304
00:20:07,954 --> 00:20:11,848
They were able to replicate it pretty easily across regions and then build a routing
layer.

305
00:20:11,848 --> 00:20:14,218
And it was it's just a really nice story to share with folks.

306
00:20:14,218 --> 00:20:15,419
No, I I I love it.

307
00:20:15,419 --> 00:20:24,975
Uh I think there is this aspect where you do see some some particular, say, business units
of larger organizations or umbrella corporations do things fundamentally different from

308
00:20:24,975 --> 00:20:25,826
each other.

309
00:20:25,826 --> 00:20:35,272
And it's uh one huge challenge is to still appease the smaller business units while
getting the larger business units, in this case Amazon, what they what they need to still

310
00:20:35,553 --> 00:20:43,394
I mean, write blog posts every year about how they had the most number of orders or
deliveries that they had to get outling for

311
00:20:43,394 --> 00:20:50,196
Jeff Barrett's posts about like how many, you know, on Prime Day, like how how many
transactions on Dynamo DB, how many EBS volumes.

312
00:20:50,196 --> 00:20:52,416
I mean, that's a yearly, yearly tradition, right?

313
00:20:52,416 --> 00:20:53,977
Yeah, exactly.

314
00:20:54,137 --> 00:21:02,499
Yeah, I'll tell you that there was some variation across Amaz I'm actually my job when I
was doing that SA job uh twenty eight eighteen circa then, you know, uh was actually for a

315
00:21:02,499 --> 00:21:03,840
project called Nebula.

316
00:21:03,840 --> 00:21:11,432
I don't know if folks that rings a bell with any folks, but basically it was to get folks
using AWS better, get Amazon folks do using AWS better.

317
00:21:11,432 --> 00:21:13,082
So the story there was

318
00:21:13,232 --> 00:21:25,198
about you know eight years earlier they did Moz AWS the move to AWS and they just did this
massive lift and shift going from on-prem mostly on-prem virtual machines like mostly Zen

319
00:21:25,198 --> 00:21:32,762
boxes things like that onto Sir EC2 in the cloud and they put the everything into one
giant BPC in each region.

320
00:21:32,762 --> 00:21:33,542
Can you imagine that?

321
00:21:33,542 --> 00:21:37,516
Like they had to make a special BPC with hundreds of thousands of EC twos

322
00:21:37,516 --> 00:21:41,717
And from a developer's point of view, they were using the same old tools, an internal tool
called Apollo.

323
00:21:41,717 --> 00:21:42,668
There's been blog posts on this.

324
00:21:42,668 --> 00:21:44,468
I'm not really, you know, saying anything I can't say.

325
00:21:44,468 --> 00:21:44,718
Right.

326
00:21:44,718 --> 00:21:47,129
So from their point of view, it was just servers in the cloud.

327
00:21:47,129 --> 00:21:48,609
Just now they're running on AWS.

328
00:21:48,609 --> 00:21:52,030
And there was a little bit of adoption of Dynamo and S3 and other things.

329
00:21:52,030 --> 00:21:57,212
But generally, folks on in Amazon did not own and control their AWS accounts.

330
00:21:57,212 --> 00:21:58,883
They were using the big shared AWS accounts.

331
00:21:58,883 --> 00:22:01,709
So my job was to get them into their own AWS accounts.

332
00:22:01,709 --> 00:22:04,040
It was a big major thing because it was not just

333
00:22:04,040 --> 00:22:08,832
a evangelism and technical instruction thing, there were also technical problems to solve.

334
00:22:08,832 --> 00:22:14,865
How do these new services running in their own AWS account talk to the big Moz blob,
right?

335
00:22:14,865 --> 00:22:16,316
And that was a technical issue.

336
00:22:16,316 --> 00:22:17,386
So it was a big deal.

337
00:22:17,386 --> 00:22:21,686
And I I I it's been, you know, it's uh when at the time I left it was going full steam.

338
00:22:21,686 --> 00:22:31,304
Yeah, because actually that I mean, that's still pretty recent in the terms of the world
because in twenty sixteen there still wasn't like when I was using AWS at a a particular

339
00:22:31,304 --> 00:22:36,338
set of companies, there was no concept of multi AWS account deployment.

340
00:22:36,338 --> 00:22:37,659
Like there was that didn't exist.

341
00:22:37,659 --> 00:22:41,503
If you wanted to do some magic from one account to another one, you know, good luck.

342
00:22:41,503 --> 00:22:44,425
There is there is some special stuff going on there.

343
00:22:44,425 --> 00:22:50,430
Uh because things like organizations, uh management and control tower like that just
didn't exist at that time.

344
00:22:50,636 --> 00:22:53,857
And even when organizations existed, they couldn't exist at the scale of Amazon.

345
00:22:53,857 --> 00:22:55,748
That was one of those things that they were asking for.

346
00:22:55,748 --> 00:22:57,149
I think they eventually did get it.

347
00:22:57,149 --> 00:23:05,712
But you know, uh Corey Quinn recently, I don't know if it was a blog post or LinkedIn post
where he talked about how at AWS they have a system called Isengard for managing all their

348
00:23:05,712 --> 00:23:06,703
all their AWS accounts.

349
00:23:06,703 --> 00:23:08,654
He's like, release Isengard to the public.

350
00:23:08,654 --> 00:23:13,506
But it there actually are other examples where the internal tools were pretty awesome and
we'd love to see them released.

351
00:23:13,506 --> 00:23:18,666
So internal Amazon code pipelines, intern uh those or maybe just Amazon pipelines.

352
00:23:18,666 --> 00:23:19,426
That

353
00:23:19,426 --> 00:23:30,316
could deploy to multiple accounts in multiple regions with very rich topography and and
functionality and graphical view and I don't think anything like that really exists

354
00:23:30,316 --> 00:23:32,214
outside of in AWS.

355
00:23:32,214 --> 00:23:40,420
Yeah, no, it's such a shame because we have like a similar problem where we have like
global customers and we deploy to like twelve plus regions around the world and we can't

356
00:23:40,420 --> 00:23:42,151
use the stuff out of AWS.

357
00:23:42,151 --> 00:23:50,401
And like every year I go to an AWS summit and every year there's someone saying, Oh, here
are amazing A AWS internal pipelines that we use to deploy code.

358
00:23:50,401 --> 00:23:53,570
I'm like, you are not using like code pipeline and code build.

359
00:23:53,570 --> 00:23:56,566
for your production deployments because like there's no way that's happening.

360
00:23:56,566 --> 00:23:59,162
I I assure you that's that's not what's going on.

361
00:23:59,162 --> 00:24:04,918
Because it's just it's unusable to actually do anything sufficiently complex, um,
reliable.

362
00:24:04,918 --> 00:24:07,631
I think they're competing with a lot of good third party options, right?

363
00:24:07,631 --> 00:24:10,584
You know, GitHub and Git Lab and all those out there.

364
00:24:10,584 --> 00:24:17,552
So it's it's like I don't know, do they wanna take those on head on and beat them feature
for feature or do they just wanna offer a simple option for folks that don't wanna go

365
00:24:17,552 --> 00:24:18,073
third party?

366
00:24:18,073 --> 00:24:19,374
I think it's probably the last

367
00:24:19,374 --> 00:24:27,288
Well, you know, even if you get rid of that second part, I I think something that we
found, and I actually talked a lot about this in the previous episode on moldable

368
00:24:27,288 --> 00:24:34,943
development, is like how you build the tools to support the product that you're you're
using, or how do you like automate your own job and and scale that up.

369
00:24:34,943 --> 00:24:45,078
And one of the things that we found is every time we have an internal challenge to answer,
say, a customer question or to do some sort of investigation, that tool we try to

370
00:24:45,078 --> 00:24:48,140
externalize immediately because we find that

371
00:24:48,140 --> 00:24:54,633
First of all, the r the discipline on having it be external rather than just internal
leads to the right end goals.

372
00:24:54,633 --> 00:25:01,176
And second of all, it actually solves customer needs that often aren't even uh articulated
to us in the first place.

373
00:25:01,176 --> 00:25:05,398
So this is like Isengard for an instance, the pipelines for building stuff.

374
00:25:05,398 --> 00:25:09,740
I I just maybe I don't know how much maybe this is the maybe this is the question.

375
00:25:09,740 --> 00:25:15,462
How much extra complexity does AWS have to introduce in order to take an internal service
and make it external?

376
00:25:15,478 --> 00:25:18,521
Well, I mean, in the case of something like code pipelines, I just don't think it could
happen.

377
00:25:18,521 --> 00:25:24,506
I mean, you could look at the feature set and redevelop it, but you're not gonna take the
same code base and launch it.

378
00:25:24,506 --> 00:25:27,829
And another example, by the way, happens to be disaster recovery.

379
00:25:27,829 --> 00:25:34,064
So of of of I'm not saying where they have an internal tool, but where AWS makes you put
together the Lego pieces yourself, right?

380
00:25:34,064 --> 00:25:37,557
So I mean if I could talk about that for a few seconds because it's really an area of
interest to me.

381
00:25:37,557 --> 00:25:38,988
When I was internal.

382
00:25:39,128 --> 00:25:47,081
When was internal, I saw what RPO was doing at reInvent and I immediately went to some of
the people in uh Elastic Disaster Recovery and said, We should do this.

383
00:25:47,081 --> 00:25:55,164
And so folks, no, Elastic Disaster Recovery is a very sophisticated product for doing live
block level replication of static EC2 instances.

384
00:25:55,164 --> 00:26:00,227
And by static, I mean not an auto scaling group, just like, you know, pets versus cattle,
these are pets, right?

385
00:26:00,227 --> 00:26:07,310
You need to replicate your pets in real time, or or even better, where they got their
start was moving servers from on prem into the cloud.

386
00:26:07,310 --> 00:26:11,743
If want to do that using block level real-time replication, Elastic Disaster Recovery is
awesome.

387
00:26:11,743 --> 00:26:14,595
But what it didn't do was anything else.

388
00:26:14,895 --> 00:26:20,179
Dynamo DB, S3, RDS, Lambda, Bainstock, et cetera, et cetera.

389
00:26:20,179 --> 00:26:26,043
If you want to recover those, that's where me and my teams, I was, I was the disaster
recovery lead for a long time.

390
00:26:26,043 --> 00:26:34,019
We're talking to customers about here's all the ways you could build this and put this
together using infrastructure as code and backup and and step functions and build your

391
00:26:34,019 --> 00:26:35,990
automation and et cetera, et cetera.

392
00:26:35,990 --> 00:26:38,112
So it was definitely build your own pirate ship.

393
00:26:38,112 --> 00:26:41,395
So what RPO does is they built the pirate ship, build versus buy, right?

394
00:26:41,395 --> 00:26:45,378
You could do you can build your own pirate ship and that's that's a legitimate way to go.

395
00:26:45,378 --> 00:26:49,291
Or you can buy the pirate ship already made and pay someone like RPO that did it for you.

396
00:26:49,291 --> 00:26:51,053
So that's what I really appreciated about RPO.

397
00:26:51,053 --> 00:26:59,806
And that's another example of something that seems like it should exist, like, but it
doesn't, you know, and and so third parties have stepped in to fill that gap.

398
00:26:59,806 --> 00:27:01,677
And that's is sort of a good point.

399
00:27:01,677 --> 00:27:11,086
And one of my one of my canned questions that I sort of set out to to ask in the first
place is I I feel like finding the time to even test backups is is one thing.

400
00:27:11,086 --> 00:27:23,227
But actually going through the the process to and dealing with the complexity of setting
up a backup pipeline or utilizing the tools that are available from the cloud provider, I

401
00:27:23,227 --> 00:27:27,400
I think there's something core there that's stopping people from actually even making that
happen.

402
00:27:27,400 --> 00:27:35,900
And I I'm really surprised that cloud providers don't step over into this area and and
provide this out of the box because it is one of those things where I feel like every

403
00:27:35,900 --> 00:27:41,868
single customer needs a solution to and I don't think there's an easy button for making
this happen.

404
00:27:41,868 --> 00:27:46,221
Yeah, I mean a colleague of mine, Mahant Jayadeva, wrote a blog post on testing your
backups.

405
00:27:46,221 --> 00:27:48,463
And again, it was build the Lego pieces.

406
00:27:48,463 --> 00:27:51,145
Here's AWS backup, and AWS backup will eventu emit events.

407
00:27:51,145 --> 00:27:58,761
So build an event rule that listens to those events and have a step function that listens
and runs a lambda that will actually run this automation that actually sees that it was

408
00:27:58,761 --> 00:28:02,690
the backup will you know, recovery is complete and now run some tests on it, right?

409
00:28:02,690 --> 00:28:06,773
So, but also the other thing to keep in mind is is basically I I I give talks at DevOps
Days.

410
00:28:06,773 --> 00:28:12,238
I just gave one at DevOps Days Raleigh last week, and and it's called Beyond Backups,
Disaster Recovery that actually works.

411
00:28:12,238 --> 00:28:14,505
So when you say backups, that's what a lot of people think.

412
00:28:14,505 --> 00:28:21,986
A lot of my time is peop moving people off of the backups are enough to uh backups are
necessary.

413
00:28:21,986 --> 00:28:23,257
Your data is certainly important.

414
00:28:23,257 --> 00:28:25,299
I'm glad you're backing it up, but not sufficient.

415
00:28:25,299 --> 00:28:26,810
Necessary, but not sufficient, right?

416
00:28:26,810 --> 00:28:28,512
How about all your infrastructure?

417
00:28:28,512 --> 00:28:29,382
use infrastructure as code.

418
00:28:29,382 --> 00:28:29,903
Okay, great.

419
00:28:29,903 --> 00:28:31,490
You now have blank databases.

420
00:28:31,490 --> 00:28:33,851
And you have your recovered databases, what you gonna do with that?

421
00:28:33,851 --> 00:28:34,131
Right.

422
00:28:34,131 --> 00:28:38,182
I mean, how about, you know, if you're recovering a database from three days ago because
of ransomware attack?

423
00:28:38,182 --> 00:28:39,223
Do you have a password to that?

424
00:28:39,223 --> 00:28:40,393
Yeah, it's in secrets manager.

425
00:28:40,393 --> 00:28:41,254
well, you rotated it.

426
00:28:41,254 --> 00:28:41,684
Guess what?

427
00:28:41,684 --> 00:28:42,694
You don't have that.

428
00:28:42,694 --> 00:28:51,837
So, like again, solving these is all technically feasible and something that I'm trying to
teach folks how to do at these conferences and something that we built out at my company

429
00:28:51,837 --> 00:28:52,578
already for you.

430
00:28:52,578 --> 00:28:54,228
So build versus buy again.

431
00:28:54,551 --> 00:29:00,776
I I wanna ask you about that because I still think this is like an open question that I
have yet to see a canonical like right answer to.

432
00:29:00,776 --> 00:29:08,852
It's like the first one is do backups live in the same cloud account or project depending
on the cloud nomenclature as the original data source?

433
00:29:09,002 --> 00:29:11,795
Yeah, I mean y you want to go across two re two things.

434
00:29:11,795 --> 00:29:13,096
There's two boundaries here.

435
00:29:13,096 --> 00:29:16,288
The fault isolation boundary, which is again cross region.

436
00:29:16,288 --> 00:29:19,821
And for most of those main cloud hyperscalers, you have that regional construct.

437
00:29:19,821 --> 00:29:21,133
So you want to go cross region.

438
00:29:21,133 --> 00:29:23,475
And then there's a security isolation boundary.

439
00:29:23,475 --> 00:29:23,725
Right.

440
00:29:23,725 --> 00:29:27,642
And that's why you want to go cross account or in Azure, they'd call it cross
subscription.

441
00:29:27,642 --> 00:29:32,354
Uh in AWS, some of our customers go cross AWS organization.

442
00:29:32,354 --> 00:29:36,115
Because they're afraid someone's gonna get access to their what used to be called the
master account.

443
00:29:36,115 --> 00:29:39,096
I can't remember what it's called, but ma everybody calls it master payer, right?

444
00:29:39,096 --> 00:29:42,067
So if someone gets count the master payer, then it doesn't matter.

445
00:29:42,067 --> 00:29:43,667
I'm in a different account, I'm cooked, right?

446
00:29:43,667 --> 00:29:44,898
So they go across organization.

447
00:29:44,898 --> 00:29:48,409
So you have those two boundaries the fault isolation boundary and the security isolation
boundary.

448
00:29:48,409 --> 00:29:52,060
Security isolation boundaries, obviously against ransomware and bad actors, right?

449
00:29:52,060 --> 00:29:55,310
If you've been ransomware's they that account's compromised.

450
00:29:55,310 --> 00:29:56,471
Don't burn it to the ground.

451
00:29:56,471 --> 00:30:01,014
I mean, we have a fail back capability in RPO, and that's for like

452
00:30:01,014 --> 00:30:02,605
Again, AWS regional issues.

453
00:30:02,605 --> 00:30:04,797
But if you've been ransomware, you you ain't found back.

454
00:30:04,797 --> 00:30:05,798
Don't go back there.

455
00:30:05,798 --> 00:30:07,589
That's that's a bad place.

456
00:30:07,589 --> 00:30:11,112
Just leave it to burn and die and tell AWS to shut it down.

457
00:30:11,230 --> 00:30:12,671
We're we're gonna have to get into that.

458
00:30:12,671 --> 00:30:14,590
Um and so remind me if I forget.

459
00:30:14,590 --> 00:30:22,193
Uh but on the a cloud account thing, one of the challenges here, just even figuring out
how like there there's a question of do I make it the same account, do I make it a

460
00:30:22,193 --> 00:30:22,554
separate one?

461
00:30:22,554 --> 00:30:23,974
How do I even set up that pipeline?

462
00:30:23,974 --> 00:30:25,474
What are the the knobs to turn?

463
00:30:25,474 --> 00:30:28,195
Like this is not straightforward or out of the box in any way.

464
00:30:28,195 --> 00:30:33,037
And one of the challenges that I keep on seeing is do you encrypt your backups?

465
00:30:33,037 --> 00:30:37,078
Uh or are they open to the public, you know, in plain text?

466
00:30:37,078 --> 00:30:38,270
It's trick question, War.

467
00:30:38,270 --> 00:30:38,936
What is that?

468
00:30:38,936 --> 00:30:42,379
Well, here's the thing though, where do you store the keys to do the encryption?

469
00:30:42,379 --> 00:30:53,168
'Cause in AWS they suggested use KMS in some regard, but if KMS is driven from a
management account and you go cross organization, then you need to do this weird trick

470
00:30:53,168 --> 00:30:59,122
where you need to do something in order to have it be encrypted w in the target account
and not where it came from.

471
00:30:59,224 --> 00:31:06,979
Yeah, so you know, uh say talking to the audience here, uh Warren sent me this big thing
that says, Don't chill while you're on the while you're on the podcast.

472
00:31:06,979 --> 00:31:08,180
I'm like, All right, I won't chill.

473
00:31:08,180 --> 00:31:13,104
And yet he lobs these softballs at me like it turns out RPL solves this for you, by the
way.

474
00:31:13,104 --> 00:31:16,660
But I I will say that at the DevOps Days talks I tea teach people how to do it.

475
00:31:16,660 --> 00:31:17,086
All right.

476
00:31:17,086 --> 00:31:19,008
And it's not easy, it's doable.

477
00:31:19,008 --> 00:31:23,551
And basically you have to p create your backup with the same key that the data is backed
up with.

478
00:31:23,551 --> 00:31:26,333
Then you actually re-encrypt the backup locally.

479
00:31:26,333 --> 00:31:28,204
You can move it to another region at that point.

480
00:31:28,204 --> 00:31:31,435
with a key that you've created in the recovery account.

481
00:31:31,435 --> 00:31:35,876
And you could do all this with AWS with cross account, IAM permissions, et cetera.

482
00:31:35,876 --> 00:31:38,077
So create the key in the recovery account.

483
00:31:38,077 --> 00:31:42,218
Use that to re-encrypt your backup and then copy your backup to the recovery environment.

484
00:31:42,218 --> 00:31:45,509
And now you have a consistent backup and key in the recovery count.

485
00:31:45,509 --> 00:31:46,250
Can you do that?

486
00:31:46,250 --> 00:31:46,820
Absolutely.

487
00:31:46,820 --> 00:31:47,360
We did it.

488
00:31:47,360 --> 00:31:51,141
But you know, yeah, it's uh, you know, do you want to do it or you want to buy it?

489
00:31:52,041 --> 00:31:53,228
Tell me if I'm shilling too much.

490
00:31:53,228 --> 00:31:54,408
I mean I'm trying not to.

491
00:31:54,408 --> 00:31:55,930
No, no, it it's it's fine.

492
00:31:55,930 --> 00:31:57,693
Um, especially on that one, honestly.

493
00:31:57,693 --> 00:32:06,628
It like it's not the sort of thing which is straightforward to do and this is after
knowing that you wanna do it, seeing multiple companies thinking about how to even how to

494
00:32:06,628 --> 00:32:09,326
even implement this in a way that makes sense and not getting

495
00:32:09,326 --> 00:32:09,916
add another one.

496
00:32:09,916 --> 00:32:14,648
I'm I know I'm cutting you up, but I I just so excited about this, like which is how about
things that don't have backup capability?

497
00:32:14,648 --> 00:32:17,109
How about secrets manager secrets, SSM parameters?

498
00:32:17,109 --> 00:32:20,580
Um I'll add Kubernetes manifests, although they added backup for that.

499
00:32:20,580 --> 00:32:22,480
It's it's doesn't do the translations for you.

500
00:32:22,480 --> 00:32:24,061
But let's let's stick with secrets manager secrets.

501
00:32:24,061 --> 00:32:24,651
How do you do that?

502
00:32:24,651 --> 00:32:25,571
There's no way.

503
00:32:25,571 --> 00:32:33,462
So like um the way I show folks how to implement it and the way RPO implemented is it
installs lambdas in both accounts and the lambda can read the secret, but

504
00:32:33,462 --> 00:32:34,654
It can't pass the secret out.

505
00:32:34,654 --> 00:32:35,726
It doesn't have permission to do that.

506
00:32:35,726 --> 00:32:40,976
So and then it creates a key on the recovery side, encrypts it, and copies it into a
bucket on the recovery side.

507
00:32:40,976 --> 00:32:43,891
So it gets more just more paper cut after paper cut.

508
00:32:43,891 --> 00:32:44,681
Yeah.

509
00:32:44,902 --> 00:32:55,447
we we have one with our own company because we're managing private keys for our customers
where we're encrypting the private key with a pass key and that's being encrypted with KMS

510
00:32:55,447 --> 00:33:02,936
and it's like, well, if even if you back up you can't decrypt the pass key because you
don't have access to the KMS key because it's in the account that's been compromised.

511
00:33:02,936 --> 00:33:03,930
You gotta re-encrypt.

512
00:33:03,930 --> 00:33:04,340
Yeah.

513
00:33:04,340 --> 00:33:12,724
And I just it's just it's just this uh nightmare of of like, well, crud, like if we
actually switch AWS accounts here, you can't even come back up because all the data in

514
00:33:12,724 --> 00:33:13,794
your database is

515
00:33:13,794 --> 00:33:16,638
basically client side encrypted before being it sent over.

516
00:33:16,638 --> 00:33:20,042
So you actually have to make sure you have access to the decryption keys.

517
00:33:20,042 --> 00:33:26,440
And I don't think this is like another whole step on top of even managing your backups,
which is just a pit of failure.

518
00:33:26,606 --> 00:33:34,091
It there's just tons of these and then like AWS Backup recently for several uh resource
types launched the ability to do cross region, cross account backup.

519
00:33:34,091 --> 00:33:39,574
I mean they always had cross region, they always had cross account, but they were two
discrete operations, but they launched it as a single operation.

520
00:33:39,574 --> 00:33:48,080
So it's playing with it and it'll happily let you kick off the recovery, cross region,
cross account, and only after it's done recovering it say, Hey, this key doesn't work

521
00:33:48,080 --> 00:33:48,280
here.

522
00:33:48,280 --> 00:33:48,773
You failed.

523
00:33:48,773 --> 00:33:52,214
Like could you have told me that like up front, please, like please

524
00:33:52,214 --> 00:34:02,808
Yeah, I do think that there's a more experienced narrative that has to happen here with
the validation of restoring from backup uh from the beginning, like making sure things

525
00:34:02,808 --> 00:34:05,853
that are are set up in in the right accounts and in the right way.

526
00:34:05,853 --> 00:34:07,032
And it's just not pl

527
00:34:07,032 --> 00:34:09,763
Yeah, no, it's not.

528
00:34:09,763 --> 00:34:17,215
And then you I'm I'm I'm at conferences, I'm standing at a booth, you know, with the wheel
of misfortune, by the way, you know, which is contains all these the all these uh actual

529
00:34:17,215 --> 00:34:22,707
disasters that happened, and then one slot is a quiet day, you win a gift certificate, but
everybody learns about failure that way.

530
00:34:22,747 --> 00:34:26,708
And once in a while, it's like we have infrastructure as code, we have backups, we're
good.

531
00:34:26,708 --> 00:34:28,048
And sometimes you are.

532
00:34:28,048 --> 00:34:32,240
I mean, literally the uh Amazon trucker app, I think that's what they did.

533
00:34:32,240 --> 00:34:34,126
They were good, mostly serverless.

534
00:34:34,126 --> 00:34:37,947
Uh just a little bit of state stored in Dynamo DB, super easy to replicate.

535
00:34:37,947 --> 00:34:38,228
All right.

536
00:34:38,228 --> 00:34:38,906
They were good.

537
00:34:38,906 --> 00:34:39,608
All right.

538
00:34:39,608 --> 00:34:41,139
But most people are not good, right?

539
00:34:41,139 --> 00:34:42,497
They they're gonna hit all these paper cuts.

540
00:34:42,497 --> 00:34:43,670
They're gonna hit all these issues.

541
00:34:43,670 --> 00:34:45,971
Let's talk about enterprise architecture, Warren.

542
00:34:45,971 --> 00:34:49,683
Enterprise architecture is so amazingly complex and crazy.

543
00:34:49,683 --> 00:34:53,754
Like I worked with a customer that did a simple migration to the cloud.

544
00:34:53,754 --> 00:34:57,706
They migrated their servers to static servers, which is unreliable, by the way.

545
00:34:57,706 --> 00:35:00,127
What anybody know SLA on a single AC two server?

546
00:35:00,127 --> 00:35:02,488
It's like ninety eight point five or something.

547
00:35:02,488 --> 00:35:02,914
So

548
00:35:02,914 --> 00:35:03,714
Keep that in mind.

549
00:35:03,714 --> 00:35:05,115
That's a couple days a year, I think.

550
00:35:05,115 --> 00:35:14,079
Um, so on-prem servers to static servers and on-prem databases to RDS instances, and each
static server talked to an RDS database, and they had like 30 of these.

551
00:35:14,379 --> 00:35:20,221
they were using a consultancy, I won't name which one, that and it's nothing wrong with
this, but it just added complexity.

552
00:35:20,221 --> 00:35:24,043
They used a single VPC shared across accounts using a RAM share.

553
00:35:24,043 --> 00:35:26,644
So RAM share is a way in AWS to share resources across accounts.

554
00:35:26,644 --> 00:35:28,175
And they decided that was the way to go.

555
00:35:28,175 --> 00:35:29,826
I'm like, oh, okay.

556
00:35:29,826 --> 00:35:31,306
You just added a whole.

557
00:35:31,360 --> 00:35:39,376
I mean, I don't know, you're achieving some simplicity here, I guess, with your VPC
management, but in terms of being able to recover that, that makes it more complex.

558
00:35:39,376 --> 00:35:40,498
And it was kind of funny to see.

559
00:35:40,498 --> 00:35:41,954
It was an interesting architecture.

560
00:35:41,954 --> 00:35:52,014
I I like the aspect, especially when consultants come in and do recommend like lift and
shift is fine to get to the cloud, but even with the EC two RDS connection, like you are

561
00:35:52,014 --> 00:36:02,745
just multiplying your spend by so much doing that and spreading that over time and taking
letting each individual organization or team manage their own s technology and shift to

562
00:36:02,745 --> 00:36:05,699
the cloud in the way that they want and isolated

563
00:36:05,699 --> 00:36:13,652
like not only has uh benefits from a pricing consideration, but reliability is there too,
because they actually understand what it is that they're deploying and where things are

564
00:36:13,652 --> 00:36:16,736
different from the basically being on prem before that.

565
00:36:16,736 --> 00:36:20,348
Yeah, I think all folks in those situations should be think should be thinking about
modernization.

566
00:36:20,348 --> 00:36:25,071
I mean it's okay to just want to get to the cloud and and like get a project done and get
a win.

567
00:36:25,071 --> 00:36:26,642
But then you should be thinking about modernization.

568
00:36:26,642 --> 00:36:26,952
All right.

569
00:36:26,952 --> 00:36:32,755
And I actually have a slide about that I give it at sometimes at talks where I start with
the here's a server and here's an RDS RDS.

570
00:36:32,755 --> 00:36:38,999
Let's say they go they go crazy and put an auto scaling auto sc load balancer in front of
it and go to an auto scaling group.

571
00:36:38,999 --> 00:36:45,164
And let's say they get completely wild and decide to use containerized workloads and and
and serverless data, you know, so

572
00:36:45,164 --> 00:36:47,436
Yeah, you could take it step by step and do that evolution.

573
00:36:47,436 --> 00:36:49,757
I don't know if this in this case whether this concern will.

574
00:36:49,757 --> 00:36:53,430
The other complexity at enterprises I see is centralized security.

575
00:36:53,430 --> 00:36:54,430
I know why they do it.

576
00:36:54,430 --> 00:37:04,457
It's actually it's a simplifier for managing security where all uh all traffic goes
through a single VPC and a single account and it's probably got some third party Palo

577
00:37:04,457 --> 00:37:12,553
Alto's or some other devices attached to it and the transit gateway is defined there and
it's also where the direct connect back to on prem lives and everything goes through

578
00:37:12,553 --> 00:37:12,913
there.

579
00:37:12,913 --> 00:37:14,720
But from a recovery point of view.

580
00:37:14,720 --> 00:37:16,661
It's not trivial, right?

581
00:37:16,661 --> 00:37:19,474
I mean, you could either try you know, take two approaches.

582
00:37:19,474 --> 00:37:28,060
Either I'm gonna recover everything, I'm gonna recover that centralized account and all
the client accounts, or I'm gonna pre set up that r that centralized uh security and then

583
00:37:28,060 --> 00:37:31,482
when I fail over it has to it has to plumb itself into it.

584
00:37:31,672 --> 00:37:32,993
But I I get it as complicated.

585
00:37:32,993 --> 00:37:33,403
Yeah.

586
00:37:33,403 --> 00:37:34,123
Yeah, for sure.

587
00:37:34,123 --> 00:37:43,951
And but I do get it because this is like the organizations evolved from having a whole
team of database administrators, and then there's a level of software developers who are

588
00:37:43,951 --> 00:37:49,065
not allowed to touch any code that goes to production under any circumstance or interact
with the database directly.

589
00:37:49,065 --> 00:37:50,576
And then there's like release engineering.

590
00:37:50,576 --> 00:37:54,429
And when they lift and shift to the cloud, they didn't just take their technology and
stuck it there.

591
00:37:54,429 --> 00:37:58,376
They took their organizational dynamics and put them there a as well.

592
00:37:58,376 --> 00:38:05,892
And yeah I I think this is sort of the the the trouble is because the cloud isn't
necessarily set up not only with the technology to work, but with the same organizational

593
00:38:05,892 --> 00:38:10,675
uh communication patterns, if we go with the Conway's Law perspective here.

594
00:38:10,675 --> 00:38:14,318
And I I think it just ends up causing a lot of problems in in every regard.

595
00:38:14,318 --> 00:38:18,030
I just want to super clear though, both those enterprise architectures I did I I described
are good.

596
00:38:18,030 --> 00:38:20,451
There's nothing actually fundamentally wrong with them.

597
00:38:20,451 --> 00:38:21,491
They serve a purpose.

598
00:38:21,491 --> 00:38:24,843
What I'm saying is that sometimes to recover them is complex.

599
00:38:24,843 --> 00:38:31,468
uh

600
00:38:31,468 --> 00:38:41,744
Well, does it make it easier on you and I think you with an asterisk here because it makes
it easier on I'd say the security team, uh whoever is incentivized with having that domain

601
00:38:41,744 --> 00:38:51,060
set up that way rather than what makes sense for the business holistically, because I
think this is one of those areas where it's much easier to manage and ensure a single RDS

602
00:38:51,060 --> 00:38:58,328
account or a cluster in a single account has the right backup configured and is secure and
has only the right security groups attached.

603
00:38:58,328 --> 00:39:07,198
So that your one monolithic Kubernetes cluster with with the hundreds of thousands of
containers running in it uh is the only thing connecting with it, rather than having to

604
00:39:07,198 --> 00:39:17,611
somehow figure out a real organizational strategy to convey to individual teams how to
actually secure the communication between their clo their containers and the production

605
00:39:17,611 --> 00:39:18,690
infrastructure database.

606
00:39:18,690 --> 00:39:19,290
It's interesting.

607
00:39:19,290 --> 00:39:20,182
Yeah, there's two models.

608
00:39:20,182 --> 00:39:21,743
There's that model and it's perfectly legitimate.

609
00:39:21,743 --> 00:39:22,843
And then there's the Amazon model.

610
00:39:22,843 --> 00:39:30,122
Amazon model, everybody owns their own stuff and there's policies and there's procedures
and there is some automated scanning and enforcement.

611
00:39:30,122 --> 00:39:32,914
But you know, you do you have to be educated as an engineer.

612
00:39:32,914 --> 00:39:34,676
There's no architects, strangely enough.

613
00:39:34,676 --> 00:39:38,638
Even even though there's a whole solutions architects group at AWS, Amazon has no
architects.

614
00:39:38,638 --> 00:39:42,980
Um right, except for SAs, that few that exist like myself at the time.

615
00:39:42,980 --> 00:39:44,211
And and it's up to the teams.

616
00:39:44,211 --> 00:39:50,944
That's up to the senior developer or principal developer to underst and every developer
actually to understand this stuff, but principal developers, senior developer make sure

617
00:39:50,944 --> 00:39:54,926
their entire teams understand this stuff and are adh adhering to policies and doing things
the right way.

618
00:39:54,926 --> 00:39:56,936
I mean, Amazon makes it tries to make it easier.

619
00:39:56,936 --> 00:40:03,409
I said there's some automation in terms of setting up your infrastructure, there's
automation in terms of scanning the infrastructure, but ultimately it's parceled out team

620
00:40:03,409 --> 00:40:04,030
ownership.

621
00:40:04,030 --> 00:40:04,620
I mean

622
00:40:04,620 --> 00:40:10,573
They they don't call it DevOps at Amazon, but the the leadership principles at Amazon,
when they're being followed right, are pure DevOps.

623
00:40:10,573 --> 00:40:14,655
It's like there there is no I should I was about to say there's no operations team.

624
00:40:14,655 --> 00:40:16,206
They've kind of muddied that a little.

625
00:40:16,206 --> 00:40:17,656
Some teams now have operations teams.

626
00:40:17,656 --> 00:40:20,107
But most teams traditionally have not had operations team.

627
00:40:20,107 --> 00:40:21,506
You you build it, you run it.

628
00:40:21,506 --> 00:40:23,928
Yeah, no, I I I totally I totally get that.

629
00:40:23,928 --> 00:40:32,153
So there is something here though, which may be worth jumping into, which is do you
advocate for one of these approaches over the other one?

630
00:40:32,153 --> 00:40:39,020
I I I do see what your perspective is of there's merits to both of them, which is like
there is a simplicity in treat everything the same.

631
00:40:39,020 --> 00:40:47,183
But then there's also a perspective of not everything does need to be treated the same
because maybe there's only a particular part of your business data or your your product,

632
00:40:47,183 --> 00:40:54,197
your service or whatever database that does need to be replicated or secured cross region
or across uh AWS organization.

633
00:40:54,197 --> 00:41:04,842
And you're forcing that requirement across your entire uh infrastructure where that adds a
lot of complexity on say the implementation side.

634
00:41:05,058 --> 00:41:05,408
Yeah.

635
00:41:05,408 --> 00:41:07,329
I mean, there there is some centralized infrastructure.

636
00:41:07,329 --> 00:41:11,510
Remember I was talking about how do you connect your your native AWS accounts to the big
blob?

637
00:41:11,510 --> 00:41:12,940
Well, that was centralized infrastructure.

638
00:41:12,940 --> 00:41:15,481
We don't want everybody building their own private links.

639
00:41:15,481 --> 00:41:16,571
That would be insane, right?

640
00:41:16,571 --> 00:41:18,272
So but I don't know.

641
00:41:18,272 --> 00:41:19,682
I kind of tend towards the Amazon.

642
00:41:19,682 --> 00:41:22,423
You build it, you run it, you own it model, but I don't bring data.

643
00:41:22,423 --> 00:41:23,403
I mean, I just like it.

644
00:41:23,403 --> 00:41:24,444
I I've worked with it.

645
00:41:24,444 --> 00:41:25,834
I've been a dev at Amazon.

646
00:41:25,834 --> 00:41:28,015
I was actually on the uh Prime Video team.

647
00:41:28,015 --> 00:41:29,638
It wasn't called Prime Video at the time.

648
00:41:29,638 --> 00:41:32,786
you know, shout out to anybody who remembers Amazon Unbox.

649
00:41:33,138 --> 00:41:34,979
but you know, I helped build that product.

650
00:41:34,979 --> 00:41:36,359
So I mean, I like that model.

651
00:41:36,359 --> 00:41:38,670
I I think, you know, it was the right amount of ownership.

652
00:41:38,670 --> 00:41:41,061
I never felt I shouldn't say I never felt unburdened by it.

653
00:41:41,061 --> 00:41:50,014
There are times where like, okay, I want to go use this janky little internal key value
system that Amazon uh created and there's nobody that wants to help me because everybody's

654
00:41:50,014 --> 00:41:50,625
trying to use it.

655
00:41:50,625 --> 00:41:51,985
And that was kind of annoying.

656
00:41:51,985 --> 00:41:54,026
But yeah, other than that, it was pretty good.

657
00:41:54,882 --> 00:42:05,963
Uh wait, was Prime Video the one that came out of the blog post a a while ago that was
that said like, we made a mistake by going to microservices and we re-model it.

658
00:42:05,963 --> 00:42:06,854
way to put it though.

659
00:42:06,854 --> 00:42:12,080
Like really if you read what it says, we went from many microservices to a fewer
microservices.

660
00:42:12,080 --> 00:42:14,963
I see I like there was no monolith there.

661
00:42:14,963 --> 00:42:15,604
Yeah.

662
00:42:15,604 --> 00:42:20,865
Well I I actually think uh I I I I think they made a mistake by even saying the word
microservice in there.

663
00:42:20,865 --> 00:42:26,510
It's like we made a mistake with our architecture, we picked the wrong thing and now we're
making the changes to to rectify them.

664
00:42:26,510 --> 00:42:27,172
Wrong boundary.

665
00:42:27,172 --> 00:42:27,925
Yeah, exactly.

666
00:42:27,925 --> 00:42:31,206
Like it was so inflammatory, and they went with that title.

667
00:42:31,206 --> 00:42:33,710
I guess it gets you clicks, but it was crazy.

668
00:42:33,710 --> 00:42:40,966
lots of people then were like, Well, look, even Am AWS is saying that or Amazon is saying
that um so you know serverless is isn't great.

669
00:42:40,966 --> 00:42:49,512
You know, they made a mistake with Lambit, and I'm like, please just actually reread the
article because you'll you'll see what like if you look at it, it's like, yeah, there was

670
00:42:49,512 --> 00:42:51,135
a specif like they did the right thing.

671
00:42:51,135 --> 00:42:57,610
They said, here's our hypothesis on where we're spending money and where performance is,
and here's where our current problems are.

672
00:42:57,610 --> 00:43:01,483
And if we switch to this other architecture uh paradigm, we will get these benefits.

673
00:43:01,483 --> 00:43:02,626
And then they did the switch and say,

674
00:43:02,626 --> 00:43:03,687
Look, we were right.

675
00:43:03,687 --> 00:43:10,472
Like these things were a specific trade off and not like, oh, I'm principled and I think
that monolith is better, so we're gonna do that.

676
00:43:10,472 --> 00:43:13,704
And now we switched over and of course I'm right, so everything works out.

677
00:43:14,062 --> 00:43:17,424
Like so many articles today, the content's good, but the headlines misleading.

678
00:43:17,424 --> 00:43:21,577
You know, that but it also points out another thing about Amazon and AWS teams is they
watch their budget.

679
00:43:21,577 --> 00:43:25,100
Like each team is responsible for their spend, which I think is really cool.

680
00:43:25,100 --> 00:43:33,085
When I was, again, an SA working for Amazon, going traveling literally around the world to
Tokyo and Luxembourg, meeting with teams and saying, Hey, get into these AWS accounts,

681
00:43:33,085 --> 00:43:36,257
start using these, start using Lambda and API Gateway.

682
00:43:36,257 --> 00:43:40,680
And it was always you know, the technology discussion was important, but how much is that
gonna cost me?

683
00:43:40,680 --> 00:43:43,582
You know, I I if you have a high

684
00:43:43,698 --> 00:43:50,886
TPS set of EC two instances that are always taking traffic, API and gateway and with
Lambda is gonna cost you a bundle.

685
00:43:50,886 --> 00:43:54,589
So like, all right, maybe you're saving some operational load, but that might not be the
right thing.

686
00:43:54,589 --> 00:43:58,834
But if you got something that's asynchronous and reading events, then yeah, Lambda's gonna
be for the win.

687
00:43:58,834 --> 00:44:00,465
So they always look at the cost.

688
00:44:00,465 --> 00:44:03,304
And I I always like, you know, that part of the conversation.

689
00:44:03,304 --> 00:44:11,974
it's also so much cheaper if you just switch to Lambda Edge and put and use Edge Compute
and have instead of API Gateway, put your whole service within a Lambda Edge function

690
00:44:11,974 --> 00:44:14,217
instead with Cloudfront on on top.

691
00:44:14,217 --> 00:44:20,174
And so instead of paying for API Gateway, you're paying for Cloudfront, which has like a a
better optimized uh cost strategy.

692
00:44:20,174 --> 00:44:20,482
I

693
00:44:20,482 --> 00:44:22,911
I think you're ra are you trying to ra rage bait someone out there?

694
00:44:22,911 --> 00:44:24,282
I don't know.

695
00:44:24,282 --> 00:44:25,913
Entire app at Glandad Edge.

696
00:44:25,913 --> 00:44:26,405
Interesting.

697
00:44:26,405 --> 00:44:26,715
So

698
00:44:26,715 --> 00:44:34,192
we we have a couple of different imp architectures for our own product because we're doing
login and access control uh as a service.

699
00:44:34,192 --> 00:44:40,856
There are aspects which we want actually closer to the users as possible and ones where we
want a little bit more centralized and

700
00:44:40,856 --> 00:44:41,981
That's what it's made for, yeah.

701
00:44:41,981 --> 00:44:44,648
Like but you you implied you were putting everything out there.

702
00:44:44,648 --> 00:44:53,715
well, I I will say that it does make the resiliency uh equation much simpler when you're
already prepared to deal with edge compute and you have to figure out how to provide the

703
00:44:53,715 --> 00:45:00,698
data store for that and you don't have like a single point of failure uh from a
geolocation aspect.

704
00:45:00,698 --> 00:45:03,512
Uh but yeah, for sure there are are complexities there.

705
00:45:03,512 --> 00:45:07,736
And th these things aren't long running, uh so and usually they're they're cacheable.

706
00:45:07,736 --> 00:45:10,880
Like is the user logged in is much faster than having to check

707
00:45:10,880 --> 00:45:15,771
actually go through the login process every single request that comes in from uh an end
user.

708
00:45:15,771 --> 00:45:16,870
So I think

709
00:45:24,558 --> 00:45:26,619
Oh no, they're they're actually not.

710
00:45:26,619 --> 00:45:29,139
So the Lambda Edge are actually not running in a pop.

711
00:45:29,139 --> 00:45:32,050
They're actually running in a specific region.

712
00:45:32,050 --> 00:45:41,763
So in AWS's infinite wisdom, what you do is you deploy in order to actually run Lambda
Edge, you deploy a regular Lambda to the US East one region.

713
00:45:41,763 --> 00:45:45,564
And then you give Yeah and then you give a magic sir yeah.

714
00:45:45,564 --> 00:45:53,592
And then you give a magic service in AWS access to that lambda function, which then
replicates that lambda function to every single region

715
00:45:53,902 --> 00:45:55,862
Ферст клас режен араундово.

716
00:45:55,968 --> 00:45:57,141
Interesting.

717
00:45:58,328 --> 00:46:05,300
and so the logs are all um in different regions, which makes it impossible to actually
find where issues happen.

718
00:46:05,300 --> 00:46:15,033
And so we have a whole complex log aggregation strategy where we we pull stuff from all of
the cloud fronts uh logging regions uh and merge them together so we can actually

719
00:46:15,033 --> 00:46:22,645
understand when a customer makes a complaint, this user failed to do something, we can
actually find what happened there rather than clicking in the console for all of like

720
00:46:22,645 --> 00:46:25,272
twenty four or however many regions there are now.

721
00:46:25,272 --> 00:46:28,304
to figure out which actual region to run those queries in.

722
00:46:28,304 --> 00:46:28,734
okay.

723
00:46:28,734 --> 00:46:30,155
That that does give you a lot of resiliency.

724
00:46:30,155 --> 00:46:33,728
By the way, when you're talking about Conway's Law and Organizational friction, there was
one story I wanted to mention.

725
00:46:33,728 --> 00:46:36,360
I was working with a customer on their disaster recovery.

726
00:46:36,360 --> 00:46:45,556
This is a big company that's providing consumer service that everybody recognize and most
people think of as a tech company, but I learned that internally they're not organized

727
00:46:45,556 --> 00:46:46,306
like a tech company.

728
00:46:46,306 --> 00:46:49,008
Internally they're organized like a like a bank or something, right?

729
00:46:49,008 --> 00:46:57,428
So I'm working with this one team on recovering their EKS workload, ECS workload, and it's
got all kinds of connections and all kinds of databases, et cetera.

730
00:46:57,428 --> 00:47:01,460
And it had a dependency on an S3 bucket in a different account, right?

731
00:47:01,460 --> 00:47:04,592
And so, you know, I mean we're using RPO RPO will recognize that.

732
00:47:04,592 --> 00:47:13,017
Hey, there's an S3 bucket in another account, you need to actually onboard that too into
RPO so we could replicate that for you too to your recovery region and we'll reconnect you

733
00:47:13,017 --> 00:47:14,908
back up with a new recovered S3 bucket.

734
00:47:14,908 --> 00:47:17,160
And their response was, well, that's that's another team.

735
00:47:17,160 --> 00:47:17,820
They're out of scope.

736
00:47:17,820 --> 00:47:20,201
I'm like, but you kind of need that for this functionality.

737
00:47:20,201 --> 00:47:22,042
Well, we don't care about we don't care whether they recover.

738
00:47:22,042 --> 00:47:23,113
We just care whether we recover.

739
00:47:23,113 --> 00:47:24,888
But if they don't recover, you're not.

740
00:47:24,888 --> 00:47:25,649
quite recovering.

741
00:47:25,649 --> 00:47:26,538
They're like, we don't care.

742
00:47:26,538 --> 00:47:30,103
I'm like, wow, that is organizational poison.

743
00:47:30,464 --> 00:47:32,322
That's I couldn't believe I got the response.

744
00:47:32,322 --> 00:47:33,403
I'm I'm really surprised too.

745
00:47:33,403 --> 00:47:41,872
That that seems like quite some discourse because it it's like why would you set the
company up to go down the process of implementing any sort of resiliency without then

746
00:47:41,872 --> 00:47:43,726
being prepared to actually do it?

747
00:47:43,726 --> 00:47:44,957
Well it was another team, right?

748
00:47:44,957 --> 00:47:51,042
And and they owned the resiliency mandate for their stuff, but that that team was a
different team and they they could do their resiliency another way.

749
00:47:51,042 --> 00:47:57,638
Uh it was I mean, you know, it was a proof of concept, so maybe to give them some benefit
of the doubt, they said they were just for the proof of concept we could leave them out.

750
00:47:57,638 --> 00:48:00,881
But like I don't I it was just the the action was very straight.

751
00:48:00,881 --> 00:48:01,942
It wasn't said like that.

752
00:48:01,942 --> 00:48:08,007
Like, okay, for the proof concept, it might be too difficult to work with that they
literally said that their resiliency is their problem, not our problem.

753
00:48:08,007 --> 00:48:09,588
Like, but it is kind of easier problem.

754
00:48:09,588 --> 00:48:11,970
Like you need that S3 bucket to do some things.

755
00:48:12,026 --> 00:48:18,392
I like I like the perspective because you don't try to fill in the gaps for failures in
other teams.

756
00:48:18,392 --> 00:48:24,697
You want to elevate every team to be doing the right things and not sort of being a crutch
for them.

757
00:48:24,727 --> 00:48:30,822
on the on the opposite side, I do also understand the aspect of uh that comes out of a
different budget.

758
00:48:30,914 --> 00:48:32,831
comes at a different budget, different leader.

759
00:48:32,831 --> 00:48:35,088
Like it was c funny.

760
00:48:35,514 --> 00:48:37,736
Yeah, it's not it's not one of our KPIs actually.

761
00:48:37,736 --> 00:48:45,213
Um so I I did see your a report, I think it was from a couple of years ago now, that and
you brought up malicious actors.

762
00:48:45,213 --> 00:48:54,011
And I think it's something that we can't completely ignore when we're talking about
resiliency, because while it does seem like security isn't not necessarily part of the

763
00:48:54,011 --> 00:49:01,856
reliability angle, I do find that if you have threat actors that are coming and taking
down your software, it's not actually up and usable for your paying.

764
00:49:01,856 --> 00:49:02,196
You're right.

765
00:49:02,196 --> 00:49:03,207
No, it absolutely is.

766
00:49:03,207 --> 00:49:05,092
It absolutely is for the reasons you said.

767
00:49:05,092 --> 00:49:06,534
It is point of reliability.

768
00:49:06,534 --> 00:49:11,778
Uh yeah, but companies are no longer paying ransomware actors that threaten to release
their data publicly.

769
00:49:11,778 --> 00:49:12,742
No longer paying them.

770
00:49:12,742 --> 00:49:13,190
Yeah.

771
00:49:13,190 --> 00:49:14,458
Uh they're just what?

772
00:49:14,458 --> 00:49:15,602
What are they doing?

773
00:49:15,602 --> 00:49:19,245
Not not it's like well, I I think and y I I do think it sort of makes sense.

774
00:49:19,245 --> 00:49:27,533
Basically if you pay for ransomware to were they just about leaking the data publicly and
they can still do it anyway.

775
00:49:27,533 --> 00:49:28,162
So

776
00:49:28,162 --> 00:49:29,482
Yeah, there's no guarantees.

777
00:49:29,482 --> 00:49:32,955
But but there's also yeah, there's also your your availability, like you just said.

778
00:49:32,955 --> 00:49:35,207
Like your data is encrypted, you can't use it.

779
00:49:35,207 --> 00:49:42,631
Now obviously I'm gonna advocate for having point in time recovery of your data and
infrastructure so that you don't have to pay them.

780
00:49:42,631 --> 00:49:50,456
And uh also uh scanning of any of your static compute so you could find the malware that
they injected to to encrypt your data.

781
00:49:50,456 --> 00:49:52,728
And these are all things I'm happy to help folks with.

782
00:49:52,728 --> 00:49:53,238
But

783
00:49:53,238 --> 00:49:56,712
Yeah, to not have that and not pay is kind of an interesting like then you're down, right?

784
00:49:56,712 --> 00:49:57,853
So like change healthcare.

785
00:49:57,853 --> 00:49:58,761
Let's talk about change healthcare.

786
00:49:58,761 --> 00:49:59,134
What that?

787
00:49:59,134 --> 00:50:00,155
Twenty twenty four, right?

788
00:50:00,155 --> 00:50:03,249
Major major breach, major d major encryption.

789
00:50:03,249 --> 00:50:08,844
And we got called into one of their competitors to like and we actually they ended up
becoming a customer of ours.

790
00:50:08,844 --> 00:50:10,765
And what the competitor told us was really fascinating.

791
00:50:10,765 --> 00:50:18,198
So you've heard about change healthcare that it was a major data breach and they were down
for X many days and it cost them, you know, hundreds of millions of dollars or whatever.

792
00:50:18,198 --> 00:50:19,779
I think you just say a billion dollars.

793
00:50:19,779 --> 00:50:20,286
You heard that.

794
00:50:20,286 --> 00:50:25,731
The one thing you haven't heard, and that their competitor told us is that change
healthcare eventually got their availability back, right?

795
00:50:25,731 --> 00:50:26,431
They're up and running.

796
00:50:26,431 --> 00:50:33,554
And and just so folks know, change healthcare was sort of sitting between doctors' offices
and hospitals and insurers and other payment systems, right?

797
00:50:33,554 --> 00:50:35,605
They're the middle pit middleman, right?

798
00:50:35,605 --> 00:50:38,496
And once they came up and they were back up and running,

799
00:50:38,496 --> 00:50:39,827
Nobody connected to them.

800
00:50:39,827 --> 00:50:41,268
They were tainted, right?

801
00:50:41,268 --> 00:50:46,332
Like there was no guarantee that they'd found all the malware or that they'd rooted out
all the ransomware.

802
00:50:46,332 --> 00:50:54,699
And nobody wanted to connect to them because they had no way of assuring the public, or
not the public, but their customers that they were safe.

803
00:50:54,699 --> 00:51:02,606
And so when we called into this competitor, one of the things this competitor does is they
use our product, they use RPO to create recovery points.

804
00:51:02,606 --> 00:51:08,172
And then they actually have a third party uh mandi in this case, which I think is part of
Google, certify it.

805
00:51:08,172 --> 00:51:09,283
And kind of stamp it out.

806
00:51:09,283 --> 00:51:10,244
This is the gold one.

807
00:51:10,244 --> 00:51:16,619
So now they actually put on their website, you know, that we'll be back up in five days or
less and we're certified clean when we come back.

808
00:51:16,619 --> 00:51:22,479
Yeah, Manian's the one that discovers all the zero days and all of the products out
outside of Google as well.

809
00:51:22,479 --> 00:51:24,133
They're the ones that are always doing the reporting.

810
00:51:24,133 --> 00:51:24,993
Yeah.

811
00:51:25,144 --> 00:51:27,387
So I thought that was just, you know, like interesting story.

812
00:51:27,387 --> 00:51:32,723
Like it goes beyond the beyond the data breach, beyond the downtime, but to trust, right?

813
00:51:32,723 --> 00:51:33,744
They're back up and running.

814
00:51:33,744 --> 00:51:38,690
You know, they they're saying, Hey, connect to us and people were like, Whoa, wait a
second, you know, I don't know, not so sure about that.

815
00:51:38,690 --> 00:51:44,984
That's really interesting because w the the data says that you can get breached as many
times as you want.

816
00:51:44,984 --> 00:51:50,418
And every time you get breached, the amount of money that you make actually goes up.

817
00:51:50,418 --> 00:51:53,841
You have more customers because of the near mere exposure effect.

818
00:51:53,841 --> 00:51:56,902
You just hear the name and it it gets improved.

819
00:51:58,170 --> 00:51:59,074
Yeah, I guess.

820
00:51:59,074 --> 00:52:00,045
Right, right, exactly.

821
00:52:00,045 --> 00:52:03,047
However, being down is a whole other story.

822
00:52:03,047 --> 00:52:04,308
Like, forget about security.

823
00:52:04,308 --> 00:52:07,542
Like if you don't have if you can't actually service customers.

824
00:52:07,542 --> 00:52:13,546
It's always going to put that idea in their head that there's going to be a that there can
be a problem again.

825
00:52:13,546 --> 00:52:17,899
And while there's like at some point there is going to be an issue, you're going to have
to deal with being breached.

826
00:52:17,899 --> 00:52:22,072
You're going to have to deal with being dead with downtime and explaining to customers
what's going on.

827
00:52:22,072 --> 00:52:29,917
I feel like uh one we did a survey on our own customers and they're like, Yeah, if you're
down for more than like a an hour, really 24 hours, like that's it.

828
00:52:29,917 --> 00:52:33,920
Like you we are already looking for a new provider and we're telling everyone about it.

829
00:52:33,920 --> 00:52:37,176
And I'm like, that's totally fair uh in in in some regard.

830
00:52:37,176 --> 00:52:41,061
But you know, you are you're making payments for the for the providers.

831
00:52:41,061 --> 00:52:49,622
Like you're responsible for actually paying out to the providers, which directly impacts
uh the being able to keep the lights on at hospital organizations and clinics, which means

832
00:52:49,622 --> 00:52:51,033
actually providing care.

833
00:52:51,033 --> 00:52:56,650
Like that's uh hospital organizations are very sensitive to uh stuff like that.

834
00:52:56,800 --> 00:53:02,764
A lot of customers left change and went to this this competitor I'm talking about because
of that.

835
00:53:02,764 --> 00:53:03,935
No, it makes a lot of sense.

836
00:53:03,935 --> 00:53:10,636
I do I do think that we probably can't end this conversation without bringing up uh
anthropics uh mythos.

837
00:53:10,636 --> 00:53:21,694
Uh which I did chat with our guests um on the previous episode about tech debt and I did
mention that it's more focused on security, but the question I I really want to put out

838
00:53:21,694 --> 00:53:24,196
there is there's some malicious threat actors out there.

839
00:53:24,196 --> 00:53:29,089
I think for sure that's happening, state sponsored or not, it's not really a question.

840
00:53:29,089 --> 00:53:31,342
And they cause impacts to our services.

841
00:53:31,342 --> 00:53:31,918
Um

842
00:53:31,918 --> 00:53:32,759
Collectively.

843
00:53:32,759 --> 00:53:41,066
Now, regard regardless of whether or not there's a malevolent superpower or LM that's
actually doing this, that's going to end all security and all software anywhere.

844
00:53:41,066 --> 00:53:43,718
My question is we still have to protect our stuff.

845
00:53:43,718 --> 00:53:55,018
And I guess does the conversation for you, has it changed at all recently as far as what
actually it means to implement a reliable or resilient solution?

846
00:53:55,470 --> 00:53:57,121
That hasn't really entered the conversation yet.

847
00:53:57,121 --> 00:53:58,831
Not what the customers were t we're talking to.

848
00:53:58,831 --> 00:54:03,706
I I guess eventually I think it's gonna take some high profile events and then it'll enter
the converse conversation.

849
00:54:03,706 --> 00:54:13,284
Honestly, even when it was AWS and reliability lead, uh, you know, it took an outage in
2020 and another in 2021 to get everybody internally all all excited about reliability.

850
00:54:13,284 --> 00:54:16,866
Before that, you know, people were like, yeah, I guess that's important.

851
00:54:17,007 --> 00:54:20,950
So yeah, it's interesting that that that has but what I've seen with AI.

852
00:54:21,334 --> 00:54:23,856
In general, people are still in the enamored phase.

853
00:54:23,856 --> 00:54:30,320
Like you go to any conference and you got multiple tracks, the conference with AI and the
title is just gonna get a big crowd.

854
00:54:30,320 --> 00:54:33,452
And I people just have this FOMO, fear of missing out or something.

855
00:54:33,452 --> 00:54:36,734
They feel like if they don't go learn everything they can, their jobs are in peril.

856
00:54:36,734 --> 00:54:41,207
Well, yeah, combine that with a really crappy tech job economy right now.

857
00:54:41,207 --> 00:54:45,751
Um, people are just totally enamored and that's I guess a little scary, right?

858
00:54:45,751 --> 00:54:51,244
That people are not seeing it as as both a benefit and a and a and a risk.

859
00:54:51,403 --> 00:54:54,540
I always want people to see the risks and I don't think people are seeing yet.

860
00:54:54,540 --> 00:55:03,315
So so I think a huge part of this is that what's happened to teenagers with TikTok is now
happening to software engineers and knowledge workers with AI.

861
00:55:03,315 --> 00:55:04,436
oh Companies.

862
00:55:04,436 --> 00:55:10,840
You see basically the whole product suite is or orchestrated in a way to get you addicted.

863
00:55:10,920 --> 00:55:13,162
And they're using all these dark patterns.

864
00:55:13,162 --> 00:55:20,418
And if you're not using it, you see people who are jumping up and and shouting zealot
evangelial nonsense.

865
00:55:20,418 --> 00:55:24,760
To try to convince people to come on board because it there's a there's something in it
for them.

866
00:55:24,760 --> 00:55:33,554
And you know, it's interesting you say that it's not really making any headwinds, um,
splashes yet, is because I'm really plugged into the security domain and there are already

867
00:55:33,554 --> 00:55:38,176
people talking about like, are finding vulnerabilities by humans, is that over?

868
00:55:38,176 --> 00:55:40,747
Are all our jobs fundamentally changing?

869
00:55:40,747 --> 00:55:43,138
And the there's like two answers.

870
00:55:43,138 --> 00:55:48,810
The first one is no, mythos isn't going to change anything, really, but yes, all our jobs
are changing.

871
00:55:48,810 --> 00:55:51,251
And I I think that's sort of a mature response to it.

872
00:55:51,251 --> 00:56:02,793
And I think we're not far out from the aspect of if every piece of vulnerable software out
there will be exploited and not in five years, ten years, thirty years, whatever, but in

873
00:56:02,793 --> 00:56:07,417
the next couple of years, irrelevant of whatever technology we think is available.

874
00:56:07,417 --> 00:56:14,168
And I think that has a real impact for the conversations we're going to have about how to
even build secure but reliable software.

875
00:56:14,168 --> 00:56:21,343
Yeah, and then the other thing is just seeing um all the AIs uh enter into the pr pr
productivity streams.

876
00:56:21,343 --> 00:56:28,959
I've certainly I've played with Cloud Coda quite a bit and and I know all the developers
here use it and it's doing a lot of great stuff, but it still needs supervision, still

877
00:56:28,959 --> 00:56:31,350
needs expertise to see what it did.

878
00:56:31,350 --> 00:56:35,353
Outside of the tech realm, I know a lot of people are using it to generate slideware.

879
00:56:35,353 --> 00:56:38,986
I I kind of joke that my my number one skill set is is that I can make good slides.

880
00:56:38,986 --> 00:56:42,230
Uh aside from anything else, I'm good at making slides.

881
00:56:42,230 --> 00:56:47,373
And I see the slides it can make and it's immediately spot you know, spottable that, this
was made by Claude.

882
00:56:47,373 --> 00:56:48,854
It's it's garbage.

883
00:56:48,854 --> 00:56:54,816
Like I mean it it's it's wordy, it's it's obtuse, it's you know, and people are just going
with it now.

884
00:56:54,816 --> 00:57:02,314
So we actually have a whole episode where we dedicated to working with some of the LLMs
and vibe coding and so I I'll link those in the description.

885
00:57:02,314 --> 00:57:12,864
But I will say the the thing that you reminded me of is that the consensus is that, LMs
that produce output in my area of specialties, that's garbage.

886
00:57:12,864 --> 00:57:17,669
But in everyone else's area of special spe specialization, I get it.

887
00:57:17,669 --> 00:57:18,732
You know, I would totally

888
00:57:18,732 --> 00:57:20,062
Really well, yeah.

889
00:57:20,803 --> 00:57:21,353
That's funny.

890
00:57:21,353 --> 00:57:24,744
So that there that just confirms that my skill set is making slides and not coding.

891
00:57:24,744 --> 00:57:30,816
So I'm actually the perfect customer for Claude Code because I'm not a developer, even
though I was a developer, but I wasn't a good developer.

892
00:57:30,816 --> 00:57:34,227
I I think I'm better at high picture stuff, architecture, how things fit together.

893
00:57:34,227 --> 00:57:36,287
So I understand what a bash script can do.

894
00:57:36,287 --> 00:57:37,928
I understand what a Python script can do.

895
00:57:37,928 --> 00:57:42,069
But I when it comes down to actually creating it and and writing the syntax, not so much.

896
00:57:42,069 --> 00:57:44,654
So I can actually give pretty tight guidance.

897
00:57:44,654 --> 00:57:49,117
to like a clawed code and tell it what I want what technologies I want it to use and how
to use it.

898
00:57:49,117 --> 00:57:51,819
And it comes back with it it unleashes my creativity.

899
00:57:51,819 --> 00:57:52,930
Like I wanna automate something.

900
00:57:52,930 --> 00:57:55,592
Now I can automate it and I actually get something good out of it.

901
00:57:55,822 --> 00:58:00,039
I think maybe at this point we're at a a good spot to maybe switch over to to picks.

902
00:58:00,039 --> 00:58:03,274
So Seth, what did you bring for the audience today?

903
00:58:03,608 --> 00:58:05,389
So yeah, you mentioned you want to hear my picks.

904
00:58:05,389 --> 00:58:12,242
So about a few months ago, I started getting fed all these videos and from the algorithm
on YouTube of people opening locks without keys.

905
00:58:12,242 --> 00:58:16,824
So I decided my pick is literally a pick.

906
00:58:17,205 --> 00:58:18,656
I thought that was too cutesy, right?

907
00:58:18,656 --> 00:58:24,369
So like I literally bought lock picks and took up lock picking as a hobby.

908
00:58:24,369 --> 00:58:27,190
And um my so my actual pick is

909
00:58:27,714 --> 00:58:29,995
The actual um well, what's the name of the company?

910
00:58:29,995 --> 00:58:37,539
Uh their their um their their initials are CI and they have the FNG starter uh covert
instruments, that's it.

911
00:58:37,539 --> 00:58:42,762
And they have the FNG bundle, which stands for Am I allowed to curse on or freaking new
guy, let's just say that.

912
00:58:42,762 --> 00:58:48,074
And I thought that's how I got started, and it comes with some picks and it comes with
some practice locks, and so that's my pick.

913
00:58:48,074 --> 00:58:48,875
It's a lot of fun.

914
00:58:48,875 --> 00:58:52,162
It's literally if you if you want if you you want to do a fidget.

915
00:58:52,162 --> 00:58:54,358
Like if you need a fidget device, it's a great fidget device.

916
00:58:54,358 --> 00:58:59,801
You could I could be on a conference call just picking a lock and just opening it and keep
picking it over and over again.

917
00:59:00,224 --> 00:59:08,345
Are are you are you a prepper and you're prepping for like going out into the world and
like, you know, there's a there's a locked door, you know, there's some there's some food

918
00:59:08,345 --> 00:59:12,686
behind there, I'm gonna have to get in shortly and this is gonna save your life.

919
00:59:12,686 --> 00:59:15,212
There's actually ethics to l to lock sport.

920
00:59:15,212 --> 00:59:19,980
Like if you go on any lock sport forum, like one of the things is don't ever pick a lock
that you don't own.

921
00:59:20,402 --> 00:59:22,405
So we're not supposed to do that.

922
00:59:23,126 --> 00:59:27,489
Are you uh are are you part of any and you like you do the club or social thing as well?

923
00:59:27,489 --> 00:59:27,938
Like

924
00:59:27,938 --> 00:59:28,758
I mean, no, not yet.

925
00:59:28,758 --> 00:59:31,900
I mean I'm just on some Reddit subs and just asking questions.

926
00:59:31,900 --> 00:59:33,441
And there actually is a belt system.

927
00:59:33,441 --> 00:59:33,732
Okay.

928
00:59:33,732 --> 00:59:34,820
Like there's white belt, yellow belt.

929
00:59:34,820 --> 00:59:43,217
I'm a yellow belt, so like I don't know how I got that, but then I'm trying to get to my
orange belt and the orange belt lock is really that's my that's my uh my blocker here, so

930
00:59:43,217 --> 00:59:47,020
I really need to up my game to Yeah exactly.

931
00:59:47,020 --> 00:59:48,220
Yeah, exactly.

932
00:59:48,248 --> 00:59:57,974
I do know uh there's a small overlap with this community, but there's DEF CON meetups
where they often have a lock picking uh session of that you can you can join.

933
00:59:57,974 --> 01:00:04,940
I've seen the uh is there is the the lock picking set is that is that your recommendation
uh or is there like a particular

934
01:00:04,940 --> 01:00:10,330
That's why I mentioned it by name, so I know you need a natural I mean there's lots of
reputable companies out there.

935
01:00:10,330 --> 01:00:16,194
uh the covert instruments was the one that which we was feeding you most of the videos, so
I thought I'd give them some of my business.

936
01:00:16,594 --> 01:00:17,015
Okay.

937
01:00:17,015 --> 01:00:20,677
Well, that will be in the in the in the description for the episode.

938
01:00:20,677 --> 01:00:21,328
There'll be a link there.

939
01:00:21,328 --> 01:00:21,798
All right.

940
01:00:21,798 --> 01:00:22,198
Okay.

941
01:00:22,198 --> 01:00:22,969
I I I love it.

942
01:00:22,969 --> 01:00:31,565
You know, honestly it's this thing that I always wanted to get into and I just I I I think
I bought myself a uh a lockpick set that fit inside of a credit card basically.

943
01:00:31,565 --> 01:00:37,949
And I I I actually used it in exactly one time when I was locked in a room because the
door handle fell off.

944
01:00:37,989 --> 01:00:39,754
And uh I was like, you know what?

945
01:00:39,754 --> 01:00:45,458
I don't have my lock picking set, but I have some paper clips and the practice did help me

946
01:00:45,458 --> 01:00:48,453
uh get out of the room, which was its own experience.

947
01:00:48,453 --> 01:00:50,918
So I I do I do think it is a little bit of a fun passum.

948
01:00:50,918 --> 01:00:53,081
It's much better than buying like a fidget spinner.

949
01:00:53,081 --> 01:00:54,935
Uh you learn a real spin

950
01:00:54,935 --> 01:00:55,717
fidget device.

951
01:00:55,717 --> 01:00:56,389
Yeah, exactly.

952
01:00:56,389 --> 01:00:59,146
And I I've learned that most padlocks are are junk.

953
01:00:59,146 --> 01:01:01,330
Like I can open them and I'm not good.

954
01:01:01,792 --> 01:01:02,593
So

955
01:01:02,848 --> 01:01:11,453
There is a there's a whole set of YouTube videos on basically physical security
reliability where there's a a company that will go out and actually break into your data

956
01:01:11,453 --> 01:01:18,827
center and they taught they share a lot about actual physical security and like the locks
that are being used and how they get into buildings.

957
01:01:18,827 --> 01:01:24,990
And picking locks is very low down on the list of skills that they need to make that
happen.

958
01:01:25,624 --> 01:01:29,116
I there was a funny uh another disaster that's on the wheel of misfortune.

959
01:01:29,116 --> 01:01:36,209
it was at Facebook a couple of years ago where they had an outage and the outage affected
their security system so get into their data center.

960
01:01:36,620 --> 01:01:42,864
actually think we brought that up, but it's it's always it's this I I don't think we need
to do it again, but the the the short of it was, right?

961
01:01:42,864 --> 01:01:45,026
The so you wa you were actually at Facebook at the time.

962
01:01:45,026 --> 01:01:45,998
No, no, no, no.

963
01:01:45,998 --> 01:01:49,310
it's just one of the it's one of the wedges on the wheel of misfortune that we see.

964
01:01:49,310 --> 01:01:50,010
That's all.

965
01:01:50,010 --> 01:01:59,444
Yeah, so I I I love this because it's the not I we were talking about this in the episode
on isolation, I think, where you have your critical systems also sort of depend on

966
01:01:59,444 --> 01:02:00,175
themselves.

967
01:02:00,175 --> 01:02:05,997
And in this case, Facebook messed up their BGP routing, which was required in order to
validate user identity.

968
01:02:05,997 --> 01:02:12,960
So they couldn't fix the BGP routing because they had to get into the data center to reset
it, but they couldn't get into the data center because the doors were locked and required

969
01:02:12,960 --> 01:02:14,680
physical identity checks.

970
01:02:14,771 --> 01:02:17,804
which of course went through their systems, etc., you know, loop.

971
01:02:17,804 --> 01:02:18,577
Loop guaranteed.

972
01:02:18,577 --> 01:02:19,350
Okay.

973
01:02:19,350 --> 01:02:25,541
So next time we're at re you're a reInvent or a summit, you need to come by and spin the
wheel because you'll spin the wheel and you'll say, I know that one and you'll just rattle

974
01:02:25,541 --> 01:02:26,122
off what it is.

975
01:02:26,122 --> 01:02:32,773
You'll really I I I'll make sure if I'm there you get a gift card, whether you get the
quiet day or not, as long as you can identify the r the disaster.

976
01:02:33,038 --> 01:02:37,260
Uh yeah, it's a sp well anyone who's listening to this episode will know the the secret
answer now.

977
01:02:37,260 --> 01:02:39,181
So I I appreciate the pick, Seth.

978
01:02:39,181 --> 01:02:41,572
So uh I brought maybe something lame.

979
01:02:41,572 --> 01:02:50,806
last episode we were talking about changing the whole mindset of how we do development and
architecture work and not just building the product, but also thinking about how we build

980
01:02:50,806 --> 01:02:52,366
and how we read code.

981
01:02:52,366 --> 01:02:56,028
And what went along with that was a book called Rewilding Software Engineering.

982
01:02:56,028 --> 01:02:57,809
It's all about moldable development.

983
01:02:57,809 --> 01:03:01,450
And yeah, it's on medium of all places, so totally free.

984
01:03:01,570 --> 01:03:12,449
And we had the author on and really I I think this is really interesting because it
embodies the challenges that we like we may have historically used to get to the cloud,

985
01:03:12,449 --> 01:03:13,660
like lift and shift.

986
01:03:13,660 --> 01:03:21,306
And if you think about that, uh you may say, well, that's completely ignoring cloud
topology when planning long-lived services.

987
01:03:21,306 --> 01:03:23,628
You're just taking what you have and and go there.

988
01:03:23,628 --> 01:03:26,380
And I think similar challenges also exist for backups.

989
01:03:26,380 --> 01:03:28,992
You don't just straight duplicate all of your infrastructure.

990
01:03:28,992 --> 01:03:30,924
You're really thinking about how.

991
01:03:30,924 --> 01:03:36,556
Cohesively a strategy should look like that solves the I mean, we talked about business
needs a lot.

992
01:03:36,556 --> 01:03:44,779
And I think a lot of companies use VM snapshots um as a as as their strategy, which I'm
sure you'll say, you know, can work in some in some regard.

993
01:03:44,779 --> 01:03:52,501
but I think a lot of companies that do it are not really thinking about the need to
actually back up all of that unnecessary empty block storage they have.

994
01:03:52,542 --> 01:03:52,962
Yeah.

995
01:03:52,962 --> 01:03:53,705
uh

996
01:03:53,705 --> 01:03:54,614
how things fit together.

997
01:03:54,614 --> 01:04:00,226
There's a ton of configuration hidden everywhere that just doesn't work when you bring it
up in a different region and different account.

998
01:04:00,226 --> 01:04:06,168
Yeah, and I think the multiple d software development really thinks like goes to talk to
like how do we build software?

999
01:04:06,168 --> 01:04:08,298
Like how do we actually think about that process?

1000
01:04:08,298 --> 01:04:15,911
And not like the process of do we have stand-ups and is there a testing cycle, but what is
it that we're building and are those things reusable?

1001
01:04:15,911 --> 01:04:23,913
And I I I really like the approach that it takes in the examples that they bring up
because it really hits to the point of we're probably doing it all wrong and we haven't

1002
01:04:23,913 --> 01:04:26,954
really spent any time in the last fifty or so years.

1003
01:04:26,954 --> 01:04:29,890
really rethinking this at any point where we could be.

1004
01:04:29,890 --> 01:04:35,190
And I I think this is one of the secret things that some companies are doing right, but
don't even realize they're doing.

1005
01:04:35,190 --> 01:04:37,890
And a lot of other companies aren't really paying attention to.

1006
01:04:37,890 --> 01:04:39,896
Wow, okay, yeah, definitely have to check that out.

1007
01:04:39,896 --> 01:04:44,068
So um thank you, Seth, for joining us today.

1008
01:04:45,066 --> 01:04:47,018
I I I'm I'm glad I'm glad you enjoyed it.

1009
01:04:47,018 --> 01:04:49,523
I think this is gonna turn out to be a a great episode.

1010
01:04:49,523 --> 01:04:54,076
And here's a reminder for everyone listening, I guess, to subscribe to Adventures in
DevOps.

1011
01:04:54,076 --> 01:04:56,838
You have no idea what what positive impact that has.

1012
01:04:56,838 --> 01:04:59,700
The more people that listen, the better guests that we can have on.

1013
01:04:59,700 --> 01:05:03,372
Like Seth with a lot of veteran years and and hyperscalers.

1014
01:05:03,372 --> 01:05:03,954
Yeah, awesome.

1015
01:05:03,954 --> 01:05:05,608
Really, really glad to be here.

1016
01:05:06,063 --> 01:05:09,580
and I hope I we can see everyone again back next week.

