1
00:00:07,950 --> 00:00:09,821
Welcome back to Adventures in DevOps.

2
00:00:09,821 --> 00:00:18,995
It's already been three years since the community discarded HashiCorp for their devious
practices, and engineering orgs everywhere have been on the lookout for TF alternatives.

3
00:00:18,995 --> 00:00:27,399
I know for us at Authorist, CloudFormation will always hold a special place in my
company's hearts when it should have probably been let go a long time ago.

4
00:00:27,399 --> 00:00:35,576
But we've recently taken a cleared preference towards OpenTOFU, I think, uh for what it's
worth, the spiritual successor of Terraform.

5
00:00:35,576 --> 00:00:41,589
But I wanted to bring in an alternative perspective and a guest who believes in a
technology that I personally refuse to try.

6
00:00:41,589 --> 00:00:49,232
Staff software engineer at Snap, previously Cruz, and AWS, all things cloud
infrastructure, Pushkar Gopala Krishna.

7
00:00:49,232 --> 00:00:50,852
Welcome to the show.

8
00:00:52,373 --> 00:01:00,417
At this point, I can only imagine that uh the last customer of Terraform Somewhere is
finally migrating to OpenTOFU at this point.

9
00:01:00,417 --> 00:01:03,798
Or maybe at least has chosen uh a more nuanced approach.

10
00:01:03,798 --> 00:01:11,678
So uh obviously you want to discuss some of your your experiences and why uh you may not
have even chosen Terraform in the first place.

11
00:01:11,678 --> 00:01:15,900
We just let service developers, you know, define whatever infrastructure they wanted.

12
00:01:15,900 --> 00:01:23,162
It worked in well really well initially, but over time um I guess uh developers started
seeing you know issues with Terraform.

13
00:01:23,162 --> 00:01:34,498
Um first off, like we have to use like a custom language like HashiCop HCL which is which
is a bit of a weird configuration language and outside of Terraform.

14
00:01:34,498 --> 00:01:35,285
Well, I wanna ask about that.

15
00:01:35,285 --> 00:01:36,955
Why is that weird for you?

16
00:01:37,467 --> 00:01:46,413
So for with a configuration language like uh YAML, right, it's much more widely used,
especially if you're in Kubernetes, you know, you have to use it much more a lot.

17
00:01:46,413 --> 00:01:50,016
Uh but HCL is only used for Terraform.

18
00:01:50,016 --> 00:01:54,990
Service developers generally don't, you know, work with infrastructure configurations day
in and day out.

19
00:01:54,990 --> 00:02:02,146
They use it, you know, once to set up some infrastructure and then they're done with that
for the next, you know, few weeks, months and they don't get back to it.

20
00:02:02,146 --> 00:02:02,978
So

21
00:02:02,978 --> 00:02:11,938
Because of that, it's like uh you know, people don't take the time to learn it well and
that contributes to, you know, why people think it's a little weird.

22
00:02:11,938 --> 00:02:22,585
So is it from your perspective there there is a huge burden to getting it right because
there's a lot of complexity in the or maybe there's a huge gap in the historical knowledge

23
00:02:22,585 --> 00:02:26,768
and jumping over that gap to get to a a place of mastery is challenging.

24
00:02:26,768 --> 00:02:27,308
Yeah.

25
00:02:27,308 --> 00:02:27,679
Yeah.

26
00:02:27,679 --> 00:02:32,072
From my limited experience, it it actually hasn't seemed like there's a lot of complexity
in it.

27
00:02:32,072 --> 00:02:39,356
So I'm s maybe I just don't see it and I I just or I'm running into all those pit h
pitfalls, you know, all the time without even knowing about it.

28
00:02:39,356 --> 00:02:41,888
Like, what do people end up

29
00:02:41,888 --> 00:02:43,298
Struggling with the most.

30
00:02:43,298 --> 00:02:45,509
With control structures and looping, right?

31
00:02:45,509 --> 00:02:51,239
So let's let's say you know you want uh a GCS bucket, you go copy paste the configuration
for that.

32
00:02:51,239 --> 00:02:53,781
Now let's say you want five GCS buckets, right?

33
00:02:53,781 --> 00:02:58,812
There is like looping in the HashiCop language, but uh people don't know that syntax very
well.

34
00:02:58,812 --> 00:03:03,384
So what they do is they just copy paste the same bucket configuration five different
times.

35
00:03:03,384 --> 00:03:08,425
Now let's say you know you have to apply uh some sort of a best practice to it later on,
right?

36
00:03:08,425 --> 00:03:12,684
Let's say you have to uh update your lifecycle policy for that GCS bucket.

37
00:03:12,684 --> 00:03:14,595
Now have to do it like five different times.

38
00:03:14,595 --> 00:03:21,943
I've seen a lot of cases where people just end up copy-pasting uh stuff from the internet
or like stuff from you know other repositories.

39
00:03:21,943 --> 00:03:23,905
It's duplicated all over.

40
00:03:23,905 --> 00:03:26,227
And now making a change is hard.

41
00:03:26,227 --> 00:03:35,105
Now being in the infrastructure space, like we wanted to enforce like a lot of best
practices, lifecycle policies with GCS buckets, for instance, or with say spanner

42
00:03:35,105 --> 00:03:38,354
databases, um we had to say set up like autoscaling.

43
00:03:38,354 --> 00:03:42,305
or like backups, point in time recovery, like the right permission models.

44
00:03:42,305 --> 00:03:44,538
service developers wouldn't normally set these things up.

45
00:03:44,538 --> 00:03:47,789
Being in the infrastructure team if we had to do it, it was super painful.

46
00:03:47,789 --> 00:03:57,715
We either had to, you know, request service developers to, you know, do it themselves or
we had to take it upon ourselves and, you know, go manually change a lot of repositories

47
00:03:57,715 --> 00:03:59,510
and that was super painful.

48
00:03:59,510 --> 00:04:08,515
I think that is a really good point because realistically, when you have a bunch of
mistakes that you make, and everyone will make mistakes, it's impossible to get things

49
00:04:08,515 --> 00:04:13,078
right the first time with application code, for instance, you just refactor it.

50
00:04:13,078 --> 00:04:22,714
And I think while maybe HDL is easy to write or easy to copy, when you make a mistake
there, the the penalty for getting it wrong or wanting to change things in the in the

51
00:04:22,714 --> 00:04:25,415
future becomes a huge challenge to actually make that happen.

52
00:04:25,415 --> 00:04:26,826
Because you can't just

53
00:04:27,112 --> 00:04:37,503
merge together five GCS buckets or AWS buckets or whatever have you uh into the same code
block with uh a for each loop or a count loop or whatever you want, you really need to

54
00:04:37,503 --> 00:04:42,448
think about how to merge that code because it's gonna impact your actual production
infrastructure.

55
00:04:42,448 --> 00:04:46,952
Exact I mean guess from my perspective, I would say the code itself is easy to learn.

56
00:04:46,952 --> 00:04:51,062
The the penalties for getting it wrong are very high and it's very easy to fall into
those.

57
00:04:51,062 --> 00:04:57,736
HCL is something that like I said, you know, people write it once, don't visit it, you
know, for another time for a few more weeks or months.

58
00:04:57,736 --> 00:05:00,908
So there's no incentive to learn the details.

59
00:05:00,908 --> 00:05:04,160
But with uh Kubernetes, you have to use like YAML a lot.

60
00:05:04,160 --> 00:05:11,906
Our team, especially, uh, we had a lot of you know Kubernetes experts, but there was like
a more inclination towards YAML when compared to uh HCL.

61
00:05:11,906 --> 00:05:22,225
I think the first thing is uh that you sort of mentioned here is that exposure to the
underlying infrastructure technology or d DSL domain specific language that is being used

62
00:05:22,225 --> 00:05:22,575
here.

63
00:05:22,575 --> 00:05:29,551
And you feel like the application engineers, whoever was building a a new microservice, et
cetera, wouldn't have a lot of experience with HCL.

64
00:05:29,551 --> 00:05:34,746
And while it was easy to copy and paste somewhere else, there's there was a challenge of
getting it right.

65
00:05:34,746 --> 00:05:40,386
At least in my experiences, there has like I don't remember an application engineer that
ever touched anything in YAML.

66
00:05:40,386 --> 00:05:45,830
whatsoever, except maybe if they had to jump into their CI C D platform.

67
00:05:45,830 --> 00:05:49,792
Like if you're using GitHub or GitLab or something else, this is almost always in YAML.

68
00:05:49,792 --> 00:05:55,686
And the interesting thing is that the YAML that GitHub uses is not even standard YAML.

69
00:05:55,686 --> 00:06:04,682
What are application developers doing that would be in YAML that would justify a different
experience for them that would be beneficial over say HCL?

70
00:06:04,682 --> 00:06:13,470
At cruise, the slight difference was that we had like a huge team of uh support resistant
reliability engineers who were super familiar with YAML.

71
00:06:13,470 --> 00:06:19,969
So the support with YAML was much more than the support that application developers
received with HCL.

72
00:06:19,969 --> 00:06:26,936
you're right, like you know, application developers they in general wouldn't use uh YAML a
lot more than they would use HCL.

73
00:06:26,936 --> 00:06:36,629
An experiment that you could make in practice could have been switching all of the areas
where you're writing YAML into writing a DSL that is, I mean, you basically using HCL, but

74
00:06:36,629 --> 00:06:39,663
not for Terraform, but for your actual configuration.

75
00:06:39,663 --> 00:06:41,965
Why not consider that as an option?

76
00:06:42,074 --> 00:06:51,221
it was just that yeah, developers across the company had uh kind of had a bad experience
with HCL and it had just less left a bad taste in our mouth and you know we just didn't

77
00:06:51,221 --> 00:06:53,843
want to, you know, deal with it anymore.

78
00:06:53,843 --> 00:07:04,111
There was this, you know, initiative where um we had to add cost tags to, you know, our
infrastructure and we had you know system reliability engineers who had to take this

79
00:07:04,111 --> 00:07:04,842
initiative.

80
00:07:04,842 --> 00:07:08,034
It took a few, you know, weeks to months and um

81
00:07:08,034 --> 00:07:13,289
During this whole time, like they initially started off with politely asking service
developers, you know, to do it themselves.

82
00:07:13,289 --> 00:07:19,424
But as you know, you know, like service developers are usually tied up with, you know,
their own projects and initiatives.

83
00:07:19,424 --> 00:07:24,188
Updating HCL code is one of their least favorite things and didn't happen for a lot of
teams.

84
00:07:24,188 --> 00:07:27,631
These SREs were forced to do it, you know, on their own.

85
00:07:27,631 --> 00:07:35,882
They tried automating it and, you know, they tried writing scripts to do that, but it did
result in in a few cases where like they

86
00:07:35,882 --> 00:07:42,895
updated the HCL configuration wrongly and it led their workspaces breaking and it it was
it was a disaster.

87
00:07:42,895 --> 00:07:52,350
So the Terraform applies stopped running and uh there were some critical services that
this happened to and essentially all their like deployments were blocked.

88
00:07:52,598 --> 00:07:56,339
I mean, I I totally see this showing up in it's not just the infrastructure world.

89
00:07:56,339 --> 00:08:07,865
I I do see a lot of organizations that don't fully embody the the DevOps mindset and try
to have like there's like a security team or a QA team or a platform engineering team or a

90
00:08:07,865 --> 00:08:18,299
marketing team, sales, whatever, they have a requirement to change the infrastructure, the
the software that basically the organization owns, but the only people that are sort of

91
00:08:18,299 --> 00:08:20,054
empowered to make the change are

92
00:08:20,054 --> 00:08:22,045
aren't part of the decision making process.

93
00:08:22,045 --> 00:08:32,470
So you have one team that's uh basically put out on a ledge, the maybe it's the at site
reliability engineers, who are told we need to have those cost allocation tags everywhere

94
00:08:32,470 --> 00:08:41,295
so that we know how much the infrastructure is costing for every or also the application
is actually running so that we can, I don't know what the end game is here, basically fire

95
00:08:41,295 --> 00:08:43,620
some engineers for making really expensive

96
00:08:43,960 --> 00:08:47,420
technology that doesn't actually impact the bottom line at all, potentially.

97
00:08:47,420 --> 00:08:48,332
I I don't know.

98
00:08:48,332 --> 00:08:54,570
Or maybe chargebacks to particular customers through some sort of uh aggregate usage
depending on what product you're running.

99
00:08:54,570 --> 00:09:04,938
Uh but the end of the story is that that initiative is pushed through a part of the
organization that doesn't have the capability or responsibility or accountability for

100
00:09:04,938 --> 00:09:06,459
impactful changes.

101
00:09:06,459 --> 00:09:12,001
And in order for that to happen, they try to plead very carefully with those in charge.

102
00:09:12,001 --> 00:09:13,752
And of course other teams

103
00:09:13,752 --> 00:09:18,656
Who don't aren't part of the same initiatives have no incentive to make those changes.

104
00:09:18,656 --> 00:09:23,349
And the teams that obviously, some of them will acquiesce and say, Yeah, sure, fine, we'll
do it.

105
00:09:23,550 --> 00:09:28,975
And do something that they're not very comfortable with doing or doesn't don't do very
frequently.

106
00:09:28,975 --> 00:09:37,762
And I think this is just any technology honestly can run uh foul in this area where or any
business making uh choice will happen here.

107
00:09:37,762 --> 00:09:43,278
And so, like what I really hear fundamentally is it isn't so much a problem on the
application side.

108
00:09:43,278 --> 00:09:51,462
where you're building stuff, but really who is fundamentally responsible for making code
changes for or whatever you want to call them, configuration changes, who is fundamentally

109
00:09:51,462 --> 00:09:52,642
accountable for that?

110
00:09:52,642 --> 00:09:59,265
And aligning the ownership with the accountability with whoever is going to perform the
action.

111
00:09:59,265 --> 00:10:04,927
If you don't have that, no matter what technology choice you've chosen, you could easily
fall in into a pit of failure.

112
00:10:04,927 --> 00:10:12,070
Uh if I could get your perspective here on exactly how cross plane works for anyone who uh

113
00:10:12,374 --> 00:10:14,830
Also has never used it fundamentally.

114
00:10:14,830 --> 00:10:17,071
Crossplane a Kubernetes native tool.

115
00:10:17,071 --> 00:10:23,192
uh Essentially, you can represent your infrastructure as Kubernetes custom resources.

116
00:10:23,290 --> 00:10:30,494
Service developers typically manage their configurations as YAML in a Git repository,
packaged up as like a Helm chart.

117
00:10:30,494 --> 00:10:40,907
When you run kubectl apply, any infrastructure configuration that you can have on the
actual infrastructure in, say, GCP or Azure, that is now represented as a custom resource.

118
00:10:40,907 --> 00:10:44,718
When you make a change to your custom resource and when you apply it, the

119
00:10:44,718 --> 00:10:56,452
Kubernetes controller is continuously watching these custom resources and uh when it
detects a change, so it essentially compares the state that is in the custom resource and

120
00:10:56,452 --> 00:10:58,823
it compares it to the state of the external resource.

121
00:10:58,823 --> 00:11:07,042
When it sees a difference, it tries to reconcile it and essentially get the state of the
external resource to match the state of the custom resource.

122
00:11:07,042 --> 00:11:12,614
the infrastructure changes is asynchronous from the deployment strategy there.

123
00:11:12,614 --> 00:11:22,783
It sounds like is you make the change, it gets deployed to the Kubernetes cluster, which
picks it up and then goes and actually applies that change to the running infrastructure

124
00:11:22,783 --> 00:11:23,263
that you have.

125
00:11:23,263 --> 00:11:33,571
So if you add a new load balancer to your configuration, then when it gets when the the
YAML file gets s consumed by the cluster, you will get that load balancer spun up in

126
00:11:33,571 --> 00:11:35,412
Azure, AWS, et cetera.

127
00:11:35,574 --> 00:11:44,124
And how do companies deal with the fact that fundamentally there's no that process is
asynchronous from the actual committed change?

128
00:11:44,124 --> 00:11:53,655
Like normally I would expect when you commit an infrastructure change, it's running in
your CI C D platform and that's being the blocking on the fact that those changes are

129
00:11:53,655 --> 00:11:54,186
getting made.

130
00:11:54,186 --> 00:11:55,886
How do you do that in cross plane?

131
00:11:55,886 --> 00:12:00,570
When a custom resource is changed, it is immediately reconciled as well.

132
00:12:00,570 --> 00:12:07,297
So when you run kubectl apply, it does synchronously try to make the change.

133
00:12:07,297 --> 00:12:17,464
And only once the change is done, the status of the custom resource or the claim or the
composite resource is updated with the status of the external resource.

134
00:12:17,464 --> 00:12:25,497
So how do you manage something like projects in in G C P or the directories in Azure or
the accounts in AWS?

135
00:12:25,497 --> 00:12:30,358
If you don't have a Kubernetes cluster, like how do you spin up a new AWS account?

136
00:12:30,358 --> 00:12:35,630
How do you, you know, get the new project created or get the right IAM profiles
integrated?

137
00:12:35,630 --> 00:12:37,241
Like how do you manage that part?

138
00:12:37,241 --> 00:12:41,412
Because obviously there's no cluster running at this point to even execute that in the
first place.

139
00:12:41,412 --> 00:12:42,132
So

140
00:12:42,571 --> 00:12:50,838
so at Cruz the way it worked was um we started off with Terraform and we built a bunch of
uh infra Kubernetes clusters.

141
00:12:50,838 --> 00:12:52,720
Those are still managed with Terraform.

142
00:12:52,720 --> 00:12:55,002
It's just the base infrastructure clusters.

143
00:12:55,002 --> 00:12:59,826
And um so we have another tool that spins up uh G C P projects.

144
00:12:59,826 --> 00:13:00,450
So that

145
00:13:00,450 --> 00:13:02,352
That part is actually outside of Crossplane.

146
00:13:02,352 --> 00:13:04,063
And there are historical reasons for it.

147
00:13:04,063 --> 00:13:12,971
So we had this tool that was built a long time ago called Juno, and that was responsible
for spinning up a service and like all the associated infrastructure, you know, with the

148
00:13:12,971 --> 00:13:13,352
service.

149
00:13:13,352 --> 00:13:16,584
And as a part of it would spin up the GCP project.

150
00:13:16,584 --> 00:13:20,358
And it also managed the RBAC or the permissions piece of it.

151
00:13:20,358 --> 00:13:22,970
Permissions were all maintained outside of Crossplane.

152
00:13:22,970 --> 00:13:27,492
Only the so we use Crossplane only for spinning up um, like I said previously.

153
00:13:27,492 --> 00:13:28,752
I mean maybe let's dive into that.

154
00:13:28,752 --> 00:13:30,450
So there's like, you know, some

155
00:13:30,450 --> 00:13:37,555
uh golden path for setting up the account and then you switch over to a different
technology for the next layer and then switch over to a different technology for the layer

156
00:13:37,555 --> 00:13:38,116
after that.

157
00:13:38,116 --> 00:13:47,212
When you get down to the application layer though, how do you handle the bootstrapping
process for even knowing about the YAML file for the application to run?

158
00:13:47,212 --> 00:13:56,378
Because historically in a lot of organizations, if you are building microservices, every
single microservice is in a separate repository and you include the infrastructure changes

159
00:13:56,378 --> 00:13:58,850
that are required for that microservice with

160
00:13:58,850 --> 00:14:00,451
the application code.

161
00:14:00,451 --> 00:14:06,944
And so if you're in a repository, it's got your hypothetical YAML file or HCL for
deploying that application.

162
00:14:06,944 --> 00:14:15,319
And so when you want to add a new database, you add it directly into your repository in
maybe a templates directory or a deployment directory, et cetera.

163
00:14:15,319 --> 00:14:26,325
But with that, how would Crossplane even know or the Kubernetes cluster you're running
realistically, know to go look in this new GitHub or GitLab repository to go pick up the

164
00:14:26,325 --> 00:14:27,816
code to even deploy in the first place?

165
00:14:27,816 --> 00:14:28,268
Yeah.

166
00:14:28,268 --> 00:14:28,878
Yeah.

167
00:14:28,878 --> 00:14:36,081
So when developers created a service with Juno, it would spin up like you know uh all the
necessary infrastructure.

168
00:14:36,081 --> 00:14:43,811
So you need the compute infrastructure, you need the observability infrastructure, you
need like secrets, you need like the networking, the service mesh.

169
00:14:43,811 --> 00:14:47,806
There's there's a lot of Juno was responsible for that.

170
00:14:47,806 --> 00:14:56,364
It would spin up all these different components and additionally it would create a Git
repository and it would also like create the CICD tooling with it.

171
00:14:56,364 --> 00:15:03,992
So when a new service was created, it would essentially take this golden path code and
write it to the newly created Git repository for the service.

172
00:15:03,992 --> 00:15:09,684
Juno sounds more like less like a tool and more potentially like a a platform for creating
applications.

173
00:15:09,684 --> 00:15:11,305
Is that yeah, yeah.

174
00:15:11,305 --> 00:15:11,745
Okay.

175
00:15:11,745 --> 00:15:12,095
Okay.

176
00:15:12,095 --> 00:15:16,117
So basically, you know, it sounds like if anything, you don't love Crossplane.

177
00:15:16,117 --> 00:15:16,967
You love Juno.

178
00:15:16,967 --> 00:15:19,228
Juno seems like the best thing ever.

179
00:15:19,488 --> 00:15:24,738
It it's one of these platform engineering platform platforms that you know creates

180
00:15:24,738 --> 00:15:26,719
platforms for application deployment.

181
00:15:26,719 --> 00:15:28,400
And it seems like it's doing everything that you need.

182
00:15:28,400 --> 00:15:32,973
It's doing the account management or project management, um, resource creation.

183
00:15:32,973 --> 00:15:39,797
It's doing the orchestration for, you know, kicking off Crossplane for generating
repositories for application code and everything.

184
00:15:39,797 --> 00:15:47,331
At the end of the day, all the only part there that's still running in Crossplane is the
maybe specific application level infrastructure.

185
00:15:47,331 --> 00:15:50,373
Like at that point, why even why even use Crossplane then?

186
00:15:50,373 --> 00:15:52,418
Like why not just like also have

187
00:15:52,418 --> 00:15:57,647
Juno, you know, kick off some sort of container running in in the cluster to go execute.

188
00:15:57,647 --> 00:15:59,598
Why wh why not customize that too?

189
00:15:59,598 --> 00:16:03,320
Uh Juno was was, you know, a beast of its own.

190
00:16:03,320 --> 00:16:08,133
It it had a lot of pros and cons and uh it was built a long time ago.

191
00:16:08,133 --> 00:16:12,125
And developers, you know, some developers loved it, some developers hated it.

192
00:16:12,125 --> 00:16:21,624
And usually have that right, like with any, you know, technology or platform that's been
around for uh a few years, like say seven, eight years, um

193
00:16:21,624 --> 00:16:23,305
people have strong opinions about that.

194
00:16:23,305 --> 00:16:31,800
It was uh getting to a point where like, you know, it was getting hard to scale up Juno to
manage, you know, all these in different infrastructure components because it was a lot.

195
00:16:31,800 --> 00:16:38,423
There were a lot of services, there were lot of environments within these services and it
it was a lot for Juno to manage.

196
00:16:38,423 --> 00:16:43,776
I mean, you can make it work given that you invest enough in a tool, you can still make it
work.

197
00:16:43,776 --> 00:16:51,520
But we had reached a point where you know, we felt that it was easier to manage
infrastructure with cross plane than have, you know, Juno do that.

198
00:16:51,732 --> 00:16:53,708
especially with the legacy code that it had.

199
00:16:53,708 --> 00:16:56,189
I mean, it makes a lot of sense if you come from the other perspective.

200
00:16:56,189 --> 00:17:02,602
Like rather than from the ground up of you need to deploy infrastructure for your
organization and that and included in that is everything.

201
00:17:02,602 --> 00:17:10,996
And if you come from the the place where you have a custom tool that's doing your
infrastructure deployments and managing repositories and and whatnot, and from there

202
00:17:10,996 --> 00:17:18,210
you're like, How can we make our custom tool, which has a terrible internal development
process lifecycle where it's one managed by

203
00:17:18,210 --> 00:17:25,474
two people on different teams that have no obligation or OKR or KPI behind actually
improving the tool to do what it needs.

204
00:17:25,474 --> 00:17:29,466
What if we just rip out all of that stuff and replace it with an open source project?

205
00:17:29,466 --> 00:17:32,928
Then I can totally see like, okay, what's available out there?

206
00:17:32,928 --> 00:17:39,288
And that list is very short for lists of technology to actually delegate infrastructure
deployment to.

207
00:17:39,288 --> 00:17:43,522
Course, there are you know cases where cross plane cross plane is not the solution for
everything, right?

208
00:17:43,522 --> 00:17:48,007
Like it it's it's still like when compared to Terraform, it's still maturing.

209
00:17:48,007 --> 00:17:57,056
When we initially tried out cross plane, of course we you know prototyped it on a small
set of resources and resources and we thought that it was great and you know we started

210
00:17:57,056 --> 00:17:57,877
adopting it.

211
00:17:57,877 --> 00:18:01,550
later we realized that you know there there are issues with cross plane as well.

212
00:18:01,550 --> 00:18:03,508
So Terraform has been around for like

213
00:18:03,508 --> 00:18:04,889
more than 10 years, right?

214
00:18:04,889 --> 00:18:15,758
Like it has matured a lot and you know there are a lot of providers, like any cloud
resource or SaaS product that you can think of usually has a Terraform provider and you

215
00:18:15,758 --> 00:18:18,103
should you know should be able to manage it with Terraform.

216
00:18:18,103 --> 00:18:19,884
Crossplane is not there yet.

217
00:18:19,884 --> 00:18:25,078
So uh we couldn't really migrate completely to Crossplane because of that.

218
00:18:25,078 --> 00:18:30,790
Not every provider out there has a controller available in in Crossplane.

219
00:18:30,790 --> 00:18:32,876
So that must have caused some problems for you.

220
00:18:32,876 --> 00:18:37,749
We kind of realized that uh a little bit later that, you know, they tend to have
providers.

221
00:18:37,749 --> 00:18:47,576
So um so we were stuck in in a world where like, you know, we had to maintain Terraform as
well, because you know, these resources couldn't be migrated to cross plane.

222
00:18:47,576 --> 00:18:54,110
Also, like I think uh our expectation was that we we didn't want to move completely off of
Terraform, right?

223
00:18:54,110 --> 00:18:56,101
Like that that wasn't the goal to start off with.

224
00:18:56,101 --> 00:18:57,522
We only wanted to

225
00:18:57,868 --> 00:19:01,382
manage like say eighty percent of the infrastructure uh if possible.

226
00:19:01,382 --> 00:19:07,919
We wanted to move over like the the hardest eighty percent of the or like the most
commonly used eighty percent of the infrastructure to cross plane.

227
00:19:07,919 --> 00:19:13,285
We kind of expected that there would be like a long tail of you know components that would
uh remain in terraform.

228
00:19:13,285 --> 00:19:16,056
That would at least like reduce our terraform cost.

229
00:19:16,056 --> 00:19:23,038
So one of the challenges that usually comes up in any infrastructure as code solution is
how you manage additional environments.

230
00:19:23,038 --> 00:19:29,550
Like I I know there's always a question of like, okay, we need to spin up a temporary
environment for every single pro pull request that gets generated.

231
00:19:29,550 --> 00:19:37,393
We usually need maybe uh a temporary one that lives a little bit longer than a branch for
extensive load testing and then will get automatically terminated when it's done.

232
00:19:37,393 --> 00:19:42,306
And then of course, in every engineer's uh personal or I mean

233
00:19:42,306 --> 00:19:51,190
company provided but dedicated AWS account or cloud account that they're using, they need
may need to deploy the exact same infrastructure in in some way to to run the individual

234
00:19:51,190 --> 00:19:51,750
services.

235
00:19:51,750 --> 00:19:55,640
And maybe there's some sandboxes as well if they're doing some sort of local development.

236
00:19:55,640 --> 00:19:58,232
So Juno was our savior at that point of time.

237
00:19:58,232 --> 00:20:04,438
Uh so essentially Juno had this notion of environments and we could have, you know, three
different environments.

238
00:20:04,438 --> 00:20:11,154
So we uh built in the mechanism to create uh sub environments within a given environment.

239
00:20:11,154 --> 00:20:12,986
So actually even called it stacks.

240
00:20:12,986 --> 00:20:17,602
It wasn't super ideal, but we we were able to get the level of isolation that we wanted.

241
00:20:17,602 --> 00:20:28,370
The challenges that always comes up within the environment management is realistically
domain separation, where like you have your production domain, which is like some company

242
00:20:28,370 --> 00:20:34,715
dot dot com or whatever, and then your non production domains, which are like test dot
production dot com.

243
00:20:34,715 --> 00:20:38,258
And it's like, well, that lives in the production environment in order to resolve it.

244
00:20:38,258 --> 00:20:41,381
I don't see any way that could have ever been completely isolated.

245
00:20:41,381 --> 00:20:43,118
And I think you were sort of hinting at that.

246
00:20:43,118 --> 00:20:53,547
Yeah, so um we did our best to provide the level of isolation that we needed, but there
were cases where say permissions for instance were definitely shared across, you know,

247
00:20:53,547 --> 00:20:55,938
different environments and that actually made sense.

248
00:20:55,938 --> 00:21:02,058
We didn't want, you know, separate permissions in broad versus development and uh so there
like, you know, there was some commonality.

249
00:21:02,058 --> 00:21:07,277
Uh at the network level, yeah, we couldn't really get the level of isolation that we
wanted.

250
00:21:07,277 --> 00:21:07,960
Between like

251
00:21:07,960 --> 00:21:18,547
Dev and staging in prod, we still had like separate VPCs and you know the services would
run across separate VPCs, but within different uh stacks uh in an environment, they would

252
00:21:18,547 --> 00:21:20,110
end up running in the same VPC.

253
00:21:20,110 --> 00:21:25,792
I I think this is the sort of thing that a lot of companies end up getting wrong though,
where they believe they need this level of isolation there.

254
00:21:25,792 --> 00:21:34,436
And the my response is always like, Do you think you're using a test version of GCP or AWS
for your non-production environment?

255
00:21:34,436 --> 00:21:34,916
No.

256
00:21:34,916 --> 00:21:39,518
You are using the the one that's on their public page and you're creating a real

257
00:21:39,638 --> 00:21:48,094
you cloud account for that and using real stuff there and paying real money for your
resources, there's no reason why that shouldn't be at every other part of the of the stack

258
00:21:48,094 --> 00:21:48,625
as well.

259
00:21:48,625 --> 00:21:57,291
That's when I see some, you know, lesser experienced engineers like mind shatter, like,
yeah, uh I guess that's true, right?

260
00:21:57,291 --> 00:22:00,353
Like we don't have different stuff there.

261
00:22:00,353 --> 00:22:02,174
I think this happens a lot in clue Kubernetes.

262
00:22:02,174 --> 00:22:07,458
I I think one of the one of the core challenges is a lot of companies fall into the pit
where

263
00:22:07,458 --> 00:22:16,840
There's a few low number managed uh Kubernetes clusters for the entire company or
organization level, and everyone deploys to the same cluster.

264
00:22:16,881 --> 00:22:26,513
And I I think while it's doable for sure, you do end up with a lot of ownership issues and
you end up with having to tag a lot of things and then running amok of the IAC tools.

265
00:22:26,513 --> 00:22:35,896
And so we we know fundamentally first class isolation at low down as you can is is
beneficial, but you shouldn't push it any lower that than you absolutely need to.

266
00:22:35,896 --> 00:22:40,898
And I think using infrastructure level services is for one uh for sure, like things like
you mentioned.

267
00:22:40,898 --> 00:22:43,072
I think login and access control is another one.

268
00:22:43,072 --> 00:22:53,170
Like people being able to log in into the test environment to validate things, it's still
useful to validate your production login and access control strategy and not use a fake

269
00:22:53,170 --> 00:22:54,041
version of it.

270
00:22:54,041 --> 00:22:55,912
Because you know those things aren't ready for production.

271
00:22:55,912 --> 00:23:05,462
I think it was uh a couple of companies ago that I was advising where they kept running
into this problem where every time one team wanted to do something,

272
00:23:05,462 --> 00:23:08,683
They were utilizing APIs from a second team.

273
00:23:08,683 --> 00:23:13,245
So like there was the mobile and app development team and the website UI team.

274
00:23:13,245 --> 00:23:22,309
And they, for their non-production version of their UIs, were calling non-prod versions of
the underlying services, which had tons of things broken in them.

275
00:23:22,309 --> 00:23:24,810
It had tons of features that hadn't been released yet.

276
00:23:24,810 --> 00:23:27,931
And you're not coding against like production realistically.

277
00:23:27,931 --> 00:23:31,513
And so it was a huge challenge in those scenarios.

278
00:23:31,513 --> 00:23:35,138
And because of that, they had a lot of hacks in place.

279
00:23:35,138 --> 00:23:43,404
Where whenever you logged into a non production environment, you would also log into
production at the same time and have two tokens that would get passed around.

280
00:23:43,404 --> 00:23:48,618
And then the services would have to like figure out in the load balancer like which token
to send at the right time.

281
00:23:48,618 --> 00:23:56,714
Uh and then there was like VPNs just for non production environments because they wanted
to use a test version of their login and access control strategy.

282
00:23:56,714 --> 00:24:00,077
Genius really when you get when you get down to the bottom of it.

283
00:24:00,077 --> 00:24:01,768
So I I think I think I'm with you.

284
00:24:01,768 --> 00:24:02,072
Like

285
00:24:02,072 --> 00:24:06,548
Figuring out where the where the isolation should be and at what level makes a lot of
sense.

286
00:24:06,548 --> 00:24:12,776
If you're running a single cluster or a set of them, isolating at the VPC level doesn't
make uh it doesn't have a lot of benefit there.

287
00:24:12,776 --> 00:24:15,919
It doesn't it doesn't align with the infrastructure choices you have.

288
00:24:16,080 --> 00:24:17,870
but I think that's a that's a trade off.

289
00:24:17,870 --> 00:24:26,465
I've seen those patterns as well where like uh in the staging environment, you know,
develop like there are services that end up calling production and so they've seen cases

290
00:24:26,465 --> 00:24:31,544
where like, you know, integration tests have ended up writing to actual production
databases because of that and

291
00:24:31,544 --> 00:24:32,124
That's perfect.

292
00:24:32,124 --> 00:24:32,985
That's what you want, right?

293
00:24:32,985 --> 00:24:39,248
Because uh the integration test should have the right permissions in the production
environment and those permissions should be none, right?

294
00:24:39,248 --> 00:24:46,611
And so when it's right, when it tries to write, you just validate the right, you know,
does a person ha or the entity that's calling have access?

295
00:24:46,611 --> 00:24:48,558
Uh I mean, we do the exact same thing.

296
00:24:48,558 --> 00:24:55,325
Uh all of our services call the production version of the other of their dependencies
during development.

297
00:24:55,325 --> 00:24:56,726
Like developers running a

298
00:24:56,726 --> 00:25:01,811
uh application client on their machine will call the prod version because it's the only
version that you can trust.

299
00:25:01,811 --> 00:25:04,003
And they go through our access control strategy.

300
00:25:04,003 --> 00:25:09,229
And if and I think there's this mistake of like you isolate to protect your environment.

301
00:25:09,229 --> 00:25:10,520
You don't isolate to protect.

302
00:25:10,520 --> 00:25:14,955
You isolate to s single out for testing purposes, right?

303
00:25:14,955 --> 00:25:17,327
Because you want to figure out where the problem is.

304
00:25:17,327 --> 00:25:20,460
You don't I I mean unless you're doing a load test, then

305
00:25:20,460 --> 00:25:22,201
Then there's a question of what happens there.

306
00:25:22,201 --> 00:25:28,337
But I love the load test that failed because the non-production cluster doesn't have the
same resources as the production cluster.

307
00:25:28,337 --> 00:25:30,048
And so of course it fails there.

308
00:25:30,048 --> 00:25:31,770
And so what are you really testing?

309
00:25:31,770 --> 00:25:40,184
So I I mean there I don't know how you do that effectively without cloning your production
environment and spinning up exactly the same thing and then running the test there.

310
00:25:40,184 --> 00:25:41,715
That's a hard problem to solve, right?

311
00:25:41,715 --> 00:25:43,904
Like there are a lot of options that you can try out.

312
00:25:43,904 --> 00:25:51,839
You can essentially uh replicate your entire production stack and uh have like a complete
replica of it running and you can run your load tests over there.

313
00:25:51,839 --> 00:25:54,780
But again at that point you're like wasting resources, right?

314
00:25:54,780 --> 00:26:00,630
If you have a complete copy of your production but you're not using it for live production
traffic, you're incurring a lot of cost for no

315
00:26:00,630 --> 00:26:01,501
It's just temporary.

316
00:26:01,501 --> 00:26:06,365
First of all, like if you're running if you just have an environment that's long lived
besides prod, you're doing it wrong.

317
00:26:06,365 --> 00:26:07,495
Shut it down.

318
00:26:07,656 --> 00:26:09,317
I guess that would be the first one.

319
00:26:09,377 --> 00:26:12,800
the second one, like 'cause you're not you're not you shouldn't be using it con
consistently like that.

320
00:26:12,800 --> 00:26:16,143
And if you are, then you must be getting the value out of it for sure.

321
00:26:16,784 --> 00:26:17,564
Um yeah.

322
00:26:17,564 --> 00:26:19,256
And I think really that's the core aspect.

323
00:26:19,256 --> 00:26:19,736
So they

324
00:26:19,736 --> 00:26:22,837
was a point of time where we didn't really use Argo CD.

325
00:26:22,837 --> 00:26:25,888
We're using this uh CI tool called build kite.

326
00:26:25,888 --> 00:26:32,730
Um it was a great CI tool, but we were using it for doing continuous deployments when it
was you know not meant to do that.

327
00:26:32,730 --> 00:26:43,463
So which meant that we had to handle uh a lot of the moving pieces ourselves, like you
know, authentication and you know, making sure you know this thing scaled up and it had

328
00:26:43,463 --> 00:26:45,284
like you know all the namespacing in it.

329
00:26:45,284 --> 00:26:46,306
so uh

330
00:26:46,306 --> 00:26:52,971
We spent a lot of time uh debugging pipeline issues and it was like super frustrating.

331
00:26:52,971 --> 00:27:01,267
You also had to configure like the authentication from your pipeline, like from the from
the build guide, people would get that wrong and until we, you know, got fed up with that

332
00:27:01,267 --> 00:27:07,221
and we said, Hey, you know, like this is not working out and we actually ended up moving
to Argo C D at that point of time.

333
00:27:07,310 --> 00:27:12,064
It's all about reducing the on call support that the platform team needs to do.

334
00:27:12,064 --> 00:27:17,778
You know, the less on call time they need to handle stuff, the the happier everyone is.

335
00:27:18,726 --> 00:27:20,128
For sure, for sure.

336
00:27:20,128 --> 00:27:21,218
There was support.

337
00:27:21,218 --> 00:27:27,543
there were like non-urgent support issues which our team got a lot and there were like
these on call issues as well.

338
00:27:27,543 --> 00:27:37,880
Uh the support issues were like super frustrating and we spent a lot of time initially we
had a lot of that and we uh actually invested quite a lot of time, you know, figuring out

339
00:27:37,880 --> 00:27:46,065
like where exactly we were spending time when supporting these different teams and Argo CD
was one of the things that we did to you know reduce our time

340
00:27:46,322 --> 00:27:49,642
The straw that broke the camel's back when it came to build kite?

341
00:27:49,642 --> 00:27:56,528
Like, how did you actually decide that, like today in 2026, the the thing that you do, of
course, is use OIDC, right?

342
00:27:56,528 --> 00:28:01,331
The build server itself generates its own JWTs automatically, injects them into the
process.

343
00:28:01,331 --> 00:28:10,446
You use a standard couple of lines of code in your CICD YAML file, which automatically
authenticates to your cloud provider of choice, and you can automatically get the right

344
00:28:10,446 --> 00:28:12,237
role and permissions and execute stuff.

345
00:28:12,237 --> 00:28:14,838
And obviously, even a few years ago,

346
00:28:15,006 --> 00:28:17,747
GitHub didn't support didn't support OIDC.

347
00:28:17,747 --> 00:28:18,937
GitLab didn't support OIDC.

348
00:28:18,937 --> 00:28:25,909
Uh interestingly enough, I actually was the one who was reviewing the GitLab
implementation uh before even GitHub supported it.

349
00:28:25,909 --> 00:28:35,792
It was amazing that these providers that uh honestly usually have full admin access to
your entire AWS account, your entire GCP account, a project, etc.

350
00:28:35,792 --> 00:28:44,362
And they're being authenticated with basically plain text credentials that are saved in a
environment variables inside a third party provider.

351
00:28:44,362 --> 00:28:49,194
Obviously no one likes what was happening at this point, but most companies don't do
anything about it.

352
00:28:49,194 --> 00:28:51,836
So we actually didn't move off of it completely.

353
00:28:51,836 --> 00:28:53,116
No no.

354
00:28:54,237 --> 00:28:57,439
So it was so build kit is a great CI tool, right?

355
00:28:57,439 --> 00:28:59,460
Like um I mean I don't want to understate that.

356
00:28:59,460 --> 00:29:03,373
It's definitely a good build CI tool, but not a great C D tool.

357
00:29:03,373 --> 00:29:10,647
You know, CI is like the you know, you meet a change right from the point where your code
leaves your laptop.

358
00:29:10,647 --> 00:29:12,839
That is where the CI part starts, right?

359
00:29:12,839 --> 00:29:15,564
It should take care of like building your you know

360
00:29:15,564 --> 00:29:22,778
image, it should take care of like you know uploading it to some sort of artifact registry
and uh should take care of running uh any sorts of tests that you have, right?

361
00:29:22,778 --> 00:29:28,442
Like unit tests, component tests, integration tests, like load tests it you know that
you're gonna deploy.

362
00:29:28,442 --> 00:29:35,374
Once it is ready, I feel like that's where the responsibility of CI stops and the
responsibility of C D starts.

363
00:29:35,374 --> 00:29:37,187
Isn't there a little bit of a paradox there?

364
00:29:37,187 --> 00:29:47,895
Because like how do you know that you're done building that artifact without actually
deploying it to production in a way or in an isolated production like environment where

365
00:29:47,895 --> 00:29:49,570
it's getting real requests in?

366
00:29:49,570 --> 00:29:53,374
The way you do that is through slow rollouts, right, or Canary, right?

367
00:29:53,374 --> 00:30:02,264
And uh there are some cases where you can't figure out issues with your code until you
actually deploy it.

368
00:30:02,264 --> 00:30:08,694
Interestingly, although we introduced like Argo rollout for deploying service
infrastructure in a progressive manner.

369
00:30:08,694 --> 00:30:17,663
There there were interesting issues that you know we faced where because of a lack of a
slow rollout mechanism we've we've caused outages, you know, widespread outages across

370
00:30:17,663 --> 00:30:19,035
like a lot of different teams.

371
00:30:19,035 --> 00:30:23,089
Case where while while I was working on this uh stacks, you know, initiative.

372
00:30:23,089 --> 00:30:30,056
So at that point of time like uh we had the ability to uh slow roll our feature to certain
projects only.

373
00:30:30,056 --> 00:30:31,692
What happened was um

374
00:30:31,692 --> 00:30:34,833
We made a change and um so there was a bug in the change.

375
00:30:34,833 --> 00:30:42,546
What that resulted in happening was like the internal state of the Kubernetes namespace
was accidentally marked for deletion.

376
00:30:42,546 --> 00:30:50,250
So after we wiped out like three different production services as soon as our change was
rolled out.

377
00:30:50,250 --> 00:30:52,160
And it was all of a sudden, right?

378
00:30:52,160 --> 00:30:55,031
Like these service teams, these are like critical services.

379
00:30:55,031 --> 00:31:01,624
Um so for for some context, like at at Cruz we had these, you know, driverless cars that
operated on the road that, you know,

380
00:31:01,624 --> 00:31:07,798
called these services and it was like the level of reliability expected was really,
really, really high.

381
00:31:07,939 --> 00:31:17,518
Yeah, we accidentally wiped out like the whole production service, for like, you know,
three different services and god, that that was

382
00:31:17,518 --> 00:31:25,286
If I go like make a internet search right now, is there gonna be like a particular date
where like all the cruise cars were getting into accidents because of this incident?

383
00:31:25,286 --> 00:31:29,130
Um I don't uh no.

384
00:31:29,130 --> 00:31:32,172
uh which of course no.

385
00:31:32,172 --> 00:31:33,754
I don't think I'll comment on that.

386
00:31:33,754 --> 00:31:34,498
Yeah.

387
00:31:34,498 --> 00:31:44,621
The I mean, with such a high reliability requirement, at least for us, we look at both the
ability to reduce the probability, but also the ability to reduce the impact because we

388
00:31:44,621 --> 00:31:46,691
see the risk as a as a product of these two.

389
00:31:46,691 --> 00:31:51,563
I mean, obviously it's just not a straight multiplication, but you can look at it
similarly, right?

390
00:31:51,563 --> 00:31:52,953
as far as expected value goes.

391
00:31:52,953 --> 00:31:58,845
So if you're thinking about reliability, and obviously autonomous driving is super
critical uh from a reliability standpoint.

392
00:31:58,845 --> 00:32:04,300
What was the sort of mentality around dealing with the reducing the impact?

393
00:32:04,300 --> 00:32:14,040
side or potentially dealing with it in a reactive nature, like how did the vehicles handle
scenarios where they couldn't connect back to the production control plane?

394
00:32:14,040 --> 00:32:19,245
Initially there were issues, but uh as you know, like we were in a R and D phase for a
really long time.

395
00:32:19,245 --> 00:32:26,541
That that was when uh these sorts of issues were identified and so we did have humans like
human drivers in the cars who would operate the cars.

396
00:32:26,541 --> 00:32:29,404
If this sort of thing happened, like there was a human to take control.

397
00:32:29,404 --> 00:32:36,401
When we got to a point where like you know, we no longer had human drivers in the car, we
had gone past all these sorts of reliability issues.

398
00:32:36,401 --> 00:32:40,204
Yeah, we we were in a state where like you know the back end was super reliable at that
point of time.

399
00:32:40,204 --> 00:32:48,969
I'm sorta curious what the autonomous car companies are are sort of doing in these regards
because it is sort of this scenario where you have to assume that you lose connection for

400
00:32:48,969 --> 00:32:49,985
for external control.

401
00:32:49,985 --> 00:32:52,518
So the vehicle itself has to be completely automated.

402
00:32:52,518 --> 00:32:57,698
Is there like a s like it's basically like a safe stop uh mechanism or these

403
00:32:57,698 --> 00:32:59,809
Cars are like pretty intelligent, right?

404
00:32:59,809 --> 00:33:04,262
Like it has all the software to run autonomously installed on the car.

405
00:33:04,262 --> 00:33:10,015
It's not like it d it doesn't operate in a in a world where like you it's getting all the
instructions from the back end.

406
00:33:10,015 --> 00:33:15,548
It's not like hey, you know, it's getting instructions to take a turn or like you know, go
on a certain road from the back end.

407
00:33:15,548 --> 00:33:24,743
The amount of information that comes from the back end is kind of actually limited,
outside of like, you know, just general like navigation or like the the pickup and drop

408
00:33:24,743 --> 00:33:26,784
off locations and you know.

409
00:33:26,784 --> 00:33:27,544
route locations.

410
00:33:27,544 --> 00:33:34,188
Like there were cases where like you know, it still had to be connected to the back end,
even though it wasn't needing any instructions.

411
00:33:34,188 --> 00:33:38,260
Like there were still remote operators who had to, you know, look at where the cars were.

412
00:33:38,260 --> 00:33:43,200
There there's a really high level of availability needed just just so that remote
operators could keep track of

413
00:33:43,200 --> 00:33:44,660
I mean it makes sense, right?

414
00:33:44,660 --> 00:33:46,011
if you're gonna deal with this scenario.

415
00:33:46,011 --> 00:33:53,799
And I I think there's like obviously in some other spaces, like when I was working in
aerospace, you have to deal with the fact that even if everything is completely reliable,

416
00:33:53,799 --> 00:33:55,521
it's not instantaneous.

417
00:33:55,521 --> 00:34:01,130
And so you still have to deal with a lot of those scenarios in in the in that regard where
there's incredible latency.

418
00:34:01,130 --> 00:34:09,976
I mean it may just be hundreds of milliseconds, but there's no way you can get that down
because you have to do real processing and you just really have to push that out.

419
00:34:09,976 --> 00:34:12,248
to wherever the vehicle is at that point.

420
00:34:12,248 --> 00:34:20,778
By that very nature, you don't need as high reliability system internally to deal with
that because it's not doing as critical actions in that regard.

421
00:34:20,778 --> 00:34:28,110
Is there an extra complexity here that you really were trying to solve specifically at
Cruz uh regarding this particular scenario?

422
00:34:28,110 --> 00:34:32,052
infrastructure side standpoint there is like a lot more to that, right?

423
00:34:32,173 --> 00:34:42,220
you're not just creating uh isolated resources like uh sure a service you know needs a
Kubernetes cluster, it needs a namespace, it needs a database, it you know needs a CST

424
00:34:42,220 --> 00:34:42,620
pipeline.

425
00:34:42,620 --> 00:34:49,245
Sure, from an infrastructure standpoint, you're not just creating these resources and
handing it over to the service, developed service team, right?

426
00:34:49,245 --> 00:34:51,186
You are stitching it all together.

427
00:34:51,186 --> 00:34:56,009
You're making sure that you know they all work in a coordinated fashion and they're
integrated together.

428
00:34:56,009 --> 00:34:57,170
And you know

429
00:34:57,272 --> 00:35:02,007
it works as a cohesive uh piece of software when you give it to the service developers,
right?

430
00:35:02,007 --> 00:35:07,272
Like imagine that let's say I someone asks me for like uh a jigsaw puzzle, right?

431
00:35:07,272 --> 00:35:12,156
I could uh put all the pieces of the jigsaw puzzles in the bag and you know just hand it
over to them.

432
00:35:12,217 --> 00:35:21,185
Or I could solve the puzzle for them and I could lay it out on my table and join the
pieces together and create the final you know puzzle and give them the beautiful picture.

433
00:35:21,185 --> 00:35:22,270
That's what we did.

434
00:35:22,270 --> 00:35:28,784
I hate the analogy because, you know, from that standpoint, it makes it sound like you're
doing all the fun work so no one else gets to do it.

435
00:35:28,784 --> 00:35:33,256
I'm just like, would you like I'm gonna give you the present of a jigsaw puzzle and you
got two options.

436
00:35:33,256 --> 00:35:40,680
Either one, it starts all together and so you need to decide, do I rip it apart first
before I get to enjoy using it?

437
00:35:40,680 --> 00:35:44,052
Or do I just leave it like that?

438
00:35:44,563 --> 00:35:45,814
And get it done already.

439
00:35:45,814 --> 00:35:46,695
Yeah.

440
00:35:46,695 --> 00:35:50,419
And from my standpoint, I'm just like, why not let other people do the uh the enjoyment
work too?

441
00:35:50,419 --> 00:35:54,674
Maybe application developers want to enjoy um writing some infrastructure code.

442
00:35:54,966 --> 00:35:57,240
So at that point of time you are reinventing the wheel, right?

443
00:35:57,240 --> 00:36:02,379
So if a common team takes a problem, they first off have the expertise to be able to solve
that, right?

444
00:36:02,379 --> 00:36:03,681
So they can do it faster.

445
00:36:03,681 --> 00:36:06,582
They're solving it commonly for all the different teams together.

446
00:36:06,582 --> 00:36:07,552
No, I I totally agree.

447
00:36:07,552 --> 00:36:11,833
I just like from that standpoint, this is an interesting analogy to make.

448
00:36:11,833 --> 00:36:12,984
Because I mean I sort of agree.

449
00:36:12,984 --> 00:36:24,107
Like I think there are f like there's very specifically, there's a lot of people that go
into the area of an organization that's more focused on building up like say developer

450
00:36:24,107 --> 00:36:30,109
tooling, um, SDKs, et cetera, or thinking about reliability from an infrastructure
standpoint.

451
00:36:30,109 --> 00:36:35,302
And I do think that they think about it being a jigsaw puzzle that they're putting
together and

452
00:36:35,382 --> 00:36:43,391
You know, when I hear that, I I really do wonder, you know, uh lots of different
organizations are saying, especially SaaS companies, they love saying this, we do this

453
00:36:43,391 --> 00:36:46,254
work so that the developers don't have to.

454
00:36:46,315 --> 00:36:55,195
And at the end of the day, I I've heard hundreds of companies say this, or thousands at
this point, and I have to wonder, what are developers doing if everyone else is doing

455
00:36:55,195 --> 00:36:56,305
their work?

456
00:36:58,105 --> 00:37:00,382
So you go higher up the stack, right?

457
00:37:00,382 --> 00:37:05,314
Like you work on more of the business problems and supposedly.

458
00:37:05,314 --> 00:37:16,457
I feel like I need to invite on a first class application principal engineer and have them
tell me what sort of work they're doing still today in an organization where everyone else

459
00:37:16,457 --> 00:37:25,130
says that they solve, you know, XYZ problems for them because I'd be really curious where
they don't end up as just basically being a product manager and doing the business

460
00:37:25,130 --> 00:37:25,570
problems.

461
00:37:25,570 --> 00:37:34,188
And if we and I think I'm gonna spoil this episode by saying AI for the first time, if
that if the developer isn't doing that much left.

462
00:37:34,188 --> 00:37:41,750
then really it can be for sure saw like if there's if if it's a null set that an engineer
is doing, you can of course replace them with an with an L O.

463
00:37:41,750 --> 00:37:45,457
Uh maybe I'll say let's switch over to picks for the episode.

464
00:37:45,457 --> 00:37:47,670
So Pushker, what what did you bring for the audience today?

465
00:37:47,670 --> 00:37:48,662
Yeah, yeah.

466
00:37:48,662 --> 00:37:59,525
So I'm into like flying drones and um so I got this uh DJI Mini 3 drone you know a couple
of uh a year ago, a little over a year ago.

467
00:37:59,525 --> 00:38:05,186
And um so holding it up, like this is my DJI Mini drone and it's amazing.

468
00:38:05,186 --> 00:38:07,017
I just love flying the drone.

469
00:38:07,017 --> 00:38:14,969
It has like an amazing uh camera and it has like a really great battery life as well and
it has like an amazing range.

470
00:38:14,969 --> 00:38:16,189
so let me hold it up.

471
00:38:16,189 --> 00:38:18,550
So this this is my pick.

472
00:38:18,550 --> 00:38:20,931
I really enjoy like, you know, flying this.

473
00:38:20,931 --> 00:38:31,324
Um so th there are like, you know, spaces where uh outside of the city you have a permit
to fly your drone and you know and of course like I so I stay in Seattle.

474
00:38:31,324 --> 00:38:37,366
around Seattle there are a lot of like you know, mountains and waterfalls and like the
scenery is like you know amazing outside of the city.

475
00:38:37,366 --> 00:38:42,567
I don't really have to drive outside of the city if I can just get my drone in the air and
you know, look around.

476
00:38:42,567 --> 00:38:44,696
So it's something that I, you know

477
00:38:44,696 --> 00:38:45,721
really enjoy.

478
00:38:45,721 --> 00:38:48,657
Yeah, just just having my tone in there and like, you know, capturing things around.

479
00:38:48,657 --> 00:38:49,302
Yeah.

480
00:38:49,302 --> 00:38:58,538
I think there was a paper that Amazon released uh almost a decade ago about this magic
little area of height where they were thinking about flying deliveries of drones onto

481
00:38:58,538 --> 00:39:02,611
people's balconies because it was like right below the commercial aircraft route.

482
00:39:02,611 --> 00:39:06,573
It had to be reported and above the height where it had been regulated.

483
00:39:06,573 --> 00:39:09,455
And so they could fly them from building to building and I I thought it was ingenious.

484
00:39:09,455 --> 00:39:13,762
I I haven't maybe if someone knows about that project, I I'd love to hear more and

485
00:39:13,762 --> 00:39:17,574
Here's a here's a free invite to come on as a guest and talk about that.

486
00:39:17,574 --> 00:39:18,634
I like the pick.

487
00:39:18,634 --> 00:39:19,144
it's interesting.

488
00:39:19,144 --> 00:39:22,706
Honestly, I haven't bought a drone myself, but if I if I go now, here's one.

489
00:39:22,706 --> 00:39:27,815
I think my requirement for a drone has to be like my keyboard, as silent as possible.

490
00:39:27,815 --> 00:39:30,899
I you know, I I racked my brain this week to to come up with one.

491
00:39:30,899 --> 00:39:33,721
Uh and then I realized that I hadn't shared this.

492
00:39:33,721 --> 00:39:38,157
Uh maybe I did and I just lost the link, but it's the Murderbot Diaries.

493
00:39:38,157 --> 00:39:41,274
Uh it's a series of books by Martha Wells.

494
00:39:41,274 --> 00:39:42,134
It's about

495
00:39:42,134 --> 00:39:49,158
An autistic android who tirelessly works to prevent their humans from getting killed due
to their own stupidity.

496
00:39:49,158 --> 00:39:55,532
Maybe you're a introvert engineer and you try to stay in your remote cubicle at work and
never interact with another human being.

497
00:39:55,532 --> 00:40:02,896
I I feel like like there's something that resonates a lot with with me personally, you
know, and I think other other introverts, other software engineers will will get this too,

498
00:40:02,896 --> 00:40:06,808
that you just sometimes want to be left alone and just do your thing.

499
00:40:06,808 --> 00:40:10,594
And you know, it's those pesky other humans that are uh

500
00:40:10,594 --> 00:40:15,087
Doing going out and making life more difficult for themselves.

501
00:40:15,087 --> 00:40:16,843
Sometimes, yes, yes for sure.

502
00:40:16,843 --> 00:40:18,658
I see I see what you're saying.

503
00:40:19,582 --> 00:40:23,064
In this c in your case, maybe it's the the application engineers, right?

504
00:40:23,064 --> 00:40:29,567
You know, it's like I have to read the application engineers from uh shooting themselves
in the foot with their own technology.

505
00:40:29,567 --> 00:40:33,059
And uh so I I think there's something very reminiscent out there.

506
00:40:33,059 --> 00:40:35,359
Uh the books, absolutely great.

507
00:40:35,359 --> 00:40:38,472
They're all most of them are very short, so it's very easy to get through all of it.

508
00:40:38,472 --> 00:40:42,813
And then there is uh one season of a show with I think absolutely fantastic.

509
00:40:42,854 --> 00:40:43,567
highly recommended.

510
00:40:43,567 --> 00:40:44,300
Awesome.

511
00:40:44,300 --> 00:40:48,288
Yeah, I'll I'll try to yeah, read the book and watch the show as well.

512
00:40:48,288 --> 00:40:48,808
oh

513
00:40:48,808 --> 00:40:53,360
it's no obligation, but I do feel like uh you know, maybe I should start making it a
requirement.

514
00:40:54,761 --> 00:40:59,712
But you know, maybe I should make it a requirement that all my guests have to try out
whatever I recommend.

515
00:40:59,712 --> 00:41:08,426
Uh and I I I promise to do the same, although I I now I fear with the number of episodes I
do that I would be under a huge obligation to try a lot of things.

516
00:41:08,426 --> 00:41:17,560
Uh like someone made like someone's like, yeah, I love my Peloton bike and I'm like I
don't have space for some more exercise equipment So I I don't know if I'm gonna make that

517
00:41:17,560 --> 00:41:18,330
work.

518
00:41:18,570 --> 00:41:22,866
Uh or just like some really expensive stuff that I'm just like I'm I'm not into mountain
biking.

519
00:41:22,866 --> 00:41:24,172
And if I was, I'm not buying that.

520
00:41:24,172 --> 00:41:25,103
Yeah.

521
00:41:25,265 --> 00:41:27,400
I have a bike similar to a Peloton bike.

522
00:41:27,400 --> 00:41:29,265
Now it's just a cloth hanger for me.

523
00:41:29,265 --> 00:41:31,170
I no longer use it.

524
00:41:31,170 --> 00:41:31,822
It's just

525
00:41:31,822 --> 00:41:32,162
Yeah.

526
00:41:32,162 --> 00:41:37,888
Well, Pushker, thank you for coming on and trying to justify uh cross plane to to to f for
for the world.

527
00:41:37,888 --> 00:41:41,541
I guess this is now the the fishing first time someone's come on and discussed it.

528
00:41:41,541 --> 00:41:48,697
We've had episodes on on Terraform and uh Open Tofu and the direction those have going, so
you can do a compare and contrast if you want.

529
00:41:48,697 --> 00:41:54,743
Uh or maybe this has can convinced you to try out a new technology, especially if you have
something homegrown.

530
00:41:54,743 --> 00:41:58,050
Uh I want to thank you, Pushker, for coming and joining us for this episode.

531
00:41:58,050 --> 00:41:58,662
Yeah, thank you.

532
00:41:58,662 --> 00:41:59,022
Thank you.

533
00:41:59,022 --> 00:42:05,890
It was yeah, really nice talking to you and thank you for giving me the opportunity to,
you know, express my opinions on my podcast.

534
00:42:05,890 --> 00:42:06,531
Yeah, of course.

535
00:42:06,531 --> 00:42:13,990
Hopefully the the audience feels the same and thanks to all of them for tuning in to this
week's episode and I hope to see everyone back again next.

