speaker-0 (00:07.95) Welcome back to Adventures in DevOps. It's already been three years since the community discarded HashiCorp for their devious practices, and engineering orgs everywhere have been on the lookout for TF alternatives. I know for us at Authorist, CloudFormation will always hold a special place in my company's hearts when it should have probably been let go a long time ago. But we've recently taken a cleared preference towards OpenTOFU, I think, for what it's worth, the spiritual successor of Terraform. But I wanted to bring in an alternative perspective and a guest who believes in a technology that I personally refuse to try. Staff software engineer at Snap, previously Cruz, and AWS, all things cloud infrastructure, Pushkar Gopala Krishna. Welcome to the show. At this point, I can only imagine that the last customer of Terraform Somewhere is finally migrating to OpenTOFU at this point. Or maybe at least has chosen a more nuanced approach. speaker-1 (00:50.318) Thank you. Thank you. Thanks for having me on the show. speaker-0 (01:03.798) So obviously you want to discuss some of your your experiences and why you may not have even chosen Terraform in the first place. speaker-1 (01:11.678) We just let service developers, you know, define whatever infrastructure they wanted. It worked in well really well initially, but over time I guess developers started seeing you know issues with Terraform. first off, like we have to use like a custom language like HashiCop HCL which is which is a bit of a weird configuration language and outside of Terraform. speaker-0 (01:34.498) Well, I wanna ask about that. Why is that weird for you? speaker-1 (01:36.526) So for with a configuration language like YAML, right, it's much more widely used, especially if you're in Kubernetes, you know, you have to use it much more a lot. but HCL is only used for Terraform. Service developers generally don't, you know, work with infrastructure configurations day in and day out. They use it, you know, once to set up some infrastructure and then they're done with that for the next, you know, few weeks, months and they don't get back to it. So Because of that, it's like you know, people don't take the time to learn it well and that contributes to, you know, why people think it's a little weird. speaker-0 (02:11.938) So is it from your perspective there there is a huge burden to getting it right because there's a lot of complexity in the or maybe there's a huge gap in the historical knowledge and jumping over that gap to get to a a place of mastery is challenging. Yeah. Yeah. From my limited experience, it it actually hasn't seemed like there's a lot of complexity in it. So I'm s maybe I just don't see it and I I just or I'm running into all those pit h pitfalls, you know, all the time without even knowing about it. Like, what do people end up Struggling with the most. speaker-1 (02:43.298) With control structures and looping, right? So let's let's say you know you want a GCS bucket, you go copy paste the configuration for that. Now let's say you want five GCS buckets, right? There is like looping in the HashiCop language, but people don't know that syntax very well. So what they do is they just copy paste the same bucket configuration five different times. Now let's say you know you have to apply some sort of a best practice to it later on, right? Let's say you have to update your lifecycle policy for that GCS bucket. Now have to do it like five different times. I've seen a lot of cases where people just end up copy-pasting stuff from the internet or like stuff from you know other repositories. It's duplicated all over. And now making a change is hard. Now being in the infrastructure space, like we wanted to enforce like a lot of best practices, lifecycle policies with GCS buckets, for instance, or with say spanner databases, we had to say set up like autoscaling. or like backups, point in time recovery, like the right permission models. service developers wouldn't normally set these things up. Being in the infrastructure team if we had to do it, it was super painful. We either had to, you know, request service developers to, you know, do it themselves or we had to take it upon ourselves and, you know, go manually change a lot of repositories and that was super painful. speaker-0 (03:59.51) I think that is a really good point because realistically, when you have a bunch of mistakes that you make, and everyone will make mistakes, it's impossible to get things right the first time with application code, for instance, you just refactor it. And I think while maybe HDL is easy to write or easy to copy, when you make a mistake there, the the penalty for getting it wrong or wanting to change things in the in the future becomes a huge challenge to actually make that happen. Because you can't just merge together five GCS buckets or AWS buckets or whatever have you into the same code block with a for each loop or a count loop or whatever you want, you really need to think about how to merge that code because it's gonna impact your actual production infrastructure. Exact I mean guess from my perspective, I would say the code itself is easy to learn. The the penalties for getting it wrong are very high and it's very easy to fall into those. speaker-1 (04:43.064) Exactly. speaker-1 (04:51.062) HCL is something that like I said, you know, people write it once, don't visit it, you know, for another time for a few more weeks or months. So there's no incentive to learn the details. But with Kubernetes, you have to use like YAML a lot. Our team, especially, we had a lot of you know Kubernetes experts, but there was like a more inclination towards YAML when compared to HCL. speaker-0 (05:11.906) I think the first thing is that you sort of mentioned here is that exposure to the underlying infrastructure technology or d DSL domain specific language that is being used here. And you feel like the application engineers, whoever was building a a new microservice, et cetera, wouldn't have a lot of experience with HCL. And while it was easy to copy and paste somewhere else, there's there was a challenge of getting it right. At least in my experiences, there has like I don't remember an application engineer that ever touched anything in YAML. whatsoever, except maybe if they had to jump into their CI C D platform. Like if you're using GitHub or GitLab or something else, this is almost always in YAML. And the interesting thing is that the YAML that GitHub uses is not even standard YAML. What are application developers doing that would be in YAML that would justify a different experience for them that would be beneficial over say HCL? speaker-1 (06:04.682) At cruise, the slight difference was that we had like a huge team of support resistant reliability engineers who were super familiar with YAML. So the support with YAML was much more than the support that application developers received with HCL. you're right, like you know, application developers they in general wouldn't use YAML a lot more than they would use HCL. speaker-0 (06:26.936) An experiment that you could make in practice could have been switching all of the areas where you're writing YAML into writing a DSL that is, I mean, you basically using HCL, but not for Terraform, but for your actual configuration. Why not consider that as an option? speaker-1 (06:42.074) it was just that yeah, developers across the company had kind of had a bad experience with HCL and it had just less left a bad taste in our mouth and you know we just didn't want to, you know, deal with it anymore. There was this, you know, initiative where we had to add cost tags to, you know, our infrastructure and we had you know system reliability engineers who had to take this initiative. It took a few, you know, weeks to months and During this whole time, like they initially started off with politely asking service developers, you know, to do it themselves. But as you know, you know, like service developers are usually tied up with, you know, their own projects and initiatives. Updating HCL code is one of their least favorite things and didn't happen for a lot of teams. These SREs were forced to do it, you know, on their own. They tried automating it and, you know, they tried writing scripts to do that, but it did result in in a few cases where like they updated the HCL configuration wrongly and it led their workspaces breaking and it it was it was a disaster. So the Terraform applies stopped running and there were some critical services that this happened to and essentially all their like deployments were blocked. speaker-0 (07:52.598) I mean, I I totally see this showing up in it's not just the infrastructure world. I I do see a lot of organizations that don't fully embody the the DevOps mindset and try to have like there's like a security team or a QA team or a platform engineering team or a marketing team, sales, whatever, they have a requirement to change the infrastructure, the the software that basically the organization owns, but the only people that are sort of empowered to make the change are aren't part of the decision making process. So you have one team that's basically put out on a ledge, the maybe it's the at site reliability engineers, who are told we need to have those cost allocation tags everywhere so that we know how much the infrastructure is costing for every or also the application is actually running so that we can, I don't know what the end game is here, basically fire some engineers for making really expensive technology that doesn't actually impact the bottom line at all, potentially. I I don't know. Or maybe chargebacks to particular customers through some sort of aggregate usage depending on what product you're running. but the end of the story is that that initiative is pushed through a part of the organization that doesn't have the capability or responsibility or accountability for impactful changes. And in order for that to happen, they try to plead very carefully with those in charge. And of course other teams Who don't aren't part of the same initiatives have no incentive to make those changes. And the teams that obviously, some of them will acquiesce and say, Yeah, sure, fine, we'll do it. And do something that they're not very comfortable with doing or doesn't don't do very frequently. And I think this is just any technology honestly can run foul in this area where or any business making choice will happen here. And so, like what I really hear fundamentally is it isn't so much a problem on the application side. where you're building stuff, but really who is fundamentally responsible for making code changes for or whatever you want to call them, configuration changes, who is fundamentally accountable for that? And aligning the ownership with the accountability with whoever is going to perform the action. If you don't have that, no matter what technology choice you've chosen, you could easily fall in into a pit of failure. if I could get your perspective here on exactly how cross plane works for anyone who speaker-1 (10:05.048) That that makes sense. speaker-0 (10:12.374) Also has never used it fundamentally. speaker-1 (10:14.83) Crossplane a Kubernetes native tool. Essentially, you can represent your infrastructure as Kubernetes custom resources. Service developers typically manage their configurations as YAML in a Git repository, packaged up as like a Helm chart. When you run kubectl apply, any infrastructure configuration that you can have on the actual infrastructure in, say, GCP or Azure, that is now represented as a custom resource. When you make a change to your custom resource and when you apply it, the Kubernetes controller is continuously watching these custom resources and when it detects a change, so it essentially compares the state that is in the custom resource and it compares it to the state of the external resource. When it sees a difference, it tries to reconcile it and essentially get the state of the external resource to match the state of the custom resource. speaker-0 (11:07.042) the infrastructure changes is asynchronous from the deployment strategy there. It sounds like is you make the change, it gets deployed to the Kubernetes cluster, which picks it up and then goes and actually applies that change to the running infrastructure that you have. So if you add a new load balancer to your configuration, then when it gets when the the YAML file gets s consumed by the cluster, you will get that load balancer spun up in Azure, AWS, et cetera. And how do companies deal with the fact that fundamentally there's no that process is asynchronous from the actual committed change? Like normally I would expect when you commit an infrastructure change, it's running in your CI C D platform and that's being the blocking on the fact that those changes are getting made. How do you do that in cross plane? speaker-1 (11:55.886) When a custom resource is changed, it is immediately reconciled as well. So when you run kubectl apply, it does synchronously try to make the change. And only once the change is done, the status of the custom resource or the claim or the composite resource is updated with the status of the external resource. speaker-0 (12:17.464) So how do you manage something like projects in in G C P or the directories in Azure or the accounts in AWS? If you don't have a Kubernetes cluster, like how do you spin up a new AWS account? How do you, you know, get the new project created or get the right IAM profiles integrated? Like how do you manage that part? Because obviously there's no cluster running at this point to even execute that in the first place. So speaker-1 (12:41.55) So so at Cruz the way it worked was we started off with Terraform and we built a bunch of infra Kubernetes clusters. Those are still managed with Terraform. It's just the base infrastructure clusters. And so we have another tool that spins up G C P projects. So that That part is actually outside of Crossplane. And there are historical reasons for it. So we had this tool that was built a long time ago called Juno, and that was responsible for spinning up a service and like all the associated infrastructure, you know, with the service. And as a part of it would spin up the GCP project. And it also managed the RBAC or the permissions piece of it. Permissions were all maintained outside of Crossplane. Only the so we use Crossplane only for spinning up like I said previously. speaker-0 (13:27.492) I mean maybe let's dive into that. So there's like, you know, some golden path for setting up the account and then you switch over to a different technology for the next layer and then switch over to a different technology for the layer after that. When you get down to the application layer though, how do you handle the bootstrapping process for even knowing about the YAML file for the application to run? Because historically in a lot of organizations, if you are building microservices, every single microservice is in a separate repository and you include the infrastructure changes that are required for that microservice with the application code. And so if you're in a repository, it's got your hypothetical YAML file or HCL for deploying that application. And so when you want to add a new database, you add it directly into your repository in maybe a templates directory or a deployment directory, et cetera. But with that, how would Crossplane even know or the Kubernetes cluster you're running realistically, know to go look in this new GitHub or GitLab repository to go pick up the code to even deploy in the first place? Yeah. speaker-1 (14:28.268) Yeah. So when developers created a service with Juno, it would spin up like you know all the necessary infrastructure. So you need the compute infrastructure, you need the observability infrastructure, you need like secrets, you need like the networking, the service mesh. There's there's a lot of Juno was responsible for that. It would spin up all these different components and additionally it would create a Git repository and it would also like create the CICD tooling with it. speaker-0 (14:45.035) Yeah, yeah. It's everything. speaker-1 (14:56.364) So when a new service was created, it would essentially take this golden path code and write it to the newly created Git repository for the service. speaker-0 (15:03.992) Juno sounds more like less like a tool and more potentially like a a platform for creating applications. Is that yeah, yeah. Okay. Okay. So basically, you know, it sounds like if anything, you don't love Crossplane. You love Juno. Juno seems like the best thing ever. It it's one of these platform engineering platform platforms that you know creates platforms for application deployment. And it seems like it's doing everything that you need. It's doing the account management or project management, resource creation. It's doing the orchestration for, you know, kicking off Crossplane for generating repositories for application code and everything. At the end of the day, all the only part there that's still running in Crossplane is the maybe specific application level infrastructure. Like at that point, why even why even use Crossplane then? Like why not just like also have Juno, you know, kick off some sort of container running in in the cluster to go execute. Why wh why not customize that too? speaker-1 (15:59.598) Juno was was, you know, a beast of its own. It it had a lot of pros and cons and it was built a long time ago. And developers, you know, some developers loved it, some developers hated it. And usually have that right, like with any, you know, technology or platform that's been around for a few years, like say seven, eight years, people have strong opinions about that. It was getting to a point where like, you know, it was getting hard to scale up Juno to manage, you know, all these in different infrastructure components because it was a lot. There were a lot of services, there were lot of environments within these services and it it was a lot for Juno to manage. I mean, you can make it work given that you invest enough in a tool, you can still make it work. But we had reached a point where you know, we felt that it was easier to manage infrastructure with cross plane than have, you know, Juno do that. especially with the legacy code that it had. speaker-0 (16:53.708) I mean, it makes a lot of sense if you come from the other perspective. Like rather than from the ground up of you need to deploy infrastructure for your organization and that and included in that is everything. And if you come from the the place where you have a custom tool that's doing your infrastructure deployments and managing repositories and and whatnot, and from there you're like, How can we make our custom tool, which has a terrible internal development process lifecycle where it's one managed by two people on different teams that have no obligation or OKR or KPI behind actually improving the tool to do what it needs. What if we just rip out all of that stuff and replace it with an open source project? Then I can totally see like, okay, what's available out there? And that list is very short for lists of technology to actually delegate infrastructure deployment to. speaker-1 (17:39.288) Course, there are you know cases where cross plane cross plane is not the solution for everything, right? Like it it's it's still like when compared to Terraform, it's still maturing. When we initially tried out cross plane, of course we you know prototyped it on a small set of resources and resources and we thought that it was great and you know we started adopting it. later we realized that you know there there are issues with cross plane as well. So Terraform has been around for like more than 10 years, right? Like it has matured a lot and you know there are a lot of providers, like any cloud resource or SaaS product that you can think of usually has a Terraform provider and you should you know should be able to manage it with Terraform. Crossplane is not there yet. So we couldn't really migrate completely to Crossplane because of that. speaker-0 (18:25.078) Not every provider out there has a controller available in in Crossplane. So that must have caused some problems for you. speaker-1 (18:32.876) We kind of realized that a little bit later that, you know, they tend to have providers. So so we were stuck in in a world where like, you know, we had to maintain Terraform as well, because you know, these resources couldn't be migrated to cross plane. Also, like I think our expectation was that we we didn't want to move completely off of Terraform, right? Like that that wasn't the goal to start off with. We only wanted to manage like say eighty percent of the infrastructure if possible. We wanted to move over like the the hardest eighty percent of the or like the most commonly used eighty percent of the infrastructure to cross plane. We kind of expected that there would be like a long tail of you know components that would remain in terraform. That would at least like reduce our terraform cost. speaker-0 (19:16.056) So one of the challenges that usually comes up in any infrastructure as code solution is how you manage additional environments. Like I I know there's always a question of like, okay, we need to spin up a temporary environment for every single pro pull request that gets generated. We usually need maybe a temporary one that lives a little bit longer than a branch for extensive load testing and then will get automatically terminated when it's done. And then of course, in every engineer's personal or I mean company provided but dedicated AWS account or cloud account that they're using, they need may need to deploy the exact same infrastructure in in some way to to run the individual services. And maybe there's some sandboxes as well if they're doing some sort of local development. speaker-1 (19:55.64) So Juno was our savior at that point of time. so essentially Juno had this notion of environments and we could have, you know, three different environments. So we built in the mechanism to create sub environments within a given environment. So actually even called it stacks. It wasn't super ideal, but we we were able to get the level of isolation that we wanted. speaker-0 (20:17.602) The challenges that always comes up within the environment management is realistically domain separation, where like you have your production domain, which is like some company dot dot com or whatever, and then your non production domains, which are like test dot production dot com. And it's like, well, that lives in the production environment in order to resolve it. I don't see any way that could have ever been completely isolated. And I think you were sort of hinting at that. speaker-1 (20:43.118) Yeah, so we did our best to provide the level of isolation that we needed, but there were cases where say permissions for instance were definitely shared across, you know, different environments and that actually made sense. We didn't want, you know, separate permissions in broad versus development and so there like, you know, there was some commonality. at the network level, yeah, we couldn't really get the level of isolation that we wanted. Between like Dev and staging in prod, we still had like separate VPCs and you know the services would run across separate VPCs, but within different stacks in an environment, they would end up running in the same VPC. speaker-0 (21:20.11) I I think this is the sort of thing that a lot of companies end up getting wrong though, where they believe they need this level of isolation there. And the my response is always like, Do you think you're using a test version of GCP or AWS for your non-production environment? No. You are using the the one that's on their public page and you're creating a real you cloud account for that and using real stuff there and paying real money for your resources, there's no reason why that shouldn't be at every other part of the of the stack as well. That's when I see some, you know, lesser experienced engineers like mind shatter, like, yeah, I guess that's true, right? Like we don't have different stuff there. I think this happens a lot in clue Kubernetes. I I think one of the one of the core challenges is a lot of companies fall into the pit where There's a few low number managed Kubernetes clusters for the entire company or organization level, and everyone deploys to the same cluster. And I I think while it's doable for sure, you do end up with a lot of ownership issues and you end up with having to tag a lot of things and then running amok of the IAC tools. And so we we know fundamentally first class isolation at low down as you can is is beneficial, but you shouldn't push it any lower that than you absolutely need to. And I think using infrastructure level services is for one for sure, like things like you mentioned. I think login and access control is another one. Like people being able to log in into the test environment to validate things, it's still useful to validate your production login and access control strategy and not use a fake version of it. Because you know those things aren't ready for production. I think it was a couple of companies ago that I was advising where they kept running into this problem where every time one team wanted to do something, They were utilizing APIs from a second team. So like there was the mobile and app development team and the website UI team. And they, for their non-production version of their UIs, were calling non-prod versions of the underlying services, which had tons of things broken in them. It had tons of features that hadn't been released yet. And you're not coding against like production realistically. And so it was a huge challenge in those scenarios. And because of that, they had a lot of hacks in place. speaker-0 (23:35.138) Where whenever you logged into a non production environment, you would also log into production at the same time and have two tokens that would get passed around. And then the services would have to like figure out in the load balancer like which token to send at the right time. and then there was like VPNs just for non production environments because they wanted to use a test version of their login and access control strategy. Genius really when you get when you get down to the bottom of it. So I I think I think I'm with you. Like Figuring out where the where the isolation should be and at what level makes a lot of sense. If you're running a single cluster or a set of them, isolating at the VPC level doesn't make it doesn't have a lot of benefit there. It doesn't it doesn't align with the infrastructure choices you have. but I think that's a that's a trade off. speaker-1 (24:17.87) I've seen those patterns as well where like in the staging environment, you know, develop like there are services that end up calling production and so they've seen cases where like, you know, integration tests have ended up writing to actual production databases because of that and speaker-0 (24:31.544) That's perfect. That's what you want, right? Because the integration test should have the right permissions in the production environment and those permissions should be none, right? And so when it's right, when it tries to write, you just validate the right, you know, does a person ha or the entity that's calling have access? I mean, we do the exact same thing. all of our services call the production version of the other of their dependencies during development. Like developers running a application client on their machine will call the prod version because it's the only version that you can trust. And they go through our access control strategy. And if and I think there's this mistake of like you isolate to protect your environment. You don't isolate to protect. You isolate to s single out for testing purposes, right? Because you want to figure out where the problem is. You don't I I mean unless you're doing a load test, then Then there's a question of what happens there. But I love the load test that failed because the non-production cluster doesn't have the same resources as the production cluster. And so of course it fails there. And so what are you really testing? So I I mean there I don't know how you do that effectively without cloning your production environment and spinning up exactly the same thing and then running the test there. speaker-1 (25:40.184) That's a hard problem to solve, right? Like there are a lot of options that you can try out. You can essentially replicate your entire production stack and have like a complete replica of it running and you can run your load tests over there. But again at that point you're like wasting resources, right? If you have a complete copy of your production but you're not using it for live production traffic, you're incurring a lot of cost for no speaker-0 (26:00.63) It's just temporary. First of all, like if you're running if you just have an environment that's long lived besides prod, you're doing it wrong. Shut it down. I guess that would be the first one. the second one, like 'cause you're not you're not you shouldn't be using it con consistently like that. And if you are, then you must be getting the value out of it for sure. yeah. And I think really that's the core aspect. So they speaker-1 (26:15.883) Exactly. Yeah. There was a point of time where we didn't really use Argo CD. We're using this CI tool called build kite. it was a great CI tool, but we were using it for doing continuous deployments when it was you know not meant to do that. So which meant that we had to handle a lot of the moving pieces ourselves, like you know, authentication and you know, making sure you know this thing scaled up and it had like you know all the namespacing in it. so We spent a lot of time debugging pipeline issues and it was like super frustrating. You also had to configure like the authentication from your pipeline, like from the from the build guide, people would get that wrong and until we, you know, got fed up with that and we said, Hey, you know, like this is not working out and we actually ended up moving to Argo C D at that point of time. speaker-0 (27:07.31) It's all about reducing the on call support that the platform team needs to do. You know, the less on call time they need to handle stuff, the the happier everyone is. speaker-1 (27:17.806) For sure, for sure. There was support. there were like non-urgent support issues which our team got a lot and there were like these on call issues as well. the support issues were like super frustrating and we spent a lot of time initially we had a lot of that and we actually invested quite a lot of time, you know, figuring out like where exactly we were spending time when supporting these different teams and Argo CD was one of the things that we did to you know reduce our time speaker-0 (27:46.322) The straw that broke the camel's back when it came to build kite? Like, how did you actually decide that, like today in 2026, the the thing that you do, of course, is use OIDC, right? The build server itself generates its own JWTs automatically, injects them into the process. You use a standard couple of lines of code in your CICD YAML file, which automatically authenticates to your cloud provider of choice, and you can automatically get the right role and permissions and execute stuff. And obviously, even a few years ago, GitHub didn't support didn't support OIDC. GitLab didn't support OIDC. interestingly enough, I actually was the one who was reviewing the GitLab implementation before even GitHub supported it. It was amazing that these providers that honestly usually have full admin access to your entire AWS account, your entire GCP account, a project, etc. And they're being authenticated with basically plain text credentials that are saved in a environment variables inside a third party provider. Obviously no one likes what was happening at this point, but most companies don't do anything about it. speaker-1 (28:49.194) So we actually didn't move off of it completely. No no. So it was so build kit is a great CI tool, right? Like I mean I don't want to understate that. It's definitely a good build CI tool, but not a great C D tool. You know, CI is like the you know, you meet a change right from the point where your code leaves your laptop. That is where the CI part starts, right? It should take care of like building your you know image, it should take care of like you know uploading it to some sort of artifact registry and should take care of running any sorts of tests that you have, right? Like unit tests, component tests, integration tests, like load tests it you know that you're gonna deploy. Once it is ready, I feel like that's where the responsibility of CI stops and the responsibility of C D starts. speaker-0 (29:35.374) Isn't there a little bit of a paradox there? Because like how do you know that you're done building that artifact without actually deploying it to production in a way or in an isolated production like environment where it's getting real requests in? speaker-1 (29:49.57) The way you do that is through slow rollouts, right, or Canary, right? And there are some cases where you can't figure out issues with your code until you actually deploy it. Interestingly, although we introduced like Argo rollout for deploying service infrastructure in a progressive manner. speaker-0 (29:58.934) Yeah, for sure. Definitely. speaker-1 (30:08.694) There there were interesting issues that you know we faced where because of a lack of a slow rollout mechanism we've we've caused outages, you know, widespread outages across like a lot of different teams. Case where while while I was working on this stacks, you know, initiative. So at that point of time like we had the ability to slow roll our feature to certain projects only. What happened was We made a change and so there was a bug in the change. What that resulted in happening was like the internal state of the Kubernetes namespace was accidentally marked for deletion. So after we wiped out like three different production services as soon as our change was rolled out. And it was all of a sudden, right? Like these service teams, these are like critical services. so for for some context, like at at Cruz we had these, you know, driverless cars that operated on the road that, you know, called these services and it was like the level of reliability expected was really, really, really high. Yeah, we accidentally wiped out like the whole production service, for like, you know, three different services and god, that that was speaker-0 (31:17.518) If I go like make a internet search right now, is there gonna be like a particular date where like all the cruise cars were getting into accidents because of this incident? speaker-1 (31:25.286) I don't no. which of course no. I don't think I'll comment on that. Yeah. speaker-0 (31:34.158) The I mean, with such a high reliability requirement, at least for us, we look at both the ability to reduce the probability, but also the ability to reduce the impact because we see the risk as a as a product of these two. I mean, obviously it's just not a straight multiplication, but you can look at it similarly, right? as far as expected value goes. So if you're thinking about reliability, and obviously autonomous driving is super critical from a reliability standpoint. What was the sort of mentality around dealing with the reducing the impact? side or potentially dealing with it in a reactive nature, like how did the vehicles handle scenarios where they couldn't connect back to the production control plane? speaker-1 (32:14.04) Initially there were issues, but as you know, like we were in a R and D phase for a really long time. That that was when these sorts of issues were identified and so we did have humans like human drivers in the cars who would operate the cars. If this sort of thing happened, like there was a human to take control. When we got to a point where like you know, we no longer had human drivers in the car, we had gone past all these sorts of reliability issues. Yeah, we we were in a state where like you know the back end was super reliable at that point of time. speaker-0 (32:40.204) I'm sorta curious what the autonomous car companies are are sort of doing in these regards because it is sort of this scenario where you have to assume that you lose connection for for external control. So the vehicle itself has to be completely automated. Is there like a s like it's basically like a safe stop mechanism or these speaker-1 (32:57.698) Cars are like pretty intelligent, right? Like it has all the software to run autonomously installed on the car. It's not like it d it doesn't operate in a in a world where like you it's getting all the instructions from the back end. It's not like hey, you know, it's getting instructions to take a turn or like you know, go on a certain road from the back end. The amount of information that comes from the back end is kind of actually limited, outside of like, you know, just general like navigation or like the the pickup and drop off locations and you know. route locations. Like there were cases where like you know, it still had to be connected to the back end, even though it wasn't needing any instructions. Like there were still remote operators who had to, you know, look at where the cars were. There there's a really high level of availability needed just just so that remote operators could keep track of speaker-0 (33:43.2) I mean it makes sense, right? if you're gonna deal with this scenario. And I I think there's like obviously in some other spaces, like when I was working in aerospace, you have to deal with the fact that even if everything is completely reliable, it's not instantaneous. And so you still have to deal with a lot of those scenarios in in the in that regard where there's incredible latency. I mean it may just be hundreds of milliseconds, but there's no way you can get that down because you have to do real processing and you just really have to push that out. to wherever the vehicle is at that point. By that very nature, you don't need as high reliability system internally to deal with that because it's not doing as critical actions in that regard. Is there an extra complexity here that you really were trying to solve specifically at Cruz regarding this particular scenario? speaker-1 (34:28.11) infrastructure side standpoint there is like a lot more to that, right? you're not just creating isolated resources like sure a service you know needs a Kubernetes cluster, it needs a namespace, it needs a database, it you know needs a CST pipeline. Sure, from an infrastructure standpoint, you're not just creating these resources and handing it over to the service, developed service team, right? You are stitching it all together. You're making sure that you know they all work in a coordinated fashion and they're integrated together. And you know it works as a cohesive piece of software when you give it to the service developers, right? Like imagine that let's say I someone asks me for like a jigsaw puzzle, right? I could put all the pieces of the jigsaw puzzles in the bag and you know just hand it over to them. Or I could solve the puzzle for them and I could lay it out on my table and join the pieces together and create the final you know puzzle and give them the beautiful picture. That's what we did. speaker-0 (35:22.27) I hate the analogy because, you know, from that standpoint, it makes it sound like you're doing all the fun work so no one else gets to do it. I'm just like, would you like I'm gonna give you the present of a jigsaw puzzle and you got two options. Either one, it starts all together and so you need to decide, do I rip it apart first before I get to enjoy using it? Or do I just leave it like that? And get it done already. Yeah. And from my standpoint, I'm just like, why not let other people do the the enjoyment work too? Maybe application developers want to enjoy writing some infrastructure code. speaker-1 (35:54.966) So at that point of time you are reinventing the wheel, right? So if a common team takes a problem, they first off have the expertise to be able to solve that, right? So they can do it faster. They're solving it commonly for all the different teams together. speaker-0 (36:06.582) No, I I totally agree. I just like from that standpoint, this is an interesting analogy to make. Because I mean I sort of agree. Like I think there are f like there's very specifically, there's a lot of people that go into the area of an organization that's more focused on building up like say developer tooling, SDKs, et cetera, or thinking about reliability from an infrastructure standpoint. And I do think that they think about it being a jigsaw puzzle that they're putting together and You know, when I hear that, I I really do wonder, you know, lots of different organizations are saying, especially SaaS companies, they love saying this, we do this work so that the developers don't have to. And at the end of the day, I I've heard hundreds of companies say this, or thousands at this point, and I have to wonder, what are developers doing if everyone else is doing their work? speaker-1 (36:57.262) So you go higher up the stack, right? Like you work on more of the business problems and supposedly. speaker-0 (37:05.314) I feel like I need to invite on a first class application principal engineer and have them tell me what sort of work they're doing still today in an organization where everyone else says that they solve, you know, XYZ problems for them because I'd be really curious where they don't end up as just basically being a product manager and doing the business problems. And if we and I think I'm gonna spoil this episode by saying AI for the first time, if that if the developer isn't doing that much left. then really it can be for sure saw like if there's if if it's a null set that an engineer is doing, you can of course replace them with an with an L O. maybe I'll say let's switch over to picks for the episode. So Pushker, what what did you bring for the audience today? Yeah, yeah. speaker-1 (37:48.662) So I'm into like flying drones and so I got this DJI Mini 3 drone you know a couple of a year ago, a little over a year ago. And so holding it up, like this is my DJI Mini drone and it's amazing. I just love flying the drone. It has like an amazing camera and it has like a really great battery life as well and it has like an amazing range. so let me hold it up. So this this is my pick. I really enjoy like, you know, flying this. so th there are like, you know, spaces where outside of the city you have a permit to fly your drone and you know and of course like I so I stay in Seattle. around Seattle there are a lot of like you know, mountains and waterfalls and like the scenery is like you know amazing outside of the city. I don't really have to drive outside of the city if I can just get my drone in the air and you know, look around. So it's something that I, you know really enjoy. Yeah, just just having my tone in there and like, you know, capturing things around. Yeah. speaker-0 (38:49.302) I think there was a paper that Amazon released almost a decade ago about this magic little area of height where they were thinking about flying deliveries of drones onto people's balconies because it was like right below the commercial aircraft route. It had to be reported and above the height where it had been regulated. And so they could fly them from building to building and I I thought it was ingenious. I I haven't maybe if someone knows about that project, I I'd love to hear more and Here's a here's a free invite to come on as a guest and talk about that. I like the pick. it's interesting. Honestly, I haven't bought a drone myself, but if I if I go now, here's one. I think my requirement for a drone has to be like my keyboard, as silent as possible. I you know, I I racked my brain this week to to come up with one. and then I realized that I hadn't shared this. maybe I did and I just lost the link, but it's the Murderbot Diaries. it's a series of books by Martha Wells. It's about An autistic android who tirelessly works to prevent their humans from getting killed due to their own stupidity. Maybe you're a introvert engineer and you try to stay in your remote cubicle at work and never interact with another human being. I I feel like like there's something that resonates a lot with with me personally, you know, and I think other other introverts, other software engineers will will get this too, that you just sometimes want to be left alone and just do your thing. And you know, it's those pesky other humans that are Doing going out and making life more difficult for themselves. speaker-1 (40:15.087) Sometimes, yes, yes for sure. I see I see what you're saying. speaker-0 (40:19.022) In this c in your case, maybe it's the the application engineers, right? You know, it's like I have to read the application engineers from shooting themselves in the foot with their own technology. And so I I think there's something very reminiscent out there. the books, absolutely great. They're all most of them are very short, so it's very easy to get through all of it. And then there is one season of a show with I think absolutely fantastic. highly recommended. speaker-1 (40:43.567) Awesome. Yeah, I'll I'll try to yeah, read the book and watch the show as well. speaker-0 (40:48.638) No it's no obligation, but I do feel like you know, maybe I should start making it a requirement. But you know, maybe I should make it a requirement that all my guests have to try out whatever I recommend. and I I I promise to do the same, although I I now I fear with the number of episodes I do that I would be under a huge obligation to try a lot of things. like someone made like someone's like, yeah, I love my Peloton bike and I'm like I don't have space for some more exercise equipment So I I don't know if I'm gonna make that work. speaker-1 (40:52.878) Interesting. Yeah, no, definitely sounds speaker-0 (41:18.57) or just like some really expensive stuff that I'm just like I'm I'm not into mountain biking. And if I was, I'm not buying that. speaker-1 (41:24.172) Yeah. I have a bike similar to a Peloton bike. Now it's just a cloth hanger for me. I no longer use it. It's just speaker-0 (41:31.822) Yeah. Well, Pushker, thank you for coming on and trying to justify cross plane to to to f for for the world. I guess this is now the the fishing first time someone's come on and discussed it. We've had episodes on on Terraform and Open Tofu and the direction those have going, so you can do a compare and contrast if you want. or maybe this has can convinced you to try out a new technology, especially if you have something homegrown. I want to thank you, Pushker, for coming and joining us for this episode. speaker-1 (41:58.05) Yeah, thank you. Thank you. It was yeah, really nice talking to you and thank you for giving me the opportunity to, you know, express my opinions on my podcast. speaker-0 (42:05.89) Yeah, of course. Hopefully the the audience feels the same and thanks to all of them for tuning in to this week's episode and I hope to see everyone back again next.