speaker-0 (00:07.832) Welcome back to Adventures in DevOps. It certainly seems most of the tech industry is running around like headless chickens, and we're hoping to bring some sensibility to the current state by focusing on observability. For that, we brought in co-founder of Grafana Labs, Anthony Woods. Welcome to the show. Thanks, Walter. speaker-1 (00:23.734) It's it's a pleasure to be here and thank you for inviting speaker-0 (00:25.486) you f we've been trying to get you on for for months and obviously it was one thing and then another first Gryphonicon and then the recent source code exfiltration attacks, which we will definitely get to. speaker-1 (00:36.002) Great. Yes. Definitely happy to get into it. It's definitely busy. I mean obviously the observability industry's got a lot of work cut out for it. speaker-0 (00:44.034) What is like the most top of mind thing regarding trying to help customers today? Like I mean, obviously Otel is almost old at this point since twenty nineteen, but you know, what is old is new again. So but maybe there's something else particularly on your mind. Yeah. speaker-1 (00:57.176) Yeah, I mean, I think it's just a lot of customers that we're working with are just going through revamping their observability strategies. and there's a lot of reasons for that. you know, the two things that, you know, their engine engineering leadership, you know, is focused on is reducing costs, right, and improving productivity. Right. Like b basically that's that's the things that they care about most. and you know, what we've seen is observability is a big part of that. Observability, you know, for a lot of our customers, you know, is a very large line item on their bills. And so they want to work out ways where they can optimize those costs. So they're trying balance both deliver more observability to the developing teams, at the same time do it cheaper. you know, so it's a little bit of conflicting priorities. speaker-0 (01:35.49) I think everyone's familiar with Coinbase reporting pre pandemic, the sixty five million dollar spend on data dog. speaker-1 (01:42.114) Yeah. Yeah. Yeah. It's interesting. There's there's it's not unusual. You know, we do see a lot of customers that have very, very large bills. and the thing that we're hearing from those customers is that they just feel that there's just misalignment between the cost and the value. You know, back in the days when you know, interest rates were zero dollars, right, and and money was free, it didn't matter, right? Like you could just go and spend money, and all you really cared about was that velocity, that engineering velocity, go and get things done. but then when things changed and we started caring about money again, it matters now, right? People want to still You have their observability, but they really want to align that that cost and value and make sure that it's making sense from a a business. speaker-0 (02:16.43) How have you been able to overcome that sort of duality there where it does seem like fundamentally, since if nothing goes wrong, it's a lot like security, that's the best you can hope for. And then when there is an issue, you are actually reporting on that effectively, so you can dive in and solve that. I I think Companies historically that have lots of problems are easy to justify a spend and observability. But the ones who need it the most tend to have fewer issues reported, right? The ones that have to have higher reliability. So especially the changeover in since we've left the money is free era, how have you been able to sort of make that justification or the mindset shift for those customers to really understand what they're getting out of the stack? speaker-1 (02:58.37) comes down to is you know really focusing on developer productivity. You know, our engineers are very busy, right? We've had this, you know, shift over the last decade or so a little bit more, right, to this DevOps motion, right? Where more and more is being put on the developers to go they have to be responsible for, right? Not only are they writing the code, they're testing it, they're deploying it, and they're operationally responsible for it. When you have them operationally responsible for things, you want them to feel comfortable that they're not going to be getting constantly paged at night, you know, that they're that things are working reliably. And if they're not the first thing that they're gonna do is they're gonna slow down the velocity of deployment, right? Just to get back to how do we keep things sane and safe, right? But for some organizations, you know, the risk is the, you know, reliability, right, and the customer impact of unreliable systems. For other customers, the risk is the lack of velocity in in speaker-0 (03:43.774) On the engineering side, how historically observability tools, monitoring, alerting, logs, traces, et cetera, are this sort of unfortunate situation that you get into only once you actually have an incident in production. how are you seeing that be a shift forward or earlier in a way so that people are involved in understanding how their stuff will be available when those incidents happen? Or when you talk about productivity, my thought is, are you seeing a fundamental shift in how development is being done or engineering is being done so that they have those tools earlier in the process? speaker-1 (04:16.558) I think it's still very similar to how it's always been, right? There's no, you know, motivator like pain. Once you're your first you have your first incident, suddenly then you realize the importance of of observability. right. There's nothing worse than, you know, a system being offline and you finding about it because a customer has told you. Or even worse, it's broken, but you don't know why and you don't have any data to understand what went wrong. Right. The the the best you can do is turn it off and turn it back on again and hope things start working again, but you just don't have the the data that you need. but what we're seeing happen now is just you know, with the the tools that we have available, right? Certainly with the AI tools for being able to build code, it is so much easier for people to, you know, to create new p features and and new software. and so we are seeing that velocity is really accelerating. The level of expertise you need, right, to be able to ship software has, you know, has dropped significantly. and then there's also questions about you know, how what is the quality of that code that is being shipped, right? How reliable is it? and so certainly there's a a need, right, to understand You know, how these systems are working and what's happening. but it's still a case of, you know, sometimes you have to wait until you know comes back to bite you first before you realise, you know, how important. speaker-0 (05:20.408) So if engineers are writing code faster, or if we can maybe we shouldn't call them engineers anymore, if code is being written faster by organizations and there is something to be said about the average level of code potentially going down or getting closer to the mean in any organization, that means more bugs will show up. It seems right on cue that we really need to be talking about observability, honestly, here. because what we're saying is worse code potentially could be getting to production faster. I I guess I'm waiting for that magical Grafana report to come out that reviews all the customers' logs and how many alerts they've got and and whatnot and actually can point to some real data about this because there's lots of people standing up and saying, we're faster, but we're better. And from my own personal unfortunate experiences, it seems like, yeah, sure, faster seems like it's the opposite of better. there is a inverse relationship here. So what's what's the solution? Is there something that we should be looking at? Is there a different way that we should be thinking about it? Is it just a matter of having the right tools? speaker-1 (06:16.078) And so there's a couple of interesting things, you know, certainly we're we're going through now, you know, in this new AI phase is that one, it's been much easier to go and ship code, but now, you know, there's a lot of questions around, you know, are they actually valuable? Right? Are they delivering value to the business? Right. And this just comes down to classic observability, right? Do we understand what the system is doing? Do we understand how our users are interacting with it? do we understand how much it costs to run this thing? And so all of these things are becoming more and more important, right? You know, people have have now realized it's not just about shipping features, right? It's understanding like Are we shipping the right features? You know, and are they doing what we want them to be doing? Are they having the the desired impact on our business? If we've gone through the the hype cycle of of everyone just well, let's just ship as many things as we can, to now where we're realizing, well, actually let's focus on things that actually are delivering value. Right. And success is not just about how many tokens am I birthed. speaker-0 (07:05.72) Yeah, I think measuring the widgets is sort of the the problem here. Like what is the appropriate metric for for use? I think historically the knowledge industry has moved off of measuring lines of code, although it seems like we're back to that as the only reliable metric left. but I always like pointing to to the original Dora metrics. It does feel like there's something that was always missing from them, and I think you really hi ran into it here, is that it's about the business impact, realistically. And we've already seen a huge impact to GitHub, the product and organization as a result of lots more code being generated. I don't think anyone's brought up so far. There's maybe I'll say even a competitive advantage of using an open source technology as your standard, or an open standard as your interface, because you have access to or everyone now has access to models which are repeating that information. And if you do something fundamentally different, then you're actually going to be at a competitive loss against the rest of the industry. Open tracing has been around for a while, but the OTEL it's I think it's twenty nineteen or so. how did Grafana Labs sort of port what they had to that? Was it an easy transition because it was something you were always involved with, or was there a lot of stuff that had to change under the hood in order to make that an effective first class integration? speaker-1 (08:19.668) Again, we're just kind of lucky with timing where, you know, certainly we started out, you know, Grafana, you know, really big focus on the infrastructure side with, you know, Prometheus. actually, you know, even Graphite before that, right? So actually I'll tell a bit of a story here that's interesting, right? So like when we started Grafana back in 2014, right, the the telemetry database at the time was Graphite, right? That was the the dominant one and that's what Grafana was originally written for. but from day one, we really wanted to focus on integration, right? And so be able to support other different Time series databases on another different data source in Grafana. And one of the nice things that we had is, you know, if part of the open source software, we'd have some anonymous phone home data that would come back from all these Grafana instances around the world and we could see which data sources people were using. so that really helped us kind of see this massive explosion of the use of Prometheus, right? Which is why we shifted back in 2017 to change and go and focus on the Prometheus ecosystem, right? Because we just saw this massive adoption of it and all of that driven by the adoption of Kubernetes as well, right? Which was the you know Prometheus and kind of a Kubernetes paired really well together. speaker-0 (09:28.126) you you almost sort of lucked out as well because Prometheus initially is more of a pull-based system, which means it doesn't matter so much what the underlying protocol was going to be or the standard that was going to be used because you controlled the whole mechanism on your side rather than being a first class API where people had to push data to, all it really required in a lot of areas is maybe flush out the push gateway or some other piece of technology as a intermediary between the technology stack or SDKs or third party products and this sort of proxy gateway into allowing Prometheus To pull from it. speaker-1 (09:59.298) Yeah, I mean, we we certainly can't take credit for that ourselves. I mean we we certainly have a a large portion of the maintainers of the Prometheus project, but we certainly didn't create it. And and a lot of it was inspired from you know from the internal Good Google project Borg, right? Borgmon, right? So Kubernetes was the real implementation of Borg, and then Prometheus was the kind of re-implementation of Borgmon. interestingly enough, Otel then came out of the new monitoring system inside Google, which is Monarch, right? And so there's a lot of I guess a competition between the two, right? And it still is, you know, basically a religious war. What's better? Push or versus pull for for metric collection. I think both have value, right, in certain situations. there's no, you know, perfect one, right? They're all good for the, you know, the right use case, right? And I'm a big believer of like pick the right, right tool for the job. but it's great to be able to see we're seeing these kind of ecosystems kind of come together. We don't want everyone reinventing the wheel and building their own, you know, observability architecture and solutions, right? We want to be able to take that decision away from. speaker-0 (10:57.998) I think it's like a lot of things, Microservice versus Monolith is a canonical one. It doesn't matter which one you pick, you're going to find a way to mess it up in the long term anyway. Yeah. I do want to ask though, if you do think that a particular architecture is going to see higher popularity in the years to come, potentially, because of the pushes that AI stack are are coming to, like not just what it's recommending, not if you just code something up and you ask an LLM. speaker-1 (11:07.019) Exactly. speaker-0 (11:25.506) You know, how should I do this? But realistically, what is actually needed by those stacks that are doing either data pipelines or model distillation or whatever have you their inference at the end of the day? Is there something that is better fitted for that? is there a particular strategy or tech stack or just ideology that comes with it, or has everything related to AI, would you say it's very similar to what we've been doing all along? speaker-1 (11:50.664) I think it's similar, but there certainly are some differences. You know, we definitely certainly from an observability perspective, we feel as though this new agentic kind of phase that we're in is like a new software architecture phase. As big a change as moving from monoliths to microservices, you know, from from what we see because it is a a new way of writing software. It's it's a new group of people writing software in some cases. but also it has some new observability channels. speaker-0 (12:12.546) One thing that I haven't been able to figure out fundamentally is whether or not the long term trajectory though is accelerating or if it's still constant. And what I mean by that is we have seen a lot of technologies or ideas or paradigms pop up in the last six years or so where they are very quick to be spun up and then very quickly get annihilated or removed because they aren't the optimal long term solution. And I'm wondering if from those major milestones one to another, if that is shrinking in time. Or if you wait on understanding what the industry actually needs and then delivering something, if you say you just released AI observability you know in 2026, is that is that late? I mean, from my standpoint, I'd say no. That's actually ahead of the curve, like you are riding the innovators wave here because most organizations don't have even running their own models or or inference in their own data centers or even interacting with providers to do that. So I'm curious how you think about whether or not it's better to be supporting the innovators and early adopters here or whether or not it'd be better to wait, see us actual standards pop up that make sense, and then build products to support those. speaker-1 (13:22.05) Yeah, it's a that's a great question. so it's very challenging right now. I mean, I think there's certainly some foundations that we want to focus on. and that's where you know, we see things like open telemetry being really important, right? Like, you know, that's a foundation technology, right? Should be part of the architecture. Because no matter what happens, no matter how the world around us changes, we still need to have visibility into what our software is doing. That's always going to be True. And who we means is certainly going to change, right? you know, historically that has been our SREs and our developers, but more and more of that is now our agents who still wanna have visibility into what our software's doing, right? But they still need that telemetry data. So for us, I think, you know, what we're understanding is that those foundations now are really important in in getting it right. speaker-0 (14:06.208) If everyone believes there's a bubble and everyone believes it's going to pop, that you really have to question whether or not that is the collective y wisdom that will follow. but th there there's a whole economic philosophy there that we're not gonna go into. But the thing I will ask is whether or not you feel like the the curve that was originally identified in crossing the chasm, that you start with these innovators and early adopters, early majority, late majority, and then laggards, are you focused in a different part of the market from where you were before. Whereas maybe one could say, yeah, we're totally early majority, late majority, those are the that's the best place to look for, you know, getting growth or customers or really just any sort of revenue. but if we think that it's changing, that maybe you're focused now in pushing to a different spot because you see that growth or the amplitude of the number of customers or the market segment there to be much bigger than it was before. speaker-1 (14:54.38) Yeah, definitely. I mean, we definitely are playing in the the early majority, right now. Like we saw that change Certainly for for our business a few years ago. and a lot of that is around open source, right? And we're like, hey, we've got this great open source technology stack. you know, you can kind of mix and match and do what you want and you can choose how you go and deploy our software and and you know have all these flexibility and choice. And then we were discovering, you know, the customers we were talking to were like, I we don't want this choice thing, right? We want you to just tell us how to do observability and just do it properly. And so we're like, okay. So we then took, you know, our learnings and we went and built, you what is our Grafinal Cloud platform today, right? Which is more opinionated, kind of out of the box. experience, right? Where we focus less on building the technology and focus on building solutions, right? What's interesting is with those users, like you know, even though they're, you know, the early majority, they're typically a little bit later to adopt technology, they're still adopting, you know, this new AI identic model just as fast as everybody. speaker-0 (15:46.926) People want to get into it because they think they're gonna miss out and that just drives more and more the cycle of of adoption. What I do want to ask you is in the new product line changes that what you're delivering now that you said the new features that are come out with the Barcelona conference this year, was there just a shift in what you're tracking or also how you're exposing that data back for agents to consume it as well? speaker-1 (16:11.16) There's a big change in how the data is consumed. you know, which is very interesting for us, right? As a as a company that's founded itself on being the, you know, the dashboarding company, right? To discover in a you know, the new world that we're moving into, it's like how important is the dashboard anymore? Right. so many are developers where they're sitting in Claude Code or Codex or Copilot or whatever they're using, and from that they're just asking the questions, right? And they're saying, hey. what's happening in production right now, you know, what is my system doing? And so we need to we've had to kind of change, right, and make sure that, you know, we're building a platform that that can do both, right? you know, one, we want to support this new model of of agents being able to consume the telemetry data and and look at it and understand it and they can, you know, make some some good decisions around it. but at the same time we still want to be able to you know, support dashboards, but 'cause they are they are still useful. right. Like especially when you wanna have them up on a screen. you can see what's happening or you want to generate report that you can share with leadership, for example. speaker-0 (17:08.715) I guess there's a two part question here, which is like, you know, is the branding should it still be Grafana Labs if the dashboards are for humans and you're transitioning to, you know, a more agentic, you know, interactive layer there and you know the The second part is is realistically this problem where do we think that even the dashboards may be the better solution for models if they are multimodal and how they're consuming data? Is there something captured in the way the dashboard was created? Or fundamentally do you need to change APIs to actually expose the data in a different way? Or is it you've already had this data all along. It's already right there. It's just a matter of having the right MCP server set up or having the user set up some configuration for access and so it's an access control situation. It's not necessarily something fundamentally has to be different. speaker-1 (17:53.196) Yeah, I mean there's definitely changes that we're making to accommodate these new use cases. There's a couple of different things and the couple of different challenges, you know, that we faced early on with AI and and one of it is that I guess for most of us, right, observability data is a lot of unstructured data. and we know LLMs don't really work well with unstructured data. It doesn't make a lot of sense to them. it doesn't make a lot of sense to individuals, right, when we when we look at it, right? We just you typically have context to kind of know what it means. And so one of the things that we've focused a lot on is, you know, how do we build that context around it, right? The use of these open standards and open ecosystems around open climate should really help, right? Because it's really well documented, you know, data that we're collecting. speaker-0 (18:30.518) you doing something complex like setting up some sort of semantic layer for it to understand the data that you've actually collected, or are you using some sort of third party product or something built into Grafana to actually utilize this? speaker-1 (18:41.314) Yeah, so we have this capability inside Grafina Cloud where one, we're leveraging open telemetry, right? Where, you know, it does have a very, you know, well structured semantic layer. and then it's a extensible system, right? So if you're using different data you know outside of open telemetry, you can then go and define rules and build upon it. You can say, Hey, you know, these labels mean these things. speaker-0 (19:00.982) Are training your own models to work through those business cases or are you pulling an open weight model or another provider to actually help you through those internal challenges? speaker-1 (19:08.908) Yeah, I mean, so right now we're just using kind of the frontier models. What we discovered, you know, that was really valuable for us is that the reason why the frontier models work so well with, you know, our use cases and our technology is because of our open source, you know, ecosystem. Right? We've we've had this more than a decade of you know building open source technology, but it's not just about the technology, it's the the community and the ecosystem that we've built, right? There's so many people that blog about you know how to use Grafana or they write tutorials or they've got public dashboards or they've got you know all this content that's just on the public internet about both what they're doing with our technology and and how to how to actually use it to do that. So the models just know, right? We haven't had to go and train our own models to to get to success. they just work out of the box. And that's a very big competitive advantage that we speaker-0 (19:50.958) And I'm very excited about asking questions to the frontier models and getting back relevant context about your your products or how to implement things from a technology standpoint. Or I you know, for us recently it was like actually understanding our business case and being able to explain it, you know, fitting in the market segment, especially when you're doing something very nuanced. So I I mean, I don't think I've ever heard anyone stand up and say, Yeah, I can't wait to, you know, build our own models. I I have a lot of extra cash I want to throw away and burn. And so this is gonna give us a great opportunity. there is something to be said about small models. that can run faster and are trained on a very specific data set, but when you are utilizing open standards or you have such a huge prevalence in the market, y you may not ever need to actually do that. speaker-1 (20:30.926) I gonna say, especially with the the how fast the models are are evolving, right? Like it's very difficult to compete against Anthropic and OpenAI and Google, et cetera, right? Like they have a a a much bigger budget, right? And an infrastructure footprint to go and train models than anybody else. speaker-0 (20:44.418) Well, you know, if anyone is actually concerned about that, I I would say that realistically, a lot of different ways in which you can utilize a model are easily utilized via the non frontier models by the private model providers. There's a lot of great open weight models where if you and I hate to say this, if you're just prompting it right, y you can get a good response without a problem. But I I I mean I do understand what you're saying, you if you're not doing anything, what I have found is that for a lot of use cases. the system prompts that the model providers, Gemini, Anthropic, OpenAI have do a l huge disservice to the value that they're offering. So if you already are on the path of managing your own system prompt and what how the agents are really focused on and the data that is exposed to them through MCP or or tool calling, I think you can get really far today and probably a lot farther if you are just running your own smaller model off the stack. Something that I I I saw and I want to get your perspective on this is that recently, especially with the increase in open source slot, is that we've seen models and people utilizing models to create a lot of pull requests on open source libraries. And on the flip side, we've seen this mistake that'll that some companies have made, most notably HashiCorp or Local Stack transitioning their license models to some sort of business license. I really appreciate how Gravana's used the AGPL version three. And I I just don't understand why anyone would ever switch. And maybe after this call you'll be like, we're we're we're gonna be switching. I mean, you basically have said how important it is to contribute publicly open source, have out people out there who are generating for their own benefit docs and guides on using Grafana who you pay zero to basically and get all the benefit from. And you lose all that by switching off. But what I do wanna ask is since the now huge amount of speaker-1 (22:10.13) Ha ha. speaker-0 (22:33.622) capability or quote unquote productivity that we see open source developers have in integr interacting or creating pull requests to open source projects. have you had to deal with any of the blowback from that on your side, how have you have there been any challenges to deal with potentially increased pull requests or spam or noise on issue tickets, etcetera? speaker-1 (22:51.97) Yeah, I mean definitely. I mean the first thing I can say, yes, AGPL, right? We have no no plans to to move to a a closed source licensing model. in fact we're really excited. I mean, like you know, we've seen a number of you know big open source well known open source you know companies and projects that move to a a closed source license, you know, typically SSPL, who are now coming back, right? Who are now adopting AGPL, right, just because they've you know and they're referencing us as as as a reason for that, which you know we we're very proud of. But you know, open source for us, you know, it's it's important, right? Like it's a deliberate business decision, right? We see a lot of value in building healthy ecosystems, right? It's not just about you know, the code itself, right? It's about building vibrant ecosystems and a community of people that can kind of share knowledge and kind of come together. And that's what's really important to us and and that's working really well. And obviously, you know, we were Apache licensed and then we moved to AGPL because, you there is still that threat of, you know, the hyperscalers, you know, taking your open source technology and and, you know, c you know, getting all the revenue from it, right? So You know, you do want to go and protect that. And but we do feel feel like the AGPL model does give us the protections that we need. We are seeing, you know, a lot of, you know, pull requests coming through and and it is problematic. And, you know, I think we're having a similar approach, you know, where we're we're kind of just advising people just don't send it, right? One of the challenges is that for large you know, open source projects and large code bases, one thing that's really important to make sure it's is to well, one thing that's really important is to make sure the code is maintainable. And so there's a lot of work that we go and do to enforce our own coding standards and our own way of doing things. And so as we're making changes, we want to make sure that the changes that are getting introduced kind of match the rest of the code base, right? And are aligned with those standards that that we've got internally. And a lot of the AI code that's generated, they're not generated based on our standards, right? They're generated based on whatever standards they've decided is common, for example. And so obviously there's a big work to go through that review cycle and have to go and change things, right? So that, you know, even though even though the pull request may implement the change and it and it may, you know, be technically correct, right? Like we still want to make sure that the code's maintainable. That is a big challenge where yeah, it's very difficult to accept pull requests that are just AI generated, 'cause they don't they don't deliver a huge amount of value to the open speaker-0 (25:07.584) projects. Yeah, for sure. I is there any sort of technical implementation or technical solution that you've put in place or non technical implementation to help combat this in any way? speaker-1 (25:16.782) And I know there's some conversations, certainly in the Open Telemetry project, right, where they've they've had some pain around this. And I think they just, you know, had the approach of just like, you know, saying, you know, we're just not accepting, you know, AI generated pull requests. right, for one. you know, they just won't get reviewed. you know, unl if it unless it looks like it's a person. I think that's just the the easier approach. what's interesting is you know, we see this challenge in in the industry, right, for software development where Yeah, we don't need new grads anymore, right? AI agents can do the work that you would normally go and have a new grad do. and so then we run the problem of we don't need new grads, we only need experienced engineers. How does how do we in 10 years time have experienced engineers if if no one's letting new grads today go and develop their skills to, you know, to become experienced. so for me, I see that's where, you know, there's a big opportunity in open source, right? Where, you know, we still want those, you know, hands-on people we don't want AI to be used, right? Where they can go and and write code. and it is a great ecosystem, right? Like you can go and write code and you can contribute it. And if you're a real person, you know, the maintainers will support you, right? They'll coach you, they'll guide you on how to go and build, you know, a better software and and you know improve things. So I see as open source as being the saving you know grace for you know how we're going to get you know new grads to develop their own you know their skills over time so that they can you know be productive in the environment. speaker-0 (26:37.518) I think that's very interesting perspective that open source loves artisanal software development, you know, hand curated by by humans. and it will set you apart as inexperienced engineer to actually go and do that because maintainers want more human approach to it. I do think there is a little bit of the Silicon Valley's the the show, Jin Yang's hot dog or not a hot dog problem though. Like ha I see a lot of these repositories popping up saying we only accept non-AI generated code, but in practice. How do you know? Like if you know what code is being generated by an AI or an LLM, I feel like that is in its own impossible problem to solve that lots of people have been trying to work on for, I suppose, now only four years. So I I I actually wonder in practice, like what like is this just a human going around being like, yep, that's LLM slot, close, close, close, close. Or if there's an automatic system that someone's actually figured out to correctly identify LM generated stuff. speaker-1 (27:32.502) I mean, I think it's challenging, right? As soon as you build a tool that can identify what looks like, you know, AI content, the AI changes and and, you know, works around it. you know, that's always gonna be the case. I think really what it comes right now, it's certainly a almost like a knee jerk reaction, right? Like we're just getting overwhelmed with, you know, a bunch of of you know AI slop and and noise. And so, you know, we're we're taking an approach to kind of filter that out. I think, you know, over time I think we're gonna be able to kind of focus on Not saying just no, right, but articulating what it is that we want, right? Because there there is probably a future where there is some AI generated code, you know, where it's going to be valuable, right? So it's really about us, you know, you know, as projects, really understanding what is the the key things that we want, right? The main thing, you know, certainly for us that we understand is we want maintainable change. speaker-0 (28:19.212) Yeah, I think I'm first of all, I think that's a super mature perspective compared to like we don't accept any of this. Because the r the realistic r result is, well, we actually would accept it if it adhered to all of our, you know, culture expectations and coding expectations. And I think just a lot of these repositories don't have that mentality of like what sort of features are actually the right thing to build and what is the right way to do it. And then you may be actually motivated to add in some automation to codify your expectations on which features get in. How big should the features be, et cetera, et cetera? Because I think we are very quickly going to end up in a world where you why would you reject LLM generated pull requests? That's not what you care about. You care about LLM generated pull requests that are useless and have a long term detrimental effect to the longevity of the project, et cetera. One of the challenges with open source technology has always been, and I sort of hinted at this before, and I'm gonna try to get you to say something on the record, is the security aspect. There's a lot of thought that what was open source had in an inherent improved security mechanism because more people could look at that code. But I feel like over time, especially with the model usage increasing, everything is being looked at for vulnerabilities. And one of one of these things actually turned out to be a problem. I don't recall exactly what Grafana was using, but was susceptible some of their open source repositories and closed store stuff to Shihalude, which is just this nasty worm out there that takes its name from Dune and basically just replicates itself by stealing MPM credentials, et cetera, and asserting itself. And I'm sort of curious, how did you actually maybe you know, how did you actually catch that this was happening? And the result was basically you're like, we don't need to pay any ransom for this. speaker-1 (29:59.448) What we're seeing is is just like the supply chain attacks, right? Is is the the path that's being exploited right now. And you know, the reality is is everybody's using you know open source code, you know, even if it's a commercial, you know, proprietary system, the reality is they're still, you know, vendoring in a lot of you know, open source dependencies, right? you know, it's a common thing, right? Like why do we want to go and reinvent the wheel? So it's a problem that's affecting everybody, right? Certainly not unique to to open source users. And you know, and and certainly, you know, we were very transparent with you know, the incident that that affected us a couple of weeks ago. The summary of it for us is yes, supply chain attack, they were able to get a credential from a developer and use that to to access our GitHub repos. All they were able to get was our code bases, right, which For us, you know, Motro, that is open source anyway. right. but you know, because of the controls that we have in place, you know, at no time were they able to access any of our production infrastructure or or access our customer data. and so you know, that's great, right? We're really proud of the, you know, the fact that we've got these controls in place that we're able to kind of protect our data. There's certainly more that we're doing now to to protect things even further. but certainly yes. you know, when they came to us and today we've we're gonna release all of your code. If you don't pay us, you know, like it was an easy choice for us to say, well we're not gonna pay you one because, you know, how do how do we guarantee you're not gonna just release it anyway. But also like it's mostly you know open source code, right? Like you know that's not where our you know, differentiation lies, right? Like, you know, much to the displeasure of our sales team, you know, most of the technology that we're building these days goes into our open source. or your lawyers. And so yeah. And so, you know, that that was a an easy choice for us to do. But certainly it was a lot of work. It was very impactful for the business to be able to, you know, have to go through and audit everything, every single change on every single repository. Like what, you know, has there been any compromises? speaker-0 (31:54.358) love the perspective and the article that was released, honestly, because it was like quick and timely, realistically and a very methodical approach to what was actually reasonable done rather than hiding behind, you know, secrecy for for months on end and saying, well, you know, something did happen because someone else reported it from a security agency or a, you know, a security tester. no, I always just think it's interesting what happens and I I'm not sure what the future will will be here, especially because a lot of different package managers are attempting to take different precautions in place. And I'm hoping still that I can get someone on the record who knows a lot about package manager security on the on the on the show. I'm just waiting for the right person. so if you think that's you, you know, please come and ping and ping me. and we'll we'll get you on. What I do want to still quickly jump into is you did mention the challenge with some of the other hyperscalers basically utilizing the solution. And we have seen in the last few years or so Grafana actually partnering directly with them to offer a managed version at a cost reasonably And I'm wondering whether or not, given the current world views, whether or not there'll also be similar partnerships we should expect to pop up for other similar cloud providers, especially ones in Europe. speaker-1 (33:11.094) to be, you know, everywhere and available for everyone so that they're familiar with it. one, because it's great technology and why wouldn't you want to use it? you know, but also from a business. you know it's great for us to, you know, for people to be familiar with our technology and and our tool set, right? for sure. And so yeah, we certainly did the the partnerships. You know, Amazon was the first one. where you know we license our technology so they could deliver the the Amazon managed Grafana. we obviously did similar partnership with with Microsoft you know the Azure platform as well as Azure managed Grafana. We're always looking at other partners. You know what's what we're finding interesting is that you know certainly for a lot of our users, you know, Grafana Cloud is going to be their preferred method just because that's where there's much richer capabilities, right? And it is a a more kind of out of the box experience and kind of solve sort of the observability problems. And we certainly, you know, have built a solution that supports all the different cloud providers, right? So it doesn't matter who you're running, it'll be, you know, it'll work for you. And we see many customers that are multi cloud, right? That are are in different providers and and are using different systems. And so they need, you know, an observability product that can work across all of those, right? And be able to correlate across, you know, all the different providers. So yeah. speaker-0 (34:17.326) It's hard it's hard it's hard to do it any other way. I mean otherwise you you build it and no users come or you know you waste a lot of money on on something that doesn't make s a lot of sense for sure. I think at this point it probably would be a good opportunity to switch over to picks for the episode. So Anthony, what did you bring for the audience today? speaker-1 (34:36.575) Yeah, yeah. So that's great. Well, so obviously for me, yeah, it was a bit of a challenge, you know, like what what am I interested in? you know, the reality is Grafana Labs, that's my my passion. but we've already spoken enough about that. but one thing that I thought was really timely was a an essay that was brought to my attention. So Raj, our CEO, actually brought it up last week when we were talking, which is it's called The Bitter Lesson, right? So it was put together by by Rich Sutton, right, AI researcher out of Google and it's really a a very an interesting read, timely read. It was written a long time ago, back in 2019. But it's really a kind of a focus around, you know, the looking back at the you know evolution of AI over the last, you know, 70 years, right? For as long as we've been doing this and kind of understanding where you know what what we're seeing playing out, right? Where the reality is is that you know, trying to build more kind of domain specific you know tools, they'll they'll always get beaten by generic tools where you are just throwing larger and larger volumes of data at larger and larger volumes of compute. right. So we've seen this with the rise of LLMs, right? you know for years everyone was trying to build you know domain-specific kind of AI tools, whether it was in you know whatever say the medical field, right? right. But now we're seeing that the large language models, right, where you just throw all of the data at it, right, and you just process it with more and more compute power, you end up getting better results you know, than these these unique tools, right? the good examples that it brings up is around like even chess, right? Like where, you know, started off they'll, you know, encoding chess strategies, right, into, you know, these I tools to make it think like a person. but then the way they made the most successful ones is where they just said, actually let's just build big neural network and just have you play chess against each other as many times as possible and you'll just you know, with the volumes of data, you'll end up being better. and that's, you know, that's the world that we're living in now. There's companies, right, you know, that have just disappeared overnight because they built these, you know, domain specific AI technologies and now, you know, these generative AI tools have just come out and just completely just, you know, overrun them. and it's always going to be the case where, you know, right now we talked about earlier today, right? Where, you know, sometimes building those more domain specific kind of things works for some things, right? For a while, right? but then what's happening is then speaker-1 (37:03.074) the larger models are then getting better at it, you know, inevitably over time anyway. And so, you know, we're always having to constantly kind of iterate. And so I think it was really good you know, kind of read to go through that and kind of reflect on that. speaker-0 (37:16.022) Right. The original deep mind approach that solved chess, I guess, realistically, wasn't wasn't trained just on chess. It was completely unsupervised training there. I do think there's an interesting perspective here, which is that realistically, the reason that companies are getting overtaken by the I hate to call it Gen AI solutions instead, is because they didn't build or innovate on their space of having an LLM model that solves that problem. They just added an LLM to their space. And so you They basically doubled down on having the best LLM, but for that area and they weren't doing anything to d to get around it. And I think there's a question of like where is the end of innovation for LM design? And a lot of people are saying, yeah, it happened in Opus four point six or or whatever. And as long as you can beat that on a benchmark in your particular area, maybe we'll come back to this particular space. But unless you're gonna innovate and buy the cutting line. GPUs to do the model training or be able to fine-tune existing models out there, I do think a lot of those companies are going to fail just because it's not necessarily that the models are better. It's just like you don't have the technology to train to the level of capacity or I reasoning that the companies that are dedicated to the actual hardware to work on. So I I I hope long long term we see an opportunity here where smaller models exist because it will be cheaper to maintain them and maybe increase the accuracy. But I think we have to get much closer to wherever the end of innovation for the LLM creation to actually apply. I agree. so what did I bring? had to rack my brain over this. since since I started running engineering teams, I kept getting into popular culture. How good is this leader on screen? So think of your favorite character and your favorite television show or movie and see how they actually perform leadership qualities in the gambit. And I can't look away now. Like every single time I see a group of people on the screen, I'm questioning like, who is the leader? How are they working? Is this even a good leader? And you go back and you'll watch some old stuff and you're like, wow, this is not a team. This is just four random individuals with no organizational structure whatsoever. And the person who's called leader is just doing the worst job ever. I'm sorry if I broke anyone with this realization. and so over time I'm just like, well, where is the best leader? The the one, the character that seems like it's the best emulation of a leader. speaker-0 (39:36.098) And it makes sense because I think leaders aren't consulted on the set, like you get medical doctors or military personnel or engineers when they want to design something or a process. But how often do you see like reference consultant for leadership on the set of some show before they're like, Would a leader even say this? and so my pig my my big is gonna be the best leader that I've ever seen on the show, which is Captain Pike of Star Trek. and it it's not the best Star Trek show, but it's definitely my favorite leader. Excellent. speaker-1 (40:05.88) Yeah, I mean I think it's it's important, right? But if you want to be a manager, right, if you it's a very different skill, right? It's a very different you know, problem set, but it's a a very, very, very different role. very different set of problems. speaker-0 (40:18.54) I I think just the single point of failure for someone doing everything in the highest leadership bot in an organization is just so problematic. But anyway, so this has ruined so many television shows for me that, you know, I'm watching it, I'm just like, This is so unrealistic. This it doesn't represent reality in any way. I guess my other theory or perhaps it speaker-1 (40:38.488) That's right. It's just they just they could be a terrible leader and that's there's plenty of those that exist, right? speaker-0 (40:43.15) yeah, yeah, for sure. I I don't remember what it was, and I probably should find a reference to this. A long time ago I read this thing like if aliens came to Earth, what would be the thing about humans that would be most surprising or shocking for them? And for for me, you know, it just goes along this angle. I think it's the great attention spent in the last a hundred years on how to manage humans effectively. But that's that's a whole different philosophical topic. So in this episode, we we we actually reference a lot about productivity and we have a separate episode dedicated to productivity. So if anyone is actually interested in that, that will be linked in the description, along with anything else that we've we've talked about, and there was quite a few number of things. So thank you, Anthony, for joining us today. it's been honestly a great episode. And here's a short reminder to everyone who's listening to Click. speaker-1 (41:27.202) No worries. Thanks Warren. I've really enjoyed it. speaker-0 (41:31.648) subscribe to the Adventures in DevOps or leave a comment in the episode. You have no idea what that means to me. I if you want me to make the right content and get the right guess on, make a suggestion or, you know, say what you like or don't like. And I hope to see everyone again back next.