speaker-0 (00:08.119) Welcome back to Adventures in DevOps. It's like the world suddenly woke up to realize that maybe testing is a necessity shocker, really. And historically one of the most challenging places to do it is in the browser. So I'm hoping we can get down to it. And I brought in a longtime test expert, previous engineering manager at Mozilla, and now head of open source at Browser Stack, David Burns. Thank you for joining us. speaker-1 (00:31.992) Thank you for having me, Warren. I'm really excited to be here. speaker-0 (00:34.786) We are in this space where the things that we're building are being utilized in browsers as a primary mechanism for interaction. And yet I always have found in my two decades of software engineering experience, actually validating what shows up there is always the most difficult part. Like I write a unit test on some back end code and it's it's a joke. But as soon as it comes to is the CSS working, is the components actually showing at the right time, is their gestures for mobile stuff actually working correctly, I I mean good luck. speaker-1 (01:04.636) you are not alone in that. Having worked on browsers for a decade, even browser engineers have that problem. it it's it's such a common problem. Browsers nowadays, or at least for the last decade, have effectively been operating systems in their complexity, right? and you can look at it, Firefox OS, Chrome OS, even WebOS before that, right? And it was amazing what you could do from rendering to interactions, and then you have all of this on multiple operating systems who do things slightly differently where again, browsers try and do a bit of massaging to it. and then you throw mobile into it and it's painful. speaker-0 (01:48.896) I mean, I love the promise of frameworks, right? Because it creates this abstraction layer that potentially you could use to get reliable results everywhere. But the whole world of how like CSS implementations are just fundamentally different between the different frameworks that power each of the browsers and the the the classes that are built in or or the flags or whatever that you're utilizing. And you need something like even if you're using a framework, I'm sure too many people are unfortunately using React, but if you're using something better, you still have to deal with this like. post CSS processing where you take what was actually generated and make it work on literally ever every browser and it d it still doesn't work out of the box for everything. So speaker-1 (02:25.486) So there's a lot of problems that I think when we're developing software that we really struggle with. CSS is a fantastic example, right? Like the Peter Griffin trying to open the the blind, and you see how that goes wrong. That is CSS. I'm like you. I've done a lot of back-end engineering. I was never good at the UI side of things. but I like the kind of automation part of it and trying to get it to do things. And CSS has kind of grown organically. from the years and then standards got in the way and they started cleaning it all up. They had to retrofit a lot of stuff. Then you've got the other side of it is that when you're interacting with all of these things, you've got JavaScript and that runs asynchronously, right? And the browser's like, you asked me to do it. I'll let you know when I'm done. No, never not not before that. speaker-0 (03:11.756) You're absolutely right. There's there's a lot there where there could just be problems at every level. The insight of understanding that JavaScript fundamentally only worked on an async level was a was a huge breakthrough for me in any sort of front-end development. Because before that it was pretty much this novel aspect of, well, there's this way you can sort of asynchronously work with stuff. And maybe if you used any sort of strong type language, you you may have encountered some sort of a weight. But when it comes to JavaScript, there's there's literally no way to get very far without utilizing this part of the of the technology. And I agree, like most engineers that interact with a website will not understand what this fundamentally means or how it works. And it does cause a lot of tests to end up with set timeout like 100 milliseconds while you wait for some execution to happen. You know that this is wrong, but especially now like LLMs will write these tests. And obviously there are some better ways, but the the default mechanism is not so obvious to even write the tests. That you want to validate the code in the right way. In doing some UI development for about five years, I got much better at writing CSS classes and attributes. I still am a terrible designer, but I I can at least picture the 2D layout and use Flexbox to make the right thing happen. And then there's like a lot of extra complexities on top that I still don't fully understand. But now I feel much more confident in making a website. But the testing is still Just gotta give up. like I'll I'll I'll pray that maybe what is it, playwright works out of the box and you can actually write the appropriate test there. But I don't have any confidence that actually what gets written will match and actually validate things correctly. speaker-1 (04:52.391) And it's interesting that you said like you would just want something to work out of the box. speaker-0 (04:56.47) know, you know that there's an issue when the mechanism to actually validate some particular test you're writing, and I I think the most common one is like use some sort of automation key IDs in your HTML or your you know, in your attributes that allow you to find certain elements during test time and and perform the right thing. So you can do it hopefully not just deterministically, but reliably, no matter what sort of elements actually show up there. And that's sort of different than like Well, if you grab it and programmatically interact with the elements, is it the same as a user interacting with them? Maybe not. And so like I feel like you brought up a bunch of different technologies, and it may be good to just go over each one of them. My understanding is Selenium is more of a like a web driver. Like you use your real browser and it's sort of like interacting with it directly. And so I can imagine that there's a lot of complexity that goes into that interaction. Like you're crossing processes. Obviously, the operating system is a huge complexity. speaker-1 (05:49.708) Yeah. So like Selenium is a browser automation framework. Let me do browser automation and do it really well. so the project works with Google, Mozilla, Apple, Egalia, like all the people working on browsers, right? And kind of the underlying technology is also used in a thing called web platform tests. Web platform tests are how all the browser vendors make sure that there is inter interoperability between each other. Right. So the test need to pass in Chrome and Firefox and Safari and on different operating systems. And that's going through the W three C. I chair the the working group on that. That's how Selenium does it. But then like Cyprus as an example, they do things slightly differently. They inject JavaScript into the page, manipulate the page as you want. So if you want to click, it will do it. But it does it in in a way that is not how a user would do it, where Selenium focuses on kind of being more real to the user. When you click, like if you're on YouTube or whatever, you click on a link, there's extra attributes added. and then they're normally for security reasons, so you can tell if it's a bot or not. But Cypress is like, well, we just want people to be able to test things and test them quickly and not have to worry about setting certain things up. And there's certain complexities between those two points. and other frameworks try fit in somewhere in that spectrum. I speaker-0 (07:15.338) the full spectrum as treat every interaction on the operating system as sort of a black box and use a web driver. And so you're interacting like through the operating system to get into the browser to the application that's running in order to interact with it to the extreme on the other side, which is basically make a fake browser that you're rendering all your components in. And I think this is like sort of the puppeteer space and I think JS DOM is here and there's there's other things. I think in my experience and in the teams that I've run, no matter which way where you go, it seems like they all have a problem in some way, right? Like if you're interacting across the operating system, then you have to trust a real browser, a real process spinning up, and the security layers that are, you know, built into that, but you get the most real experience there. And on the extreme other side, you don't even know if JSDOM is rendering your HTML correctly, let alone, you know, finding the places to click on and whatnot. And then, as you mentioned, it doesn't necessarily correctly match. you know the production reality because different browsers will of course respond differently. And I think the the biggest one that you've mentioned, which has caught me a lot of times, is things like Safari will block link redirects if they happen asynchronously and take like after more than like a hundred milliseconds after a user has actually interacted with the page. And so if you have these like back end things that are processing and then it makes an API call and then a redirect, good luck. Because while it will work in every single test you write at runtime Who knows what actually is happening for users using Zavar? speaker-1 (08:38.722) Yeah, and and the same can be said. Like, I know a lot of people are moving to Playwright. Playwright ships its own versions of Chrome and Firefox and WebKits, doesn't ship Safari, and they have different security model to what the proper versions of each of those have. You know, if you went to a page and it had a hidden iframe with malware installed on it. The normal browser would block that speaking to the normal page and then obviously getting onto your machine. But like Playwright will go, Well, you know, nine times out of ten, 'cause testing, you don't care about that thing. So we'll just turn it on. It makes testing a little bit easier. But then you've got this thing where it's like, it's not exactly how people would use it. And then you get your situation like you said with Safari and you're like, Why is this happening? speaker-0 (09:26.914) The one that caught me most recently was BF cache. Like I didn't know this was a thing. And now I'm like, I don't understand. How how is it like much later that a user leaves a page and comes back to it and they're in the exact same state that they were? There was an expectation that you load a web page and you have no memory allocated. Obviously what's in a local store or in cookies, those are preserved and obviously whatever state you have in the back end. But how do they have these things populated? It's like, well, shocker. Now browsers like try to save the actual memory state and repopulate it. And you know, the introduction of this obviously breaks every single expectation you have in any test anywhere. Is everyone just testing their applications wrong in browsers? Is there or is there like no actual right way to have a conclusive non-flaky test? speaker-1 (10:17.934) Truthfully, I think flakiness is down to a mismatch in expectations. and some of it is unknown unknowns. So BF cache, right? A lot of people have no idea what a BF cache is. BF cache is this like very specific cache that browsers have. Like if you hit the back button, it needs people don't want to have to wait for a page to reload, right? And especially if you're on like slow internet, think mobile phones is a perfect example. It loads super fast, instantly. The downside is is that like if you were hoping for it to insert new data, like you have to have code that kind of goes, I'll run on a like a time and just see if some there's new data and things like that. Again, it's a mismatch in information that people have. and I think it comes down to like browsers are super complex and hard, like they're operating systems, like you know speaker-0 (11:09.002) I can see both sides of the argument, right? From an application standpoint, if you're engineering a new one, you only want to use the features that you know about, right? Unexpected features are likely going to cause a problem at some point for your users. However, on the flip side, if you know about the functionality that you want to have, like, over time, we know users are going to leave our page and come back. It would be great if we stored their state automatically so they didn't lose it the next time they loaded the page. Especially for applications where they use it for more than a few minutes a day or hours a day, you know, maybe it's a professional application, a business application, quickly loading that state is valuable. And so now you have all all the applications we have online require basically storing state. And at that point, you're like, well, wouldn't it be great if every single website needs to implement this functionality for us to push down the ownership down a level, right? Rather than having to implement your own store to cookies or local store or whatever, or you know, some sort of SQLite database. that's in the browser to automatically repopulate quickly, what if the browser did it and then it implements that? But it does make it extra challenging to make sure that the application is actually doing the right thing and that requires, you know, throwing these tests. speaker-1 (12:18.99) If you're very interested to have a look and see what the Mozilla CI looks like, there's a website called treeherder.mozilla.org. It's open to the public because it's Mozilla, everything's open source. And you can see like how things are being built, how things are being tested. When I was joined Mozilla, which was early 2010, we had there was a guy working there, Joel Maher. It was a project called War on Orange because the flaky tests were called Orange Tests. the key thing that we kind of got to the point was that we can't get a green CI. It's impossible. We're running millions of tests every time we make a push. Mozilla was also trying to like bring down costs for running all of our infrastructure. So it's not run everything. And so you might make a change and then eight hours later a full test suite runs and suddenly there's your a failure. And you gotta know if it's related to any of those changes or things like that. And so we kept working on it. and we got to the point where flakiness we knew we were never going to get rid of. We can't make it any better. So we'll just monitor that. And if we see it often enough, then we'll kind of do something about it. Either turn off the test, get an engineer to rewrite the test. Because as I said earlier, when you write your test, you're making the assumption that like you you know where you're starting from and where you're ending. But three tests above could have forgot to do any cleanup or someone put in a set timeout, like you got all these things. You could turn off one test and it could cause another test to go flaky. Right. So it's it's whack-a-mole. So you were we were trying to bring it into a state where you could clean this up in a meaningful way. The way we monitored it was a group of people, they were called the sheriffs. they were watching the CI. It just became such a point where you no matter what you did, you we knew it wasn't gonna be green. So we just try to go for how how much is the delta? Because like this test, it should only fail maybe once in a hundred times. That's okay. Or maybe five in a hundred, no more than that. Right. and if it's that one in a hundred, the effort to go find and fix that problem becomes a lot more than just letting the test run and fail. Because we go, yeah, we we're expecting that one in a hundred. It's okay. Is it speaker-0 (14:40.362) Actually, unattainable to have a perfectly working test suite continuously for any sort of large project? Or is that a like a fool's error? And like it would take an infinite amount of time to ensure that you had this, or the complexity that would have to be built to make it happen, or the the running time of the test would have to be a multiple of what it is. For instance, you could have a more guaranteed result if you spun each one up in a private container that is. completely isolated, is this like the middle ground? Or is this just a matter of at the time, you know, this is the problem that we had, and the only way we could take the next step is by watching the del speaker-1 (15:18.922) When I was at Mozilla, I had friends at Google and Facebook th who I'd regularly speak to because some of the open source projects that Mozilla was working on, they were working on it too. so like there was work on Mercurial, the version control system, and like we had talked about trying to stick these flaky tests into kind of an area and go like let's see how it often it happens. Do it all automatically. So like there's no humans really until someone needs to fix it. And so the like Facebook philosophy was that they were aiming for 100% green as they were growing. And after a while they were like, well, you know, we keep turning off as many tests as we're adding to the system. Where Mozilla was like, well we're adding to the s system, like we need to have tests run within a certain time frame. Our time frame was kind of long compared to like your average web application where we're talking, you know, potentially eight to twelve hours from you doing a commit to the CI finishing because you gotta build it on Windows, but Windows is not just one operating system, it's multiple operating systems with multiple architectures. And I'd do the same on Mac and on Linux. And on Android to solve this. I think it's becoming a little bit easier nowadays. Each tab is its own kind of profile in the way it works so that there's no data that's shared between those. they they're called user context. speaker-0 (16:42.808) Why do any sort of browser testing? Why does it take an extensive amount of time to sort of execute? And if it's like, well, spinning up the process is most of the work, or there's some single overhead and it's the container spin up or user context spin up that's that's, you know, problematic. Then if we reduce that cost in that area, then you can more easily use the same tests. You know, the thought that comes to mind for me is that it's like a completely different set of challenges for building a browser and running the test against it. Like desktop software or OS software than cloud deployed stuff, where I think about you want to avoid breaking production applications as much as possible, so much so that you end up slowing teams down by using some sort of release train. Because you wanna obviously you only want to release one thing at a time. And at the same time, you want to make sure that you're only releasing it to a small subset of the total customer base so that you can watch for changes that you wouldn't have been able to predict after all. I mean, if every change only had predictable, you know, potential impacts, you wouldn't need to do this, but there was always edge cases that that no one really thought of. and in specific customer environments for sure. And this feels like a whole like completely orthogonal or mutually exclusive set of problems with a completely different solution. The delay is caused by the cost of actually running the test suite speaker-1 (17:59.678) I think a lot of people put their tests in the wrong place of the testing pyramid, where you have a nice wide base of unit tests or small tests, slightly narrower integration tests or medium tests, and then you have these long end to end tests and you should have as few as possible. I remember this from Selenium Conference in London. Simon Stewart, the creator of WebDriver, he was doing the keynote and he said that what people should be doing Is making sure that they have one test at the end, one large test. No more, no less. And what he was trying to say really was not that you should have one test, but that one test should only do one thing really well. So if it's like your workflow is log in, do three steps and then log out, that's fine, right? That's the most important thing. It stops where bugs being in the riskiest parts of the system. You can still have another test where it's like, well, This next risk level is here, right? But it the key thing is here is not to have 10 or 100 or whatever, is to try think of just one risk area, have a good test for that. Not the longest test, because you then have other problems, and then build out your test suite. Then people tend to have faster tests and they have faster feedback loops. I also am a big advocate of trying to see if there's other ways you could get the same information. without necessarily testing. I'm sure you've had other people on talking about say observability as an example, right? Adding that to your system. and you can add observability into your test. So you can trace it from your test, or your one test, all the way through the system and back out again. Then you you have this like knowledge that you go, I know that that part of the system is bulletproof. And then we can still move fairly fast because we've got all these backup layers of A lot of unit tests, as many integration tests as we can probably get away with, but not too many they mustn't outnumber the unit tests. And then high quality for the risk areas, end-to-end tests, large tests. And for those who do a lot of UI stuff, you could also, you know, between the integration and the end-to-end tests, you can throw in component tests. If you're doing React or any front-end framework, we'd only test that component alone and make sure that's doing really well. speaker-0 (20:25.09) No matter what it is or what the advice is, there's always a way to mess it up. And I I think the the mess up here is what exactly makes a unit test versus not. And I mean, obviously we have canonical definitions of each one of these for recommendation. Even with a defined setup and teardown for a unit test, there's maybe cruft left over as far as configuration goes, or a desire to refactor and reuse. And there's some sort of shared memory or dictionary that ends up getting ported and then it causes a lot of extra complexity. Because, wouldn't it be easier if we just did this slightly differently so that we wouldn't have extra code or whatever the reason is? And you end up with a problem here. I think another aspect, especially with the test pyramid, which is my preference personally, because I've heard that there's other geometric shapes that are often utilized as a justification for how you should make your your tests. How many though is it? Like how many unit tests should you have? How many component level tests? And you know, it doesn't really matter what these things are called. There's the the very cheap tests and the very expensive tests. and I actually really am behind the only have one because it does tell people like if you have more than one like total black box test, you probably are doing you're probably doing too much there. And it's not going to tell you what you really need about your system when it fails. Or when it fails, you don't have the sort of observability that you want. But I haven't really heard any good arguments for what is the right what is the right percentage of of testing you have? And is at like at the unit test level, is it ten percent, twenty percent, a hundred percent? Like w where is this number as a recommendation like speaker-1 (22:02.968) So a lot of people will go, I need as many unit tests as possible. I need every bit covered. They're only going to take three hundred milliseconds, so it doesn't really matter. The key thing is that you're trying to strike the balance between refactorability of your tests and this underlying system. Right. So I'm a big proponent for test driven development. I hate it when people go, I need to have a hundred percent test coverage. I can guarantee you will have a hundred percent test coverage and there will be bugs in that system. It's just the way the systems are built. A guy wrote a website that got a lighthouse score of a hundred, but it was the worst website. You couldn't read it. You couldn't do anything, right? But it said accessibility is amazing. It loads super fast. Everything. It was just the worst website in the world. Python had this major problem around like ASCII versus Unicode before it moved to Python 3. And people were like, I've got 100%. test coverage, it's amazing. And then it's like, okay, cool. And my favorite way to test that was I would put a Unicode snowman into a test and then just watch as the entire system crashes. speaker-0 (23:11.192) There is a good thing about first deciding up front what the risks are, what your risk model is, where the highest ROI is, and then attacking it directly. And, you know, my personal thing is like, is there complexity in this unit test? Like things that you wouldn't expect to have been written a particular way, a need for the implementation to be something or or something else as far as the business case goes, or is it capturing like actual expectations on how the business works? And If it's not any of those things, or there's like some you know complicated math that's involved with, like this has to be a a perfect merge sort or whatever, and you've written that function and you actually want to test it out to make sure it does actually sort correctly. Those things make a lot of sense. I think where things start to break down is as you pointed out, like just asking for 100% test coverage. My question is like how much how much are you willing to spend on making that happen? And it's not just an upfront cost, right? You said maintainability. Every additional test you write is first of all, it's extra code. Which means it's additional stuff to maintain. And extra code means obviously more space storage updates. Like if you're cloning using Git or whatever, your more space on your CI, CD, more processing power, just straight out of the box. Then there's like understanding of what's even there. There's the complexity of where are my tests? Like if you're working in a monolith project versus if you have a microservice architecture, it's so much easier to work in microservices when you can just control shift F things. And find whatever you're looking for just by typing in the search box. Whereas in a monolith, you'll get like a thousand different like top level projects with all the exact same thing. Good luck finding what you're actually looking for. and let alone the the tests themselves. there's a thing in that I really like from let's say architecture, which deals in constraints. Whereas like if you if you or geometry really, if you pick two particular points, then the rest of your object is constrained in some way. You you can't define it differently. It it's only that, right? If you have like a A square and a two-dimensional plane, and you pick two points of where the square is, like how many squares can you draw with those two points being the corners? And th there's not there's a finite number of answers. You've constrained the way it works. And unit tests work the same way. So I think these are the things that definitely get missed when talking about how many tests we should have and and really percentages. I mean, obviously for some people, saying we need to have higher unit test coverage, it may be the right thing. speaker-0 (25:33.25) but at any sort of mature level you end up starting to really look at like, okay, what tet areas do we really need to test? speaker-1 (25:39.478) I th I think kind of unless you're working in finance or health or aeronautics, a hundred percent is not where you wanna be. speaker-0 (25:48.062) You know what? I'll even disagree with that honestly, because it's like if you look at where most of the issues come from, they're hum human written code or LM written code today. Right. When you make a change, that's where the liability sends, right? You know, if everything's working today, likely it's going to keep working unless you fundamentally make a change. Or or it's an outside input into the system, right? You put a snowman in a in an input box and it it you know blows up the system. You know, it's it's human input outside that, which you know, it causes a lot of the problems. Like I think a a lot of mistakes in say the healthcare world didn't did didn don't come from the system saying something. It's the humans are dealing with it, right? speaker-1 (26:26.902) Yeah. And the the other side of it is that way. Like you and I have been talking about correctness, but there could also be tests in kind of the unit space and things like that that are focusing on performance, security, things like that, right? and if you keep adding to your point about maintenance, It's a lot. I'm sure most people nowadays who are kind of using other lambs for their work get to the point where they're like, I need to restart my computer. I need to delete this because like everything's working in work trees. You run out of hard drive space very quickly, even on the smallest amount of like smallest repositories, right? 'cause you're like, I can move faster now. I can hopefully not break things, right? Like, but you know, you're moving as fast as possible and then you're just like, man, this is really painful. I've gotta slow down, I've gotta do this, I've gotta clean up, right? Like nowadays I think I'm I'm deleting build caches at least once a month because I just run out of hard drive space in the work that I'm doing. speaker-0 (27:26.114) Yeah, yeah, and I I don't see that getting better any time. speaker-1 (27:29.262) No, especially with the cost of hard drives growing exponentially, so it's gonna get worse. I've do you know what? When I started working at Mozilla, I had that feeling of like literally everyone can see my code. The thing that I I learned very quickly and I had really good mentors was that all code is mutable. Like you can make mistakes and you can go back and fix it. And we shouldn't like chastise people for making mistakes. I think that's the the key thing. Yeah. One thing I've learned from having spent a lot of time in open source and then kind of moved into kind of a private company now is the amount of the good practices I did learn from being in open source. So like You could never put secrets or anything like that into a repository because like all hell would break loose, right? And that's actually like really good practice that we can do. So I when I moved from private into open source, I was like, I want to do it. And there's like, no, I can't do it. Everyone can see it. And it's like, I think that kind of solidified it. But also to me, I think I didn't have any of the really bad things. I probably have. Like there's probably awful code that I've written. And you know, I should be ashamed of it. But at the same time I can always go back and fix it. and I think that's the far more valuable side of things with that I think most people should get used to is that like making mistakes is bad, but you know, we're human, we're fallible, right? speaker-0 (28:57.752) going from private or companies with open source to closed source strategies and the other way around, vice versa. I I feel like there's there are learnings you can take in either direction. And also, of course, mistakes that you you should have dropped, but you know are are still there because you haven't sort of reevaluated where the the highest ROI is. Really from that perspective, it's all about learning from whatever experience you've had rather than not trying to think about it and just doing what's ever there. speaker-1 (29:25.098) Nine times out of ten, walking into any meeting, I felt like the dumbest person in the room all the time. So when I was at Mozilla, when I first joined, Brendan Eich was there. He was the CTO. He's the guy who created JavaScript. There were John Risick, the guy who created jQuery. He was working there. There is huge amounts of complexity that I never fully understood. and I learned a lot of it as I was going through it. So kind of how CSS is rendered because I had to learn it to understand how to do automation. And it's super complex. And half the time the the what the browser just goes is like, hey, operating system, this is our best effort. Can you clean it up, please? And the operating system goes, Yeah, sure. Let me do that, right? But then there's the JavaScript, but JavaScript has to be compiled down to bytecode to be able to be run properly. And I looked at that and I was like, there's a reason why I didn't do the compiler class in my computer science degree. I feel like I was working on like the the nice, easy parts where other people like their job was just security and performance while doing their like normal job. JavaScript as an example, you gotta make sure it's secure so no like people can't hack it by doing memory management when you shouldn't be able to do memory management and like things like that. And it's like, yeah, I I'll I'll just go work over there. It's the like it's kind of complex but it's nowhere near as complex as that part. speaker-0 (30:49.504) I like the the architecture design above everything and I gravitated towards the security area. I honestly feel like it's so much easier to r understand and write secure code or deal with algorithms than it is to do some memory dumps to figure out issues with with caching or, you know, performance or the hacks that you have to put in to make things more performant. It's just I don't know, that for me is just too disgusting to ever want to work in that area. Yeah. speaker-1 (31:14.478) Yeah, no, I I was the same, right? Like no one really cares about performance of test automation framework. Like they just want it to run as fast as possible. Yeah. But you're relying on like how JavaScript or those things work and it's like, Well, yeah, it's a bit that test is a bit slow. I don't think it's my code, it must be other people's speaker-0 (31:30.702) Okay. Well then, you know, maybe it's worth diving into a little bit your experience at at Browser Stack. speaker-1 (31:36.758) Yeah, so fortunately, like a lot of what Browser Stack does is built off a lot of technology in the open source space. People doing their tests in browsers and things like that. A lot of it's like Selenium or Playwright. And so we were able to commercialize it, make it more secure, things like that. And it builds all that out. And then we have interesting kind of proprietary systems in our data centers. around how we can have mobile phones because we use real devices for everything. Mobile phones in data centers, which then have other problems that like 'cause they've all got antennae and they're all trying to speak to the outside world and you're like trying to limit who it can speak to and things like that. So we have had to create interesting Faraday cages around certain things so that we can keep things on the right track. And again, you don't want like stuff to bleed through or have other problems. And the interesting stuff is not necessarily the software, it becomes more the hardware when it comes to that space. mobile phones, when they're always on, have a problem that like the batteries don't like to be charging all the time. As an example, right? We have data center engineers who sit in each of our data centers. And what we've had to build out is infrastructure equipment. Equivalent to like if you go to a sports venue, you've got these like thousands of people all trying to use the same Wi-Fi or or cell tower and things like that. And we know that like in football stadiums, they have their own infrastructure that kind of help spread the load. So we've had to build out a lot of tooling to kind of handle those situations. We have Mac minis for running Safari because there are rules and terms from Apple that Safari can only run on Apple hardware. And so we don't want to be breaking any rules around that because we can be sued on to speaker-0 (33:26.51) Do you have a lot of physical automation in play to be able to control the state of the physical devices like remotely through some sort of interface? speaker-1 (33:33.852) yeah, as much as possible. I think the last time I saw we had something like thirty thousand mobile devices spread across all our data centers across the world. The reason why we started noticing certain flakes happening on a certain machine is because memory started dying on that machine and that created a flaky test. And so we have similar problems on mobile now where we like have to build this out and then buy kind of the devices and then buy extra devices for when situations like that happen where we can go, Well, you know, we can't repair this one. It's it's dead now. Let's swap it in for a o a new one and do kind of a hot swap on those. The speaker-0 (34:12.002) The testing on mobile devices has always been really confusing for me. Like I don't understand why the experience is so like atrocious. I mean, obviously there's like money involved and some large organizations or companies that are make a lot of money from this. So controlling that makes sense. But I just don't understand why they don't come out with like full emulators offered from a proprietary standpoint that allow you to do this effectively, like you know, steal your customers basically. Like and Then they don't. It's like instead you see them have contracts with the hype other hyperscalers, and then you see the hyperscalers create ridiculous boxes or any company that does large-scale mobile device validation. And not necessarily for testing applications on the device, but sometimes using them to interact with the actual whatever it is, so any sort of robotic process automation. And they're they're quite complex. It's like a server rack size sort of thing where the mobile device sits in there and then the rest of it is sort of pneumatic you know, gadgets interact with it because you can't interact with the device from a software standpoint in any way. and now I think Google is taking away ADB on Android for non rooted devices. And you're like, well, that's it. Like everything is I need like a you know a pressure to actually activate, you know, mobile parts of the device. And then you're not just dealing with the the issues of the physical device that you're working with, but now the physical device that you've built to interact with that. I don't envy this situation. speaker-1 (35:39.244) No, I and it it's really fascinating. So like I've been in conversations with both people from Apple and Google going, like we need to solve this. We need a simple way, 'cause like in browsers it's a solved it's a relatively solved problem, right? Like we've got the W three C doing the web driver specifications and the W three C like have these rules of like you need two implementations of anything to become a standard. To get two implementations in kind of In the mobile space, other than browsers, you would have to have Google and Apple both agree. and they've both gone, Well, this is a moneymaker for us. We have our ways, that's your problem. And that's why like the people who work on Appium, the open source project, they are literally doing God's work to try solve all of these kind of problems. but it's never gonna happen. So the other side of it is there's the guy who created Selenium. his name's Jason Huggins. He created another company called Tapster. and Tapster is doing what you said, which is testing mobile devices, but he's built like got three D printers and Arduinos to figure out where certain things should go and then he's got a pen driving the mobile device for testing. speaker-0 (36:54.966) bit of a sobering story because like no matter how bad you think writing tests are and having no correct way of doing it in the browser, there is another whole set of applications, a platform out there where, you know, good luck. speaker-1 (37:08.354) Yeah. if people want to see other horror stories, I highly recommend looking at the differences between screen readers between operating systems. That's another nightmare in itself. So Linux will do one way, Windows will do another way, Mac OS will do another way, iOS will do another way, Android will do another way. speaker-0 (37:28.75) I'm very thankful for anyone who does work in the accessibility space. I as someone who doesn't super need it myself, but does maintain some open source software, when we have users come to us and be like, Hey, this is broken, I'll I can make a pull request to fix it. And then I learn about a a part of the web that I had no desire to and it's horrific. Like it is not it is not nice stuff to overcome these things. And I think part of the challenge is, as you mentioned it earlier on, we just kept stacking stuff on top of history. in effort not to break things, et cetera. And it makes me wonder, do you think that we would have ended up in a much different position than we are now if we, you know, started to invent the tech today? speaker-1 (38:11.032) So I wish I could say yes, but I think I think we would do all the same mistakes, right? Our systems are just improving exponentially. That's a a problem for tomorrow, from testing to performance to security. We now have an interesting problem of like, do we solve quantum in the next 10 years? How's that going to handle for security? Who knows? Right? Like, is it gonna break, you know, crypto to kind of banks, right? It's gonna be everything in between. and so there's interesting spaces there. I think we've lived through a lot of the golden age of things, to be brutally honest. I think kind of we've got this really interesting space now where like from say 2000 to 2020, loads of people were doing open source and open source allowed us to push new ideas into wonderful spaces. And then in the last six years, that has contracted so much. I'm not sure if we had start if we start now, would that be a case? If we started twenty ten, maybe we could have done something different. But I think kind of then I think people always go, 'Cause like humans are inherently lazy, right? Like I'll I'll I'll do enough. And then carry on. speaker-0 (39:22.712) You're definitely spot on there that some of the challenges that led us to the current state, you know, unless you eliminate those, we still likely end up in a in a similarly challenging spot. or maybe we don't have any of the current problems, but we just how long it was. speaker-1 (39:39.736) Yeah. Yeah, definitely. speaker-0 (39:42.082) Yeah. browsers are in this weird space where they're almost really great operating systems. They're pretty much functioning as that mechanism. And like on mobile devices, they are almost fully that. We have very little control over the OS there from a user standpoint. And so if we look at applications as like mobile experience, websites, then it's almost entirely websites and the browser experience. I don't know. Like I'm of I'm of two minds here. I guess my perspective is I would love it if the all the sandboxing effort, everything for interacting with websites really became a full experience. but then I still see what application developers are doing if I look at, you know, you look at the chat tools we have or email clients that are available, and these are like literally the the worst. speaker-1 (40:33.41) Yeah. No, so like I I was a big advocate for kind of the idea of a PWA. I was always a big advocate because I thought it was really interesting. And like part of my work through W3C, I get to see some of the interesting bits and pieces that are happening at the web level. Google and Microsoft had this project called Project Fugu. And so it's based off this idea of the sushi fugu blowfish. But you've got to make sure you get the best part of the blowfish. Otherwise you get poisoned. and so that's why it was called Fugu, because they were trying to take some of the concepts of mobile and put it into the web. So Bluetooth, address book, USB, all of that. some of that sounds awesome. Part of that also sounds incredibly scary, right? speaker-0 (41:19.606) the project name. Like I I I you know, as soon as you call it that I I feel like you you already know where speaker-1 (41:23.478) It's going. Yeah, exactly. We're all taught do not stick random USBs into your kind of computer. imagine if you stuck a random USB and a your browser goes, I've got a new USB. let me see what's inside it. And then suddenly like all your banking information and everything is stolen, right? Like it there's some because browsers try to hide all of that stuff in encrypted parts like the BF cache, they try encrypt a lot of that stuff. So no one can just randomly steal the stuff if they're on the computer. and so Project Fuga's trying to do that. I think there are good parts to it and there are bad parts to it. So there is an interesting space there that I think we should try make the web better on mobile. But then we come back to the original problem that we were talking about earlier of like how do you test on mobile devices? An interesting thing is that I can s shrink my browser on my desktop to that right size, but it doesn't mean it's going to render that way on a mobile device because the way rendering w works on my Mac or my Windows machine is very different to a mobile device, by design, right? But people go, that's this is good enough and you're like, it it it is, but it's not at least admit there's going to be differences and if you're okay with those differences, then fine. speaker-0 (42:36.568) There there are some of them are not trivial. Like this is just a warning for anyone doing this today. Like, even if you change the size, even if you go to the developer tools and click the, you know, convert to responsive mode. And even if you change out of responsive mode and select specific like phone sizes, mobile device sizes to use, you're still not getting the gestures to work correctly. And the biggest ones that I've seen break are like media query sizes, or like VH or like partial heights there, because the mobile apps will do. ridiculous things to the size of the viewport when you change the size the application on mobile because you can support that or the keyboard pops up and then the height that's remaining changes. It is it is not the same. I am I like that there is like very it's very difficult to debug without plugging your phone into a computer somewhere and like exporting some ADB or getting the inspector to run specifically for the application in order to actually see what's going on. speaker-1 (43:34.518) Yeah, it's it's fascinating. I've been telling people that for years and most people go, No, David, like just chill, you're okay and they're like, No, it is very different and it's nice and refreshing to have someone go, David knows what he's talking about speaker-0 (43:47.522) Yeah. and it's not just on the display side. There's a lot of things with like how the security keys for I would say pass keys, but FIDO two, web auth and stuff, will say that it works the same everywhere, but on device. No, like I can use my phone as the pass key, but if I'm run a mobile web s like a website or an application and try to use the pass key like on device on it, it doesn't work. or like on Bluetooth it works, but if you try to use a special connection where you scan a QR code and does it doesn't go. And it's not it's not just the implementation of Web Auth N that some application developer turned off because you can specify for Web Auth N like only blue Bluetooth is allowed, or only like removable like physical hardware is allowed, like a UB key or whatever or whatever the other ones are called. So assuming they didn't do any of that, which could have caused a problem, there's still additional other issues which it just don't work. And that's like a very small part of the story. So speaker-1 (44:44.588) Yeah, definitely. Hundred percent agree with you. speaker-0 (44:47.63) Yeah. So, you know, that that that's the future there. I I think maybe Chrome OS was this, to try to expand the browser into being the full operating system in a way. obviously it took off for a a group of people there. I certainly always like the idea of a thin client to do the right thing. And I feel like one of the huge challenges with browser stuff and when we had Sam go to on here from Google in a previous episode, he was talking a lot about the security challenges, especially in the browser world. And it was that was that was interesting. So anyone who wants to go more in the security space, there is that there is that episode which we'll link in the description. But I will say that one of the huge challenges was like sort of jumping over this gap of the operating system, because the the browser is so limited because it doesn't have full control. It's not like a super user running on device. And yeah, but on the flip side, as you mentioned, there's the whole problem there of because you bridge that gap, now the issues go back the other way. And Where the applications running in the browser would have been insulated from things that connect to the operating system where sort of increasing that attack vector basically. speaker-1 (45:55.586) And I think the easy way for kind of non security people is the wherever kind of two systems meet, that's where the security problems will be there. And it's not a if they will be there, they will be there. And speaker-0 (46:08.238) And those are where your test break too, because it's where all your assumptions lie, right? Like the expectations of whatever that other interface goes. I feel like one of the questions that always come up is should you include in your testing the actual third party service or library, not just their interface, but the actual implementation there. And I've been going back and forth over this on a lot of for a lot of years, like, okay, you would want to respect the tests of that third party that you trust for the open source that they make. That it does the right thing that it expects. And if it changes, you can see those tests that change. But on the flip side, people make non-semantic version changes, call it semantic version changes that will break your application and your usage. the interface doesn't match what actually the implementation is. here's an example that happens a lot for us. We found out that there's some great libraries in in Python which like to complain if the response format from an API doesn't match the classes that it's it's defined. And I'm just like, don't throw on that. Like, there's no reason to throw an error because it doesn't match on a field that you don't even care about. But a lot of libraries do this. And there is this fear that if you don't test it, something can change and break. But if you do test, you're wasting a lot of resources, potentially we talked about, but also time and effort. And you can't narrow your test down just to the part of the code that you you want to think about. So what's the what's the advice here? speaker-1 (47:29.462) I think the the advice really is back to that. Let's write one test and then go from there. Right? Right. So small, medium, large tests. I only test what like the where I think the risk is. And I go from there. I don't I don't need to test everything because you know, if there's a security problem or kind of performance or any other bugs that happen in that system, The next time I update, I'll likely be fixed, right? And then do I want to maintain all of those tests in the first place? So let's just focus on one test. and that one test is the thing I care about for that area, right? And it could be purely down to one API call or whatever, right? But I think kind of that mentality of just focusing your where the value is, right? You don't want to have to go and update twenty tests. yes, you can do it a lot faster nowadays with an LLM, but then like you've got to make sure that it's still doing the right thing afterwards, right? Where if you've got one test and an LM makes a change, you can review it quickly and then move on. A lot of the problems that I've seen are not necessarily with the third party frameworks or whatever people are using. It's how people are using those frameworks. I was working on Browser Stack's test management tool. It uses React. When we were developing it and we got to this one point where UI updates were super, super slow. We had loads of component tests, we had loads of unit tests, we had everything, and we had a bunch of end-to-end tests. none of them were flaky. we noticed that they were taking a long time. And what it was is that like whenever data would update this one component, it would cause the entire page to refresh. And the tests would be like, okay, cool, cool. I I've handled this, so there's no failure. We're good. But that's a performance problem, right? And it's not a React problem. It's not a whatever. It was a like our developers. When we started putting all of these things together, because everyone was working on the individual parts and then one person was like, I'll smush them all together. Hooray. speaker-0 (49:41.134) I feel like architectures and frameworks which cause the wrong thing to happen in the long term, which are very difficult to identify, can be the wrong thing for your business or your technology stack, et cetera, and take years to sort of identify by an expert who's been in the in that area, that there is a fundamental problem with that technology that you'll keep running into. I feel like that's another episode, and I think I may need a specific guest to come on to talk to that. Okay, then then you know, so there's there's one last area that I I definitely want to get into and I haven't read the the spec or wherever the proposal is at, but I think you brought up LLMs a couple of times and I try to avoid anything related to AI because it's it can be a rat hole or everyone's worst nightmare, or I'm sure there's some listener who's jumping up for joy right now because I'm gonna bring this up. But it does seem like the ability to interact cross-device or cross-process in a way that can support LLM interactions. Is something that's gone avoided for long for too long. I I think a lot of them when you want to interact with a particular web page or application on the mobile device, it's just not possible today. And I have hopes in the future it will be. But in the browsers, I saw this thing called Web C P and when I hear that it makes me think the right things are gonna happen. And I have no love for LLMs, but I I do see the the value there. speaker-1 (51:01.262) Is an interesting potential overlap between testing and how people do it. So WebMCP is this set of APIs that allow you to build an MCP into a web page. And so MCPs are these tools that you can LLMs know how to interact with. flights between this place and this place, right? And you just tell your LM, I need you to go to Kayak, do this. You load c it loads kayak in a browser, it goes, there's a tool. I don't need to try to figure out the UI. I can just speak to the tool. The tool will then figure out the UI for me. Right. And that's where MCPs are really good is that you can have these set of tools and they can do very specific things really well because they've been programmed that like to do small succinct things. The thing I like about Web MCP is more of a mindset thing because you don't want to have like I'm sure we've all tried to install a like say GitHub MCP and you do it and it'll go, There are now ten tools available, right? If you create a web MC like web MCP and you load it up and your LM finds it and goes, there's an MCP, let me load the tools, there are three hundred and two. as a random number, the value of it disappears really quickly. And so to me, it's like, well, you would only have a tool for where the high value traffic should be going. The the high risk, the things that you want people to do correctly, first time correct and so you could do that. I've been playing with that. I built a a tool. It uses WebNN, so the web neural network APIs built into Chromium to then generate an MCP. And so that it can go, hey, these are the high traffic areas. it's got a whole bunch of heuristics around what it considers valuable. So like if it sees a header, it'll go, that's not valuable. If it sees a a link, it'll go, that is valuable, right? And tries to balance it all out. speaker-0 (53:12.63) On mobile, there's things like called intents, at least on Android, that allow you to switch between applications. But most apps don't expose them. And even if they did, other apps don't actually try to utilize them. But I can see a lot of opportunities there with wanting to interact with the applications or the value that your phone is providing without going through the first class UI. And the browser side, I worry what's happening is where bypassing the value of having a first class API. In the first like you say the GitHub MCP, you know, wasn't great. And then people spin up their own MCPs on top of GitHub for a very specific value that they want to extract from the API. That's both good and bad. being able to directly define what your value is and what the use cases that you will need it for and optimize it, you know, that's good rather than having, you know, a hundred thousand different tools that are exposed. And then you have this whole problem of how does the LLM even make sure it's picking the right tool with the right parameters to call it in the first place. Yeah. But at least it would force GitHub to think about how to expose that data in the right way and not add in an extra complexity of a whole user session with an expensive application like a browser in order to just have an LM on your machine interact with it. And not to mention, if you don't have an LLM on your machine because you have a GPU that was more than one year old at this point, you're gonna be running you could now have to run an expensive browser process in a virtual machine somewhere and interact with the MCP that's created in that way. But on the flip side, I get it, right? Not every application developer understands how to build a first class API that can be consumed and then expose it with the right documentation, et cetera. And if you're doing anything interesting, that maybe that website is just a static page. And being able to effectively pull out the information anyway is going to be much cheaper for everyone than to force it to parse the actual DOM tree that is. is available at that moment in order to skip the HTML and get out the parts. Because as you mentioned, it's not just the number of tokens. There's a lot of formatting there for human benefit that offers zero for actual value. speaker-1 (55:17.566) The one thing to remember is Web C P is relying on the web developer of that site to create the right tools in the right space. and if they like this is where I think the interesting part from the testing point of view is like if they don't do that part correctly, they're probably not doing the testing part. correctly either. They just don't realize it. And I and it's and it's not a shade on anyone, right? Because like both of those things are incredibly hard, right? Like there are people who focus their whole career on user experience, but this is a different type of user experience, right? so it's not a natural thing to do. It's the same with testing. Like lots of people focus on testing, not developing, because like that's where their skill sets are. speaker-0 (55:59.992) Yeah, actually in the episode we did on the Wikimedia Foundation and open source Wikipedia tools, our the guest Morel actually really talked to us about the benefit of thinking about what how your users are actually utilizing your application. And if they want to use it through this way, they're still your user, you're still getting the value out. Just the value isn't the UI structure. And so don't don't effectively curtail or limit a potential user because they have a unique way of approaching it. I think it was a pretty interesting insight that I hadn't can I hadn't considered. And I think it's a much smarter and mature approach than saying we don't want LLMs to use our website. It's you have to admit you're never going to stop that. Whether or not it's right or not, there are some users who want to do that. And the question is do you want a user that is trying to use or you can get the value out of your your site or application or tell them to go use a competitor who is more open to I I could have a whole conversation about this. Like we ran into tons of trouble with our documentation where like links in Markdown could be relative, but do you want the links in the LMs.txt to be absolute and like what's the actual domain and kind of piece those things together, especially if you have some sort of base path. But I'll I'll I'll leave that for for another episode potentially. And maybe now is a good time to move over to picks. So David, what did you bring for the audience today? speaker-1 (57:19.798) so I don't have it with me right now, but the thing that I wanted to talk about was that I'm a big into retro gaming. I am an elder millennial, so I still I still like to play kind of these older games where I kind of have so I play a lot of Pokemon. That's one of my favorites. speaker-0 (57:42.978) I didn't expect you to say that onto the list of all the retro kicks. speaker-1 (57:46.894) so like I play a lot of Pokemon. I've been a Pokemon fan. and so I have there's a company based in China called Amber Nick and they ship a like device that's shaped like the old retro devices. So if you're like want a Nintendo, like the old Nintendo, I've got a a Lego version here on so I'll show that to the screen. it's a Lego version of the original Game Boy because I really like it. but there's also the Nintendo SP, which was the one of the flip ones. so I I like collecting those and then collecting all the ROMs to play it. So kind of but I play everything from Mega Man to Pokemon to kind of full metal solid, like all of these things. Cause like one of the things I liked about gaming back say 20 years ago is that like there wasn't you couldn't just use the internet to download updates. Right. So games had to be fully shaped and correct. And if there were bugs, there were bugs, right? and that was it, right? And you just kind of had these things that you could just play. and they were designed for you to like play for hours and hours or kind of very quickly, right? Where I I'm not a big fan of kind of microtransactions, things like that are that are in games nowadays. So speaker-0 (59:12.92) You reminded me of this great video about the manufacturing of a of a old retro Prince of Persia game and I there's quite some interesting lore there about what you have to do when you only have so many bits available of of memory and processing power to actually generate things. And the genius that is of the the inverse fast square root in in Quake, right? No one would would bother trying to identify what that is and understand how that even executes and maybe we'll And if you don't know what those things are, I'm sure I'll have to find a good YouTube video on on on both of those. I had this conversation actually not too long ago about like what classifies something as a retro game. And I feel like is it the eight bit shift like before after eight bit it's no longer retro? Or is there like a particular year for you or type of game that falls into this category? speaker-1 (01:00:01.726) kind work how the but it it's kind of like anything from Nintendo advance backwards. Okay in that range, right? my kids will tell you kind of anything before twenty twenty. it's probably retro. we have arguments, but it's it's that type of thing. and it's speaker-0 (01:00:08.952) Yeah. speaker-0 (01:00:23.638) They're allowed to be wrong. They're just kids. speaker-1 (01:00:25.676) Yeah, exactly. and so like but it's for me it's like Game Boy Advance and it's those games like I say, it's that point where you couldn't just download fixes. so kind of remember being like just spending hours as a teenager just like tr grinding, where now it's like, you know, as an adult I've got responsibilities and when I do play I can like do these things and do it quickly, and still come back to it later, with not having felt like I've Wasted in the whole evening 'cause like I'm just old and tired. speaker-0 (01:00:58.488) play Pokemon go by by Niantic. They had a previous iteration called Ingress, which their whole point was to capture geolocation data and pictures from around the world before they came up with Pokemon. And then they just released recently that they had been doing it with Pokemon. So basically anyone who's playing the game has been funneling the LLM training machine in order based off of where they go and their phone orientation and any pictures that they've taken, which is I mean I I knew they had some genius strategy, but it's quite may lose some of the Pokemon fandom. speaker-1 (01:01:34.958) But that that's how Google Maps got all that information 'cause originally Niantic was an internal company within Google and then they bought themselves out of Google and then carried on. But they we still had Google as a customer. speaker-0 (01:01:49.036) You know, honestly, for interesting ways of collecting data for training LLM agents as as it relates to sort of the whole software engineering pipeline, I'd love to get a guest on the show. So if anyone listening has someone who they think would recommend for this particular segment, I I'd be happy to connect with them. I think I have to share my pick, which I don't know if it's gonna be as riveting, but it's this paper that was released. by my university Cornell called the bullshit the corporate bullshit receptivity scale. And it basically is a pure metric for calculating how BS a particular statement is. And I mean it's sort of interesting because it's what I've been able to figure out is like a lot of the statements which are BS, they sound very similar to non-BS statements, except like instead of at the end of the statement where you like inter introduce the actual point or a word that sort of completes the thought, you throw in something else that just loops back on itself. And it's it's the really interesting thing is that there's a people can very readily detect it, that it's happening. And it causes a lot of negative sort of results in organizations. They people can will say that organizations aren't great or that it's negative to be in. however, being a exposed to it actually causes lower analytical thinking and insight and believing that BS was said, and you believe the person to have been more intelligent in a way, for a small portion of the population. And I this sort of scares me that like some people actually like hearing BS and they will spread it in an organization even if they know it's bad. And I think it's all it's sort of very connected to how I see the the narcissistic language that comes out of LLMs today. which sort of justifies for me why so many people think that LMs are so great in in what their what they say or or their capabilities, which is another really scary thought. speaker-1 (01:03:54.294) I I'm gonna go look that up. I think yeah, I think you and I have very similar ideas around LLMs, th by the sounds of it. they I hate how much of a people pleaser LLMs are. And so I think it's it's very interesting. And it if you're enjoying good books, and you like to understand how people think, you should read the book, Thinking, Fast and Slow. Go read that. Then now like if you read it before. reread it and think look at how L L Ms work and it's very interesting. speaker-0 (01:04:28.174) Daniel Kenneman. yeah, no, no, I I I agree. it's been a while since I read that, so I'll maybe have to put that back on my reading list. yeah, we'll get that in the show notes. Honestly, this has been a great discussion, David. Thank you for joining us today and talking us through the browser stuff and testing. speaker-1 (01:04:43.01) Well, thank you so much for having me. I've really enjoyed this. speaker-0 (01:04:45.934) It's been it's been great. And hopefully everyone liked this episode and they'll be back again next week.