From Chrome DevTools to Loop Engineering: Addy Osmani on Understanding How Things Work
The Pragmatic EngineerAddy Osmani spent 14 years at Google, most of it on Chrome, where the work included Chrome DevTools and Core Web Vitals. The role grew from developer relations engineer to director of engineering, and the last stretch was spent in Cloud AI and Gemini. In this conversation on The Pragmatic Engineer, Osmani traces that path back to a teenager in rural Ireland who built a web browser from scratch. The conversation then turns to how AI agents are changing software engineering.
One idea runs through the whole discussion: Osmani values understanding how things work, one layer beneath the surface. That idea shapes how Osmani describes the early browser project, the DevTools work, and the current worries about "cognitive surrender" when working with agents.
How agents have changed Osmani's day-to-day work
Asked how their workflow has changed recently, Osmani says agents have let them "take the improbable and turn it into the possible" on a daily basis, and that they now manage a lot of their life with agents. The example comes from the week of the recording. Osmani was giving a closing keynote at the AI Engineer World's Fair in San Francisco and wanted two things: to avoid repeating points other speakers had already covered in depth, and to connect the talk to earlier sessions.
In the past, Osmani says, you could only hope the sessions had been published online, or skim the abstracts. This time Osmani sent out a batch of agents to collect everything they could about the previous days' talks, including abstracts and social media posts. Osmani then used the results to reshape the talk they had already planned. Osmani calls this empowering, because the task would have taken a very long time, if it had been possible at all.
The host notes that many people describe the current moment as both fun and chaotic. Osmani agrees. People with many ideas used to be limited by time, and agents have in some ways "unchained" them. When people ask what engineers do with the time agents free up, Osmani's answer is that they do more work. They add that people wouldn't fill that time with more work unless they enjoyed it, so for them it is "fun thriving in that chaos."
A teenager's web browser and a national science prize
Osmani grew up in rural Ireland in the dial-up era and first used a computer at about eight or nine, when their father bought the family's first desktop. The question Osmani kept asking was how any of it worked: you type something into an address bar and text, photos, and videos appear. Osmani moved from building websites to programming. The first language was Pascal, with Borland's tools, and C++ came early as well.
The project that made their name started with slow downloads. Osmani recalls that a song could take hours and a music video a night or even two days. Download managers of the time sped things up by opening multiple connections to a server and requesting the file in chunks, which worked when the server supported it. Osmani wondered whether anyone had applied the same technique to loading web pages.
To test that, Osmani first had to build a browser. At 15 or 16 they started reading the HTML, CSS, and JavaScript specifications. The hardest part, Osmani recalls, was not parsing documents or loading images. It was handling pages that ignored the specs, because real browsers still render them. Osmani still has great respect for how much "weird crap" developers throw at browsers. As a form of therapy for anyone who feels bad about their own code, Osmani suggests opening DevTools and browsing the web for ten minutes to see how much goes wrong while sites still work.
JavaScript support was very difficult. Osmani then decided a browser wasn't complete without applets, Flash, and embedded Windows Media Player content, and added those too. Only after that could Osmani test the speed-up idea. It worked, and at the time many servers supported it, so pages loaded somewhat faster. The motivation was personal. Before they had a decent connection, Osmani would wear cargo pants with lots of pockets every weekend, fill them with floppy disks, and walk to the local library, which had a slightly faster connection.
Osmani entered the project in a national science competition. That was the first year the competition took computing seriously. Osmani expected to show some people some cool stuff and go home, but won the overall prize in a live televised finale. The following Sunday the phone kept ringing, with calls from the Wall Street Journal and CNN. The host compares it to going viral before Twitter existed.
Osmani's main lesson from the experience is that they still didn't fully understand what they had built. Getting an app to work on your own machine doesn't mean you understand the layers underneath. Osmani says this started a lifelong need to understand how things work, and jokingly calls themselves "a one-trick pony" in that respect.
Peeling back the onion
Osmani then worked at startups and at AOL. On the first day at AOL, a manager told Osmani to log into the browser and help the team. The AOL browser app immediately asked for a credit card before any debugging could start. The manager was busy, so Osmani entered it. Osmani calls it "a very different time."
Osmani describes computing as an onion with interesting things under every layer. With browsers, the layers include the network, compositing, and the JavaScript engine, and below those are chips, memory, and the GPU. In Osmani's view, the more of these foundations you understand, the better you can build for constrained environments. If you know why a slower phone is slow, you know how to design differently for it. That is why Osmani encourages people to learn how things work.
jQuery, TodoMVC, and the path to Speedometer
jQuery was Osmani's first contribution to a large community open-source project. For a few years it was the most widely used library on the web. Osmani credits its creator, John Resig, with building a welcoming environment where people could become better contributors. Osmani started with issue triage, then wrote blog posts and contributed code. Because jQuery was used in so many ways, people often argued hard over whether a feature belonged in core or in a plugin. Osmani learned how to work with a community while still holding a line on decisions that affect long-term maintainability.
TodoMVC began with Osmani's own confusion. As frameworks such as Angular, Backbone, YUI, and Ext JS appeared, every landing page promised easier app development, and it was hard to see how they actually differed. Osmani built the same to-do application in each framework and standardized its functionality so people could compare architecture, syntax, and each framework's approach to components and UI. The app was meant to be simple enough for anyone to understand, but interactive enough to exercise features like state management and routing.
Osmani had no expectations for it, but it quickly gained thousands of GitHub stars. Framework authors started sending pull requests to add their frameworks, and Osmani met some of their first close open-source friends that way, including Sindre Sorhus, who went on to write a large number of Node modules. For years, framework tutorials used a TodoMVC app as their baseline, and Osmani says they still saw labs demonstrating features with TodoMVC apps as recently as last year.
While the project was growing, the Safari and WebKit team at Apple contacted Osmani about building a browser benchmark for responsiveness. They meant how quickly a browser responds to clicks and taps, not responsive layout for mobile. The collaboration became Speedometer. Osmani says Speedometer has become the main web application responsiveness benchmark for all browsers, and that browser vendors now maintain it together and update it as new frameworks and architectural patterns appear. Osmani sees it as carrying TodoMVC's legacy forward.
Joining Google in the early Chrome era
Osmani first thought about working at Google while visiting their wife's parents in the Midwest, when a documentary about early Google engineers came on TV. Years later, after Osmani had been publishing free educational material for front-end and JavaScript developers, Google reached out to interview them for a DevRel and builder role. Google wanted help with some tooling and with evangelism in the tech community. Osmani ended up on the Chrome team.
Osmani joined around 2012–2013. At the time, Chrome was focused on encouraging developers to push the platform and expose its gaps so Chrome could build better APIs. One example was the Chrome Experiments site, which showcased work such as WebGL demos and inspired some people to learn about shaders. Front-end tooling was still immature. Meta-frameworks like Next.js didn't exist, JavaScript modules weren't standardized across browsers, and people used AMD, UMD, and CommonJS alongside build tools like Grunt. The host adds that debugging meant using Firebug in Firefox and hoping things worked in Internet Explorer, which for a while had almost no good debugging tools.
Google's contribution to this tooling included Yeoman, a scaffolding tool with a CLI wizard. You chose a UI library, a testing library, and a deployment target, and it generated a starting point. Osmani doesn't call it the first meta-framework, but describes it as an attempt to bring some organization to how projects start. These ideas are standard now, Osmani notes, but they didn't exist then, and Osmani wouldn't be surprised if modules from that era still run inside today's tools. Osmani also recalls the constant churn from Grunt to Gulp to Webpack, Rollup, and Vite, and is glad things now seem to have stabilized somewhat.
Inside Chrome DevTools
Osmani credits Pavel Feldman, the DevTools tech lead, with much of the product's early direction. Chrome was deciding how to set its developer tooling apart from the WebKit Inspector, and Feldman's team took a developer-centric, ecosystem-focused approach. They worked closely with developer advocates like Paul Irish and later Paul Bakaus (whom Osmani mentions as now known for the "impeccable" design skill). Those people, Osmani included, were web developers who built things on the side. They brought their friction points to the DevTools team. Sometimes the tools couldn't be built yet because the browser lacked the underlying instrumentation.
Osmani highlights the performance panel, where you can hit record, interact with the page, and get a flame graph with detailed tracing of where time is spent. Memory is the counterexample. Osmani suspects few developers understand memory management, which makes memory problems even harder to debug, and says the state of the art in memory debugging hasn't advanced much over the years.
Osmani describes several eras the DevTools team adapted to:
Frameworks and libraries. On an interactive page built from many libraries, which code do you care about: the framework, third-party plugins and components, or your own code? DevTools invested in source maps so developers could trace behavior back to the original files, even through the long toolchains large sites use. Osmani thinks people underestimate how many tools those sites run for a single task. DevTools also added ways to ignore specific code, so you could say, in Osmani's example, don't flag issues inside the React library, but do flag issues in the React code I wrote.
Mobile. At first there were no tools for testing viewport widths, tap-target sizes, or sensors. DevTools added a device mode that previews a site at different viewport sizes, roughly as it would appear on an iPhone or Pixel. Osmani acknowledges that testing on real devices is still best, but says a quick check of whether you're heading in the right direction was valuable.
Progressive web apps. As the web tried to compete with native apps through offline caching, push notifications, and background sync, each of these features needed debugging support. That led to the Application panel for inspecting service workers, caches, and related features.
The host says Google never seemed strong at building IDEs, except inside Chrome, where DevTools offers the kind of breakpoints, conditional breakpoints, and performance and memory tooling they were used to in Visual Studio. Osmani says whether DevTools is or should become an IDE was a constant debate on the team. The team settled on meeting developers where they already work, because everyone will keep a favorite editor, or now, a favorite control plane for their agents.
The latest era, under tech lead Yang Gao, has been about AI. Osmani describes two goals: helping people make sense of the huge amount of data the browser produces, and letting developers connect agents to Chrome and DevTools to automate workflows. On the first, Osmani recalls that when helping large sites with performance, they could spend half a day reading traces before writing a single fix. LLMs, Osmani says, can now work through huge traces quickly and get to actionable fixes.
Core Web Vitals: turning how a page feels into numbers
The host notes that metrics like Largest Contentful Paint (LCP), Cumulative Layout Shift (CLS), First Input Delay (FID), and Interaction to Next Paint (INP) link how a page feels to measurable values, and asks how the team came up with them. Osmani says the Chrome team has always relied on user experience research and revisited it whenever the web changed.
Performance used to be judged by vague ideas like whether a page had "loaded" or was "ready." Does ready mean you can see it, or that clicking does something? The team broke the user journey into moments: Is anything happening, like a header or spinner? Is something useful there, like a hero image, video, or the main content? Is it usable? Each moment can map to a metric. A hero image might correspond to LCP, but Osmani stresses this doesn't generalize. On some pages, the article text is what matters.
For interactivity, Osmani describes clicking "add to cart" while shopping and getting nothing. The host explains what's usually happening: the JavaScript or event handler hasn't loaded yet. Users then tap repeatedly, and once the handler attaches, the item may be added two or three times. Osmani compares it to a pedestrian pressing a crosswalk button over and over.
Layout shift came from something the team saw as clearly harmful. Sites monetized heavily with banners and modals. Even if businesses need revenue, Osmani says, that shouldn't ruin the experience, and you shouldn't be reading an article when an ad finally loads and pushes everything down. CLS reflects the idea that the page should stay stable.
Because the web is so varied, the team ran many experiments with different metric definitions. They worked with the standards community and developers to check whether the metrics matched how developers themselves judged their pages. Osmani says some companies had already thought through what matters to users starting from a blank white screen, and many had not. For the latter, Core Web Vitals started a more nuanced conversation about making sure users can complete their key action quickly.
Google's engineering culture
Osmani worked on both developer-facing and consumer-facing projects, such as Chrome performance. When a change reaches billions of users, Osmani says, engineering velocity and experimentation work very differently than at a fast-moving startup. A change may need several solutions for different markets, and success has to be measured while 20 or even 100 other experiments are running at once. Osmani describes Chrome's A/B testing process as stable and rigorous.
Osmani also says many people at Google cared about developer goodwill, even if that didn't always show from outside, and admits that in a company that large, not every group can talk to every other group. Asked to name their main contribution to Chrome's culture, Osmani says it was "meeting developers where they are at": accepting that people will use whatever tech they want, and helping them succeed on the platform. In practice that meant working with framework teams and taking their feedback on which APIs to build, instead of guessing what the community needed.
Osmani also valued how Google shared learnings across teams, pointing to the book Software Engineering at Google and similar internal write-ups. Osmani liked comparing their team's views on testing or user experience with, for example, YouTube's. The Chrome team did help YouTube improve its Core Web Vitals, which Osmani describes as very nuanced, very educational, and very time-consuming. The host notes that this kind of cross-org collaboration isn't a given at large companies, where orgs can be focused on their own goals. Osmani attributes it to people on both sides with enough agency to make it happen and who saw mutual value.
From DevRel to director
Osmani joined Google in the UK at level 4, a mid-level engineer, as a developer relations engineer. Over time they were promoted to L5 and L6 and became a manager within DevRel. About five or six years in, Osmani felt that, despite loving developer relations, they were "a builder at heart." Having been an engineer before Google, Osmani moved back to software engineering as an engineering manager, a role that allowed both engineering work and managing teams.
The team started small and grew to about 45–50 people at one point. Osmani's goal, also described in the book Leading Effective Engineering Teams, was an organization that mostly runs itself, so the leader only needs to check in occasionally and correct course, and can focus on the next important problems. For Osmani, that meant thinking through what improving model quality meant for developers, developer tooling, benchmarks, and partnerships with third parties, then bringing those ideas back to help the team make DevTools and Chrome more useful for agents.
Osmani stresses that getting there takes a lot of work: building a team structure with managers of managers, and handling time zones and coordination across a global team. Osmani then helped other managers make similar space for themselves. Osmani went from L6 to L7 to director, which is L8 at Google. For a long time the focus was the work itself, but once the org was healthy, Osmani wanted the next challenge and aimed for director for a while. As with any promotion, Osmani notes, you have to do the job before you get it, and what got you to your current level won't get you to the next one.
What changes at the director level
The host calls director the first executive level at many companies and asks what changes. Osmani recalls that when they were coming up, the director was the first executive contact. Directors held teams to their annual and quarterly goals, sponsored large programs and new projects, and ran regular reviews when things went off track. The role carries more accountability, and Osmani says you can't succeed in it if you let go of the details. A self-running org doesn't mean letting go. It means having systems that bring information, decisions, and blockers to you quickly. A central part of the job is helping people who don't always see how their technical work connects to business goals to see that link clearly.
Osmani held the director role while working on Gemini and Cloud AI. They were responsible for one of the organization's top goals for the year and had to report on it every week or two. Osmani also highlights a recent change: as models and tools improved, many directors, VPs, and SVPs started building things themselves. Every week there were conversations about what people built on the weekend, which models they were trying, and where they hit friction. Osmani says that didn't happen before, because executives usually focused on big company problems.
Cognitive debt and cognitive surrender
Osmani separates two ideas. Cognitive debt is the gradual loss of your memory and understanding of the problems you work on as you rely more on AI. Cognitive surrender follows from it: you accept whatever the AI says, so "its answer becomes your answer," and critical thinking and problem-solving fall away. Osmani is enthusiastic about harness engineering, loop engineering, and software factories. Even so, Osmani wants engineers to understand enough to fix things when they break, instead of hoping the agent figures it out.
The host notes that about a year ago Osmani advised reading an agent's reasoning and reviewing its code. Now agents are much faster and people run several at once. Osmani says their view has changed. A year ago you might see one "thinking" message, expand it, and follow the trajectory at a readable pace. Now, with Claude Code or Codex, 20 or 30 subagents may have run, and Osmani won't click through 30 trajectories. Instead, Osmani does two things.
First, Osmani reads the final summary of decisions from start to finish. If there isn't one, they ask for it, while staying aware that a model can make up decisions. The host adds that this isn't deliberate and can come from context-window limits.
Second, Osmani relies on what they call mutual amplification: working so that both the agent and the engineer get better every day. That can be as simple as asking the agent to log what it learned in a session, the decisions made, the friction encountered, and anything unusual about the approach. Osmani calls this intentionality. As long as you stay a little curious about how things work underneath, Osmani believes you can keep some of your understanding while working with models.
Loop engineering and software factories
The host brings up "loop engineering," which people including Peter Steinberger and Boris Cherny have written about, and asks what loops really are. Osmani sees loops as part of the move toward software factories. Instead of prompting toward an outcome yourself, you build a system that does the prompting, produces the outcome, and handles testing and verification. Osmani calls this the next step in a "rising tide of abstractions."
Osmani acknowledges the obvious question experienced engineers will ask: what about quality? You have to decide deliberately where humans stay in the loop. For example, the system can flag when a change touches a critical part of the codebase and needs human review. Letting loops build everything without guardrails on blast radius and quality, Osmani says, is "a recipe for disaster."
The host challenges the factory analogy. A factory produces a finished car or screw, but software, especially SaaS, isn't finished at release. Production is where it crashes and bugs appear, so a factory that doesn't connect to how the software runs in production is a different kind of factory. Osmani says this is exactly the next phase. A system that can decide what to build and how to verify it can also connect to telemetry, user feedback, and the product backlog, and might eventually become proactive. The host gives an example: an automation that triggers a coding agent when an existing error recurs in production, which attempts a one-shot fix and sends it for review. Letting it merge automatically is possible, the host says, but sounds like a bad idea today. On whether "loop" is the right word, as opposed to "workflow" or "feedback loop," Osmani expects new terms to appear every month. Some will fit and some won't. "Workflow" works too, but Osmani personally pictures it as a loop.
Osmani's concrete example is an app of their own where users can submit bug reports. Osmani used to go through them manually when time allowed and could only address a few. Now Osmani connects other data sources, such as Google Analytics and hosting provider logs, so the system can set priorities and implement fixes based on more than one signal. Suppose a view is very slow for users in India and the app gets heavy traffic from India. That raises the priority. Osmani asks whether priority still matters when an agent can work through the whole backlog, and answers that it does. Every change still has to be checked: what the agent actually changed, and how much manual testing on real devices is needed beyond emulated testing. For Osmani, the benefit is making sharper product decisions without sorting through every signal by hand.
What remains of the engineer's job
The host quotes Ryan Dahl, the creator of Node.js: "The era of humans writing code is over." The host recalls telling Uber candidates that at least half the job was writing code, and notes that this is disappearing. What replaces it?
Osmani frames the answer around "alpha," meaning your advantage, or what models still can't do well. Alpha shrinks with each model release, so it changes over time. Right now, Osmani says, engineers' alpha is taste: are we building the right thing, and is it good? Osmani pushes back on the claim that an agent can judge quality. An agent can tell whether something looks correct or matches a spec, but "good" can mean delightful, something people want to come back to. Osmani thinks it will take time for models to fully catch up on that.
Even if models catch up on judgment and verification in a year or two, Osmani argues that engineers will still need to be answerable for systems, and that kind of trust comes from understanding a system and building expertise over time. Osmani points to Chromium, one of the largest codebases in the world, where key directories have OWNERS files listing a few people accountable for that area. Those people haven't written all the code, just as engineers won't write all the code agents produce. They are responsible for understanding it and deciding what ships, what's blocked, and what's deferred. The host compares this to lawyers and to hiring a professional for a major home renovation.
Osmani adds two points. First, every time software has become easier to create, far more of it has been created. The host notes that iOS app releases and new websites are rising sharply. Osmani acknowledges that not every app has the same value and that an app one prompt away from being copied faces a different market. Still, many more people can now build, which means more potential businesses. Second, historically automation has eliminated some jobs and created others. Osmani says a big open question is what the new jobs in this knowledge economy will be. We may not yet know what they'll look like, but Osmani believes they will come.
Writing with AI, and the pull toward sameness
Osmani says they have published about 18 books, many with O'Reilly, including Software Engineering: The Soft Parts, Leading Effective Engineering Teams, and Vibe Coding, and writes regularly on social media, a blog, and a newsletter. Now that anyone can generate text with an agent, Osmani thinks it matters more than ever that published ideas are worth reading. If you ask someone for 5–15 minutes of their time, you should put real effort in.
Agents help Osmani most with research. For the loop engineering piece, Osmani sent deep-research agents to Hacker News, Twitter, and other sites to find what people had tried, where opinions were strong, what was contested, and what excited or confused people. Osmani says the goal isn't to have agents write the text. It is to form a thesis about what people are struggling with and what educational content would help.
For drafting, Osmani often writes their own version while having agents or different models write theirs, then compares them. Did the models take the thesis somewhere different, or did everyone arrive at roughly the same place, in which case the models add little? Osmani then often runs even handwritten drafts through a model for readability. This is where it gets hard. Osmani says they once spent about three hours trying to improve the workflow, but one model kept producing triads and other telltale patterns of AI writing, which made Osmani wonder whether it was worth the effort.
The host raises the loop engineering article directly. Navi from NeetCode read parts of it in a video, including a section called "Automation: this is the heartbeat," and said it felt abstract and likely AI-generated or AI-polished. The host says they also felt it lacked specifics like the ones Osmani gave in this conversation, and found it wordier than necessary, though they acknowledge that many readers appreciated it and its structure was clear. Osmani explains the process: a handwritten draft, several self-edits, then a model pass for readability. Osmani likes highly structured writing. Ten years ago, Osmani often wrote articles in the 15 minutes before a meeting, adding paragraphs without the structure they wanted. Osmani felt the earlier version of the loop engineering piece probably lacked a clear through-line and felt better about the reworked version, but recognizes that some readers would have preferred something rawer. Osmani says they genuinely struggle with the trade-off between polish and authenticity. The host mentions Michael Novati, who was criticized for an AI-sounding post and said the fix was to edit the model's output much more heavily.
Osmani says a published article usually takes three to seven days of work, and a fast one means they had an unusually clear vision. Osmani usually reviews every line several times before publishing. Osmani also admits to being in the middle on this: models may push writing toward a uniform style, and Osmani sometimes wonders, "wait, what was my writing style?" Osmani rejects the idea of training a custom model on their old writing, because they were a different person then, with different perspectives. Osmani also tested AI detectors like Pangram and was frustrated when paragraphs they had typed themselves were flagged as not human-written. Osmani says this isn't a criticism of Pangram, but sees an opportunity for tools that flag such patterns and help writers find their way back to a human voice.
What's next, and advice for engineers
Osmani recently left Google after 14 years. They say they want to help developers and businesses adapt to how software engineering is changing, since every company they talk to has many questions about the future. Osmani plans to share their next step in the coming months, but will stay in this space and in a role that involves working with developers and the ecosystem.
Asked what experienced engineers and managers should invest in, Osmani predicts that engineering, product, and other roles will increasingly converge: engineers with product sense, product managers with engineering or UX sense, and UX people who care about product. Osmani's advice is to look beyond engineering. People who haven't yet thought about product, technical evangelism, or go-to-market, the other pieces of how businesses succeed, should start. As these role boundaries blur, being able to show employers that you are more than a builder will matter. Osmani's summary: "Don't just be an engineer." Asked whether this comes back to curiosity, Osmani agrees: be a lifelong learner, stay endlessly curious, and there will be roles for you in the future.
One area that you and the team push is Core Web Vitals.
If I'm on a page that has a bunch of ads, I shouldn't start reading an article and then suddenly everything gets pushed down just because the ad is finally loaded. And so, that's where cumulative layout shift kind of comes from.
One question that comes up from a lot of people is this idea of cognitive surrender.
The first is cognitive debt, the erosion of your ability to have good memory. Cognitive surrender is where you blindly give in to whatever the AI says. Its answer becomes your answer.
Software factory meaning
Instead of focusing on the prompting, you're building a system that can do the prompting, but simply having your loops build everything without guardrails around the blast radius is a recipe for disaster.
What do you think will be important in the next couple of years to stay at the front of the industry?
What we're very likely to see happen next with engineering careers is
If you've ever opened Chrome DevTools or optimized a page for Core Web Vitals, you've used software built by Addy Osmani. Addy spent 14 years at Google, most of it on Chrome, going from software engineer to director of engineering. And when starting out in tech as a teenager in rural Ireland, he built his own web browser from the ground up. Today, we talk about the inside story of Chrome DevTools, why it became the closest thing Google has to an IDE, and the problems that are still unsolved like memory debugging. What becomes different when you become a director at Google, and how Addy stayed hands-on building software while building an engineer org with more than 50 people. How AI is changing software engineering, cognitive debt, cognitive surrender, loop engineering, and software factories, and many more. If you want to hear from someone who has spent decades helping developers understand the web and is now thinking deeply about how AI changes software engineering, this episode is for you. This episode is presented by Antithesis.
If you work with agents, your job is no longer just writing code, it's specifying and testing it. And Antithesis is the most effective method of verifying agentic code today.
Before we get into Addy's journey at Google, I wanted to talk about a really cool product at Google, Google Cloud Run, and their recently launched Cloud Run Sandboxes. When you're building AI applications or AI agents, you often want to run untrusted programs like execute some Python code the model generated, or run a headless browser to fetch data from the web, or even execute code submitted by a user. But, how do you make this fast and secure? This is exactly what Cloud Run Sandboxes do.
Cloud Run Sandboxes are ephemeral, isolated gVisor environments that spin up extremely fast. They were built with security in mind. They enforce credential and environment isolation. Basically, they have no access to your service environment variables or secrets. Sandboxes also operate with lockdown network egress, denied by default, to the internet. Accelerate development, eliminate infrastructure toil, and run untrusted workloads with confidence on Google Cloud Run. Try Cloud Run Sandboxes today at cloud.run.
Addy, it's so nice to have you in person on the podcast.
Oh, thank you for having me.
And before we kick off into your career and how you got started, I wanted to ask, just before we started recording, we're talking about how has your day-to-day workflow changed recently a bunch thanks to all these tools?
So, agents have allowed me to take the improbable and turn it into the possible in ways that are kind of weird and wonderful every day. I manage a lot of my life now using agents. And one example, just from, you know, this week, we've got the AI World Fair happening in San Francisco. I'm doing a closing keynote a couple of days in.
And so, I wanted to make sure I wasn't repeating any beats that other speakers had gone into depth on. And I also wanted to make sure there was good connective tissue from my talk to a lot of the other sessions that had happened. Now, normally in the old days, you kind of pray and hope that there was any content from these sessions online, and maybe you'd look through the abstract. I was able to fire off a bunch of agents, you know, go through everything that you can find about the talks from the last couple of days, look at the abstracts, any social content, anything around that that can be useful. And that was able to help me kind of sculpt what I already wanted to talk about into something that I hope is refined and will give people a way to connect from other parts of this conference back to, you know, the way that I'm going to close it up. That for me just feels very empowering. You know, something that would have taken a very long time, if I would have even been able to do it at all, is now very much within reach.
And then you mentioned that it feels like, and a lot of people tell you like, there's the kind of like both fun and chaos at the same time right now, like everywhere, right?
Yeah. Yeah, absolutely. Fun and chaos. I think that, you know, for many of us, when you have a lot of ideas or a lot of vision, you're often bounded by time or how much can I actually do? And in some ways agents have unchained us. And this is one of those reasons why, you know, you keep hearing, "Oh, hey, what are you doing with all of that time agents have freed up?" Well, I'm doing more work. I think for many of us, if we didn't enjoy it, we wouldn't be filling that time up with work, but we're having fun with it. And so, it's fun thriving in that chaos.
Yeah, but now take me back to the very beginning. A lot of us know you and got to know you through your work at Google, through your books, but I'd like to start from even before. Where did you start out? How did you have your first contact with computers? And how did you build a web browser when you were a teenager in high school?
So, I've always been fascinated with understanding how things work. And I grew up in rural Ireland, which was at times, you know, we didn't necessarily have the best internet connectivity. This was back in the days of dial-up.
And so, my first contact with computers was, you know, I was probably 8 or 9 years old. We were very fortunate that my dad was able to get us our first desktop machine. And, you know, I'd play around with apps, I'd play around with just like trying to browse the internet, and it was always fascinating to me, like, how does any of this work? I'm just typing in something into an address bar, and all this information is just rendering somehow. I'm getting back text, photos, videos. Like, how does any of this stuff work?
And so, over time, like, I would build out websites, I'd start to get into programming. My very first programming language was Pascal. So, I'm a big big fan of the Borland tool suite. I learned C++ when I was fairly young.
And there was one year when I noticed that we had a kind of popular national science competition, and traditionally, that competition was very much about, you know, hey, do students have interesting breakthroughs or thoughts on physics or chemistry or any of those things? But the year that I'm talking about was the first year where they actually started to really take computing seriously. And as I mentioned, I didn't have the best internet connection. This was also during the time when, especially if you were a teenager, you started to get into, you know, learning about downloading stuff. And we didn't have fast internet connections back then. If you cared about, you know, checking out a song, you could be waiting hours for that to download. If you cared about trying out a music video, man, that could be a night, two days sometimes to download it.
And so, I tried to study how these kind of download managers that were popping up worked. And download managers kind of offered this one hook. Well, rather than making connection to a server, what if we spawned multiple threads and made multiple connections to a server and we kind of chunked content? You know, it's classical computer science, you know, break down problems into smaller chunks. And so, that was one of the ways, and if the server supported, you know, chunking, you were able to in some cases actually get your file downloaded a little bit faster. And so, it dawned on me like, "Hey, we're using this technique for downloading individual files. Has anyone applied this to how we browse the web pages?" Yeah, and so, obviously I couldn't, you know, do something complicated before I, you know, took my first baby steps. And so, I thought, "Okay, I'm going to try exploring how you build a browser." I started to read, you know, specifications. I was probably 15 years old when I started this, 15, 16. But I started to read specifications. Okay, HTML, CSS, JavaScript. One of the things that I gained a great deal of respect for, and I still have a lot of respect for, is developers throw all kinds of weird crap at browsers. And yet they still render something, right? If you ever want, you know, an interesting experiment in therapy, if you're ever feeling bad about your code, open up the DevTools and just browse the web for 10 minutes. The number of things that will go wrong and yet you'll still be able to probably interact with the site is just wild. And I had that experience.
going wrong, you can just look at the warnings and errors, honestly. Just options.
And so, that was one of the most complicated things when I was trying to build a browser. It was like, yeah, you can parse HTML, you can parse documents, you can load up images, but as soon as you run into pages that stop following those specs and they take a very loose interpretation of what's supported, you have to really, you know, roll your sleeves up and try to behave the way that actual consumer browsers did. And so, I had my fun building out a browser, adding interactivity, and JavaScript support was very difficult. Managed to get it working, and then I'm a sucker for pain because I decided, well, I guess applets and Flash and, you know, we had Windows Media Player back then. So, people were like embedding all kinds of interesting content. I told myself, well, it's not a complete browser if it doesn't support all these other things.
And so, I added support for them. And then I could finally get on to what I actually wanted to work on, which was exploring if I could speed up web browsing for this point in time when you were kind of constrained by the hardware and bandwidth that was available locally.
Back then, this was a personal pain because before I could get a good internet connection, I would literally every weekend specially wear cargo pants that had a lot of pockets and I would fill my cargo pants with floppy disks. And I would walk down to our local library that just happened to have a slightly faster internet connection. And I would try to like save as much as I could, then go back home, check it out on my computer. And so, this was a personal mission for me. I really, really wanted a faster internet connection.
But I finally got to explore this idea. It worked, and back then, in many cases there were servers that supported this idea. It did make things a little bit faster. And so, you know, I ended up building something that worked for me. I took it to this national science competition. I was nobody. I am nobody, but I was nobody. I was just this kid and I was kind of half expecting to just leave the competition and go back home at the end of it and say yeah, you know, I showed some people some cool stuff. There was like a big live audience at the tail end of this whole event. It was, you know, on live TV and everything. And when they called out the overall winner, I was shocked cuz I didn't expect to win this thing.
And that was my kind of first taste of media attention. It was very strange. The weekend right after, you know, you're a kid. You kind of like want to sleep in on a Sunday morning. You don't really have too many people back then like calling your cell phone or whatever.
Yeah.
And my cell phone is just like not stopping ringing. The first call is like from the Wall Street Journal and
So basically you kind of like went viral in the time where there was not even Twitter, right? Like there was like none of this just yet.
Yeah, it was like, yeah, it's Wall Street Journal, CNN. I didn't really understand what was happening.
But it was my first taste of that world. One of the things that I learned from that experience was I still didn't fully understand everything that I was doing. You know, you're a teenager. Just because you can build an app that runs on your machine and accomplishes a goal doesn't mean that you understand all of those layers behind the scenes. And that I think kicked off for me a lifelong thirst for knowledge and understanding how things work. In some ways I can call myself a one-trick pony. I care about understanding problems and, you know, how to fix them, how they work behind the scenes.
I would go on to work at startups. I worked at AOL at one point. Again, continuing this theme of working with browsers. When I joined AOL, at a very AOL moment my first day, my manager was a very kind guy. Said like, hey, yeah, you know, go help the team. Just log into the browser. Go help the team. It's like, okay, cool. I'll do that. So I fire up the AOL browser app and I have my work machine. This app is loaded up. And before I can debug anything, the first thing it asked me for is my credit card. I was like, "I'm sorry, what? I need to enter in my credit card to even start my work?" And my manager was busy, so I don't know if there was like some workaround. I was like, "Okay, I guess this is what I need to do to get started with work."
Wow.
Different time. Very different time.
But years later, I would join Google and I would work on Chrome. But something for me that's been very interesting is, you know, if you treat computing as this onion where you just keep peeling back the layers, there's always something interesting behind the scenes. And the more of those layers that you can peel back and understand, I think in some ways the better you can optimize for that world. You know, there was a time when I maybe only understood the surface of how things render. But if you're talking about browsers, you know, there is the network, there is compositing, there is the JavaScript engine. You go down another layer, there's chips, there's memory, there's GPU. And the more you understand about all of these different foundational pieces, the better you can then optimize and build something that, you know, can serve people even on constrained environments. So, if you have a slightly slower phone, well, I understand now why the phone is slower and what constraints might require us to think slightly differently about what we're building there. So, I'm always a big fan of encouraging people to understand how things work.
And then in the spirit of understanding, one project that you got involved on early on, you know, you built your custom browser that did some cool stuff and you understand how to render, parse, do some of these things. And then you join into the jQuery project. How did that happen?
Yeah,
And jQuery, for a while after you joined, it became for a while the most used library in the JavaScript ecosystem. So, basically across the web. There were a few years of that.
Yeah. I think full kudos goes to John Resig, the creator of jQuery. jQuery was really my first contribution to a big community open-source project. And John was someone that was very welcoming and created an environment where people could, you know, learn and become better open-source contributors. I started off working with the team of people that would deal with triage and issues, and then moved on to like working on blog posts and contributing to code and other ways. But, you know, as you said, it was so widely used that you end up with so many
different use cases people have. And a lot of the times back then, people would have very strong opinions about like, "Hey, this thing should be in the main library in core versus being a plugin." And I got a lot of respect for how you effectively like work with a community while also holding a line in terms of, you know, what decisions should be made to optimize for long-term maintainability, for example. But, it was a great experience. And that helped me kind of carry those lessons on when I worked on my own open-source projects. I was very grateful for that opportunity.
Yeah. And one of your popular open-source projects back in the day was called TodoMVC. Can we talk about what it was and why you started it?
There was a point in time back in the dark ages of JavaScript when we didn't have frameworks and we didn't have libraries. Over time, those things started to pop up, and we began to have quite a few of them. And they all tried to accomplish in some cases overlapping goals, sometimes adjacent goals. And we're talking about libraries like Angular, Backbone, YUI, Ext JS. And if, you know, if you're too young for any of these terms to mean anything, that's also totally okay. But there was this burgeoning community of libraries and frameworks that were starting to pop up. And one thing that I personally struggled with was, well, how do these things differ? You know, you can go and you can check out the landing page for any of these projects and they all say like, yeah, we're going to help you build apps, you know, easier. But I was very big into education and trying to understand how these things worked.
So, I started off by creating basically the same application in every one of these frameworks and try to standardize the functionality so that if you were in the same position I was and you just wanted to get a sense of, okay, well, how does the architecture philosophy change between these things? How does the syntax differ? If they're telling you to build a component or a piece of UI, what is the position they're taking on it versus somebody else? And so, I got a lot of personal value out of the way that I was building this thing up. And so, I put it out into the world. I had no expectations of it being useful.
But basically it was, you implemented a to-do app, or the same to-do app with different frameworks, and you could kind of compare how they differed.
Yeah. Yeah. And the idea was I wanted an application that was simple enough for almost anybody to be able to use and reason about, but it needed to have enough interactivity and enough functionality that you could really kind of stress test at least some of that functionality a framework offered. In some cases, you know, that would be state management or routing or other things. And so, I put this out into the world.
I was kind of shocked at how many other developers were running into this exact same challenge. And the project quickly took off. It started to get a lot of stars back in the day. It got thousands and thousands of stars very quickly and I didn't quite know what was happening. And before long I had people who were working on new frameworks or new versions of frameworks reaching out to me saying like, hey, this is cool. Here's my pull request with my framework. Can you add it? Can we work together on standardizing it? I met some of my first true open source friends through this project. People who are now, you know, very well established in their own means, like Sindre Sorhus, who's written quite a lot of Node modules over time.
This idea of just giving people a simple enough application ended up becoming in some ways a standard for a number of years. I began to see that, you know, if a framework was giving people a tutorial about how to use them, they would actually use a TodoMVC app as their baseline. It's been so many years. That was at the start of my career in many ways. Even this last year, I still see labs sometimes showing off TodoMVC apps when they're trying to test out features, and the longevity of this thing has been very surprising to me.
Another thing that was surprising was at one point, when the project was taking off, Apple reached out to me. Yeah, Apple reached out to me. And specifically the people who are working on Safari and WebKit, and they said, you know, hey, we're interested in working on a browser benchmark to help browser vendors understand, like, are they doing a good job at being responsive? And responsive here doesn't mean responsive in the mobile sense, but responsive in terms of interactivity, and are we responding to clicks and taps quickly? They reached out to me and they said, hey, would you like to collaborate with us on this thing?
And what that turned out to be was Speedometer. Speedometer over the years has become the primary responsiveness benchmark, web application benchmark, for all browsers, and it's continued to be for a very long time. Browser vendors now collaborate together on it. They've kept it up-to-date, so as new frameworks, as new architectural paradigms have come out over the years, they've kept updating it, and that in many ways has carried the legacy of that project through to today, and I've been just very happy that it's given people value of any kind.
And you were building stuff on the side. You were also working at consultancies, AOL, at different startups. How did Google come along?
So Google was an interesting one. I remember one of my first longer periods of time spent in the US was when I was visiting my wife and her parents out in the Midwest, and I was sitting, I remember, in their room watching TV, and there's this documentary about Google that came on, and they showed, like, you know, early engineers that have been working there and why they enjoyed the environment, and I told myself, you know, I would love to work in a place like that someday.
I continued to put out free education into the front-end world and the JavaScript world, web app world, over the years, and at some point I guess Google noticed that it was useful to some people, and so they reached out and wanted to interview me for a DevRel and builder role. There was some tooling that they were trying to build out at the time that they thought could be a good use of some of my skills, but also some just general evangelism they wanted to do in the tech community. And you know, the stars aligned, it just happened to work out, and I ended up working on the Chrome team.
And then when you joined, can you tell us a little bit more about when you were on the Chrome team? What was Chrome like? What kind of work did you and the team do? Because now Chrome is synonymous with web browser. I know there's other browsers, and every now and then, of course, they have some market share, but Chrome has largely won the market. But back then, when you joined, this was not the case just yet, was it?
I remember back when I joined, it was a period when we were very excited about developers bringing their creativity to the platform. So, what can you do to push on the platform and show us both what's possible as well as the gaps, so that we can potentially help fill those gaps and build better APIs. So, I remember there was this great Chrome Experiments site that we had back in the day where we would, you know, sometimes work with studios or work with developers and just showcase, like, "Hey, here's a cool WebGL example that maybe you wouldn't have otherwise come across." And that served as inspiration for some people to maybe even go and then learn more about shaders or, you know, different libraries. It was also a period of time when I would say front-end tooling was still very much heavily evolving. But you know, for
About 2012, 2013?
Talking 2012, 2013. This was at a time prior to what I would now call meta frameworks. So, like Next.js, for example, a meta framework. It tries to... Yeah, it didn't exist. So, we're going all the way back to a time when we didn't have the best build tools even for front-end. We didn't necessarily have well-standardized JavaScript modules, you know, in all browsers. People were still using, you know, AMD and UMD, CommonJS, things like that. And you know, the build tooling and the scaffolding tooling was still very much evolving. And so, this was the period of time when you went through things like Grunt, for anyone that, you know, maybe we're dating ourselves, but Grunt as a build system.
And also when you debugged in the browser, you would use Firebug. You'd open it in Firefox and then hope that in IE it would work, but if it didn't, there weren't many good debugging tools in IE specifically. Later they became better, but back then there was a time when there was none.
Yeah, and I think that, you know, back in the heyday, there were a lot of workarounds people were trying to apply to still have a toolbox of some sort before things got much better. We put some work into working with, you know, the folks who were building out build tools and test runners and scaffolding tools. We worked on our own contribution called Yeoman back in the day. And Yeoman was really about... I don't know that I'd call it, you know, the first meta framework, but I would call it an attempt at trying to bring just a little bit of organization to your starting point. Yeoman was a scaffolding tool we created where you would get a wizard in your CLI and you'd kind of say, well, yeah, I'm trying to build this thing, and maybe I'm interested in using this UI library and this testing library. And maybe I'm interested in deploying to this target. Now, for folks who are listening in, those ideas might now sound very standard and things that you will find in all of the tools you're regularly using. Back then they didn't exist. And I wouldn't be surprised if many of the modules we created back then are still being used under the hood for some of your favorite tools.
So, it was very fun getting to be a part of that moment where we were trying to figure things out and reduce friction. But I will say that, you know, there was this long period where we kept changing tools what felt like every once in a while, right? You went from Grunt to Gulp to Webpack to, you know, to Vite, Rollup. All these things kept evolving, and I was happy to see the evolution, but I'm also happy that things in some ways feel like they've stabilized.
Yeah, there was, I think it was churn, but I mean, that's when innovation happens. Did you work on Google Chrome DevTools?
Yeah.
How did that start? Because I remember in 2012, I'm not sure if there was DevTools, but again, the state of the art was Firebug. I think it was open source. It was actually just superior debugging on the web to anything before. And I'm not sure at what point, but I do remember, you know, Chrome DevTools slowly started to emerge, and it started to bring a bunch of new stuff, like you could do performance monitoring, some of those things. Can you tell me from the inside how did it start?
Yeah.
What you built, how you figured out what to build.
Yeah. So, I have to give a shoutout to Pavel Feldman, who was the tech lead for Chrome DevTools and really played a very large role in helping it come to be originally. There was this period of time when, you know, Chrome was trying to figure out how it differentiated its developer tooling story from WebKit, where we had the, you know, Safari Inspector, the WebKit Inspector. And I think there's a very specific direction that was developer-centric and cared about the ecosystem that Pavel and his team were trying to help out with. And I noticed that they had a very good relationship talking to not just developer evangelists. So, this was the time when we had really sharp minds like Paul Irish around, also working very heavily with the Chrome DevTools team. We would later have folks like Paul Bakaus, who is now known for things like Impeccable, the Impeccable skill for design. And I feel like one of the nice things about that period of time was you had these people who were web developer archetypes and were builders on the side, myself, all the Pauls. And we would try to bring those insights to the DevTools team and help them understand, well, here are the areas of friction that we're running into. In some cases, you can't just build tools to help you out with them because you don't have the underlying instrumentation.
And so I was very happy to see things like performance tooling heavily evolve over the years. Like the DevTools performance panel is just an amazing piece of technology. The fact that you can just hit record, start interacting with your page, and you get a flame graph, you get very deep tracing about where all of the time is being spent. And that continued to evolve over time.
And then we had, you know, really hard problems. You know, some of the hardest problems have been around memory, right? I would say, sometimes, I don't know if it's controversial, that very few developers understand memory management. And that makes it even harder to debug memory problems. And so the state of the art around memory debugging hasn't evolved all that much over the years. But it's a hard problem. The DevTools team tackled a lot of interesting hard problems.
Can we talk about a part where you brought in something new? Because, you know, debugging memory back in the day, it's pretty much, I mean, if you have a language that has, let's say, a heap, you can try to visualize what's on there. You can attempt and maybe succeed at locating which variables there are. And then you can try to also, I mean, variables are the easy part. There's also stacks. And, you know, it gets a little bit messy, but you basically have a memory and you're typically interested in what is growing. And there's a part that I don't know that we got too far on, but you're kind of trying to see, is this getting bigger? What are the loops, where's my stack?
I would say that there are a few interesting arcs where we were seeing, you know, ourselves and developers externally running into certain kinds of friction, and, you know, worked with the DevTools team to try evolving some tooling in that direction. One of the big arcs was embracing the fact that developers were increasingly using frameworks and libraries to build for the web.
Yep.
Now, for anyone that remembers those dark ages, imagine that you have a page that's very interactive. It's using lots of different libraries, and you're trying to debug what's happened. What part of that code do you actually care about? Do you care about the framework code that is powering things behind the scenes? Do you care about the plugins or the components sitting on top of it that you haven't written? Do you care about the code you yourself have written? And so you have all of these very nuanced aspects of debugging that need a solution. One of the things that we tried to introduce was just this respect and understanding that, yeah, developers are going to be using these different tech stacks. You know, we had a source maps story sitting there
Yeah.
where potentially we can start to reason about what's in
You can map back to what part of the code
Exactly.
which is not trivial.
Exactly, which is not trivial. And a lot of kudos to the team, because I think we ended up on a source map story that really helps you reason well about, you know, even if you were using a long tool chain of things. Like, if you take a look at any tools that developers use for any big site, you know, whether it's Uber or Netflix or Twitter, any large site, you probably underestimate the complexity and the number of tools that you are running at any one time for any one task, you know, and being able to still allow people to see, well, hey, here's actually the files that you care about. It's a hard problem. I think that allowing people to get that view was part of the value
that we brought. We introduced different kinds of black box views over the years so that you could say, "Well, hey, actually, I know that I don't care about you telling me there's an issue with, for example, the React library, but I do want you to tell me that there's an issue with the React code that I wrote." And so giving you even those toggles, those controls, I think was very powerful for people.
Mobile was another big moment that changed everything. And you know, if you think about mobile, today I would say there's probably established best practices around the things to test, right? Like you want to test out your viewport width, your tap targets. Like, you know, if I'm tapping on something, exactly, is it big enough? Exactly. You know, there are all these different kinds of sensors even that mobile devices have. We didn't have tooling around any of this stuff originally. And so we ended up building out a nice device mode in DevTools that would allow you to preview what your site would look like at different viewport sizes. You can very quickly kind of toggle and say, "Yeah, this is what it roughly looks like on an iPhone or a Pixel device." And of course, the absolute best kind of testing would be trying it out on one of those actual devices, but even to quickly get a sense of whether you're headed in the right direction was very valuable to people. And we would evolve that over time as more of those best practices started to establish.
I guess the web apps growing up, so PWAs, progressive web apps. There was a period of time when people really wanted to make the web competitive compared to native. And so you think about, "Well, what are the things that are missing?" Well, you need a really good story for offline caching, push notifications, background sync. All of these capabilities that we didn't necessarily have a strong story for, and because these are non-trivial features, you need to have a debugging story around all of them. And so, we helped build out the application panel so that you can go in and for any of these features, whether it's debugging service workers or it's debugging your cache or debugging any of these things, you're able to do that.
And so, even though the tool set has expanded over time for each of these eras, I feel like DevTools has been able to keep up, especially as the APIs in the browser have also been evolving over time to meet these moments.
Well, this is interesting because usually when I look through different companies and their strengths, Microsoft is amazing at building IDEs and so is, for example, JetBrains. But for Google, I never felt that Google was any good at building IDEs except for inside of Chrome. Like whenever I have to debug a web application, the past many, many years, I use Chrome DevTools because it had the kind of debug functionality I'm used to having Visual Studio have, which is breakpoints, conditional breakpoints, so many debug options from, as we just said, performance, memory, being able to simulate some of those things. So, it's very interesting for me to see that it's almost as if, I'm not sure if this was you, your team or Google as a whole, but they realized the browser is very important and so they built almost like an IDE inside of it. You can edit the things inline and I think as engineers or as developers, unless you work in front end, you never really notice this, but when you do, it's fascinating how it came together.
Yeah, it's really fascinating and I think that "Are we an IDE, aren't we an IDE, is that a direction we want to go in?" was always a hot topic for the team. And I think that where things kind of landed was, well, we want to meet developers where they're at because you're always going to have your favorite editor. Now, we're talking about your control planes for your agents. You're always going to have a different surface, right, that you want to primarily work in. And as long as DevTools can meet you where you're at and be useful, I think that's been something the team has tried to do.
We continued having other eras. Yang Gao became our next tech lead after Pavel and helped us through the era of trying to figure out AI is now in the picture, and we want to both be able to help humans reason through this massive amount of data that the browser generates for you, as well as make it possible for you to connect your agent up to Chrome and DevTools and be able to have it just automate a lot of these journeys for you. And so, I think that for the first of those problems, I remember anytime I would work with a big site on their performance problems, you could easily spend half a day just looking at traces before you've even written any fixes at all. And now that we have LLMs, it's very quick to reason through massive stack traces and actually be able to get down to fixes you can make. And that's just been really, really wonderful to see happen.
Addy just described using LLMs to go from massive stack traces to working fixes, which is a perfect moment to talk about our season sponsor, Sentry. You probably already know what Sentry is because you're a developer. If not, just ask a dev and they'll tell you. I use Sentry to monitor the back end of the Pragmatic Engineer for any and all errors. Of course, Sentry doesn't only do errors. They also have logs, replays, spans, profiles, and more because they're all connected by the same trace.
One new capability Sentry has built that I'm really liking is the ability to fix errors. Let me show you. Here's the list of errors on my admin back end. There's a recent error on auth that I want to check out. Let's have Seer run an autofix for us. Seer is Sentry's AI debugging tool. First, it generates a root cause analysis. It's finding some problem with HTTP versus HTTPS URLs. Cool. Now that we know what's going wrong, Sentry can create a plan on how to go about fixing it. I could go and edit this plan, but I'm happy with it. So, let's create an actual code fix. Here's a code fix that Sentry generated. Assuming it looks good, and in my case it does, let's draft a pull request. And boom, the pull request is created, ready to merge. What I love about Autofix is how Sentry went from showing a list of errors inside my application to offering me a fast way to fix it and close the loop while I stay in charge of this bug fix the whole time. Debugging just got a whole lot faster and a whole lot easier. Check out Sentry at sentry.io/pragmatic and start detecting errors, diagnosing their root causes, and fixing issues and regressions today.
Addy mentioned things that change when we work with LLMs. One thing is for sure. If you work with agents, your job is no longer writing code, it's specifying and testing it. And this leads us to our presenting sponsor, Antithesis.
Antithesis is the most effective method of verifying agentic code today. Let me explain how it works. Antithesis runs your whole system in a hostile simulation. By doing so, it finds every bug before users do. And because the simulation is fully deterministic, Antithesis doesn't only find bugs, it gives you a perfect reproduction of every issue. To create such a tool, the Antithesis team needed to invent new kinds of debugging tools as well. For example, here's what's called a bug probability graph. The X axis is virtual time and the Y axis is probability. As Antithesis runs the hostile simulations, it plots time frames when the bug probability increases, which greatly helps with finding root cause of bugs. And Antithesis also has a log visualizer. Vertical lines going down represent events branching off from the same state, and the purple dots are where the bug happens. Antithesis is as good as it gets being able to ship agent-written code. It's what teams at Jane Street, fly.io, and the etcd community use to ship with full confidence. Head to antithesis.com/pragmatic to learn more.
And with this, let's get back to Addy and talk about Core Web Vitals.
And one area that you and the team pushed the industry together is Core Web Vitals. You know, these are a standardized set of metrics to just figure out the real world experience of web pages. You know, before this, again, as a developer, you would measure like, right, how quick does it render or how quick does it download? It was very simple stuff. But you introduced things like LCP, largest contentful paint, CLS, cumulative layout shift, FID, first input delay, and then INP, interaction to next paint. Like you were there, how did the team come up with these things? If you're not a web engineer, it takes a little time to understand them, but it does actually explain how users feel. I feel you somehow inside of Google managed to connect the kind of feel to a number.
I think that the Chrome team has always had an appreciation for user experience research. And again, every time there was a new moment for the web, we'd re-consult that research to understand, well, what are users' expectations and how can we help meet them? The way that we used to reason about performance was very much like, hey, is a page loading? And what does that even mean? Well, for many people, is the page ready? But what does ready mean? Does that mean that I see it? Does it mean that I can click around it and anything actually happens? And so, I think for a very long time we had this almost nebulous way of thinking about page load times. And the team felt like it was finally time to come up with a more nuanced perspective around how we reason about performance.
And so, if you break it down, there are a number of key moments across the user's journey that they care about. Is it happening? Is anything loading? You know, do you see a header? Do you see a spinner? Do you see anything at all? Is there something useful there for you? So, maybe that's a header image. Maybe it's a hero image. Maybe it is a hero video. Maybe it's the core piece of content on the page. Is it useful? Is it usable? Right? All of these different moments can correlate to these different metrics. So, for things like your hero image, you can think about that as your largest contentful paint. And that's not going to generalize across every page. In some cases, the image may not be the most important thing. It might be the article text. There may be cases where you want to be able to interact fairly quickly with a page. I can remember many times over the years when I might be shopping and, whether it's on my phone or on my desktop, I will click the add to cart button and just crickets.
Yeah.
Nothing will happen.
Cuz JavaScript did not load or maybe the event handler was not attached because not all elements finished loading. We know as engineers what's happening, but as a user it's like
Yeah. As a user, like, wait, what's happening?
Then stuff can happen where you just tap tap tap, the event handler gets attached and now you're adding it like twice or three times, but you don't know and yeah.
Humans are shockingly simple, you know. If you think about the experience you have with somebody that's just trying to cross the street, if the light doesn't say they can walk fast enough, they'll just keep hitting that button. That's the same experience they have on the internet.
I think that there were other aspects of user experience that I think we acknowledge were actually kind of problematic. One big one was, over the years, obviously sites tried to monetize as heavily as they could, and so you would see not just banner ads, but you'd see modals, you'd see all of these things thrown up in front of your face. And you know, even if you set aside, maybe there's some validity around a business needs to monetize, those things shouldn't cause a really bad experience. If I'm on a page that has a bunch of ads, I shouldn't start reading an article and then suddenly everything gets pushed down, right? Just because the ad is finally loaded. And so, that's where cumulative layout shift kind of comes from. It's this idea that, "Hey, we should be trying to keep that page stable so the user has a good time."
There was a lot of iteration around how do we define these metrics in a way that captures a few of these different use cases that are very nuanced, because the internet is not all that homogeneous. There are lots of different ways that a person can think about the value of a page and what's important. And so, the team did a lot of experiments, experimented with lots of different ways of thinking about these metrics, and worked very heavily with both the standards community and developers to validate like, "Hey, do you actually believe that these things line up with how you would say you think about the value of your pages?"
One of the things that I always find fascinating is there are some companies where they have thought from the ground up, like if you start from a blank white screen, what is actually important to the user end to end. And then there are many companies where they haven't, for whatever reason, time or they just didn't think about it, they haven't gone through that journey. And so, Core Web Vitals allowed them to finally get a more nuanced conversation going about like, "Hey, what's actually important to us? How can we make sure that whatever key action the user has to take, they can do it pretty quickly and be guaranteed that they're not going to have a bad time?"
Now, you spent 14 years inside of Google. Most of it was inside of the Chrome organization there. You moved over to Cloud AI and were with generative AI as well. We've done research before on Google's engineering culture, but can you summarize what it felt like working there, especially comparing to the startups that you worked at before? You also talk with companies now outside of Google. What were things that were uniquely Google?
As part of my Google journey, there was a lot of work that I did on the developer side, but there's also a lot of work that I did on the consumer side. So, working on Chrome performance, for example. And when you're working on something that goes out to billions and billions of users, which is the case for many Google products now, the way that you think about engineering culture, velocity, experimentation is very, very different, I think, than sometimes how a startup might approach things, especially if you're trying to move very fast. When we try to make a change inside a browser that has a global audience with a lot of people, there's a lot of experimentation that has to happen, and a lot of experimentation that also requires just testing out, well, hey, does this problem not have one solution, but actually a couple of different ones depending on what market you're in? And how do you evaluate success when maybe we have 20 other experiments or 100 other experiments happening at the same time? And so, the A/B testing culture, I would say, was a very big thing, and kudos to the Chrome team for having what is now, I would say, a fairly stable and rigorous process for being able to try those things out in the real world.
I also felt that, even though sometimes from the outside it didn't necessarily always come across, zoomed out at the Google level, I did feel like there were many people that cared a lot about developer goodwill and developer sentiment. But when you have a very large company, obviously it's going to be very challenging to have every group talking to every other group.
Impossible.
We always made best efforts to try getting to a place where we were doing the right things for developers as best we could. But I was glad to see that sentiment and that level of care for the community and for our users. I also appreciated that Google was open to change.
So, I would say if I had to summarize my one big change contribution to Chrome's culture, it would be meeting developers where they are at. And that embrace of people are going to use whatever tech they want to use. You can't tell people what to use very often. They're going to use whatever they want, and your job is to help them be successful on your platform, and to help your users have a great time.
I think that there were a lot of decisions we made over the years that help make that a little bit more possible. And, you know, we had good collaborations with different framework teams. We took their feedback about APIs they would like to see in the platform. It became a lot more of a collaboration with the community rather than kind of guessing what we thought, you know, the community needed to be successful. And so, I was very happy to see that happen.
I would also say that Google was very good at allowing different parts of the company to share their learnings towards some point of convergence. So, for example, the Software Engineering at Google book, one of my favorite books.
I was very happy to see different flavors of that over the years internally at the company because you work at such a big company, you're always curious, well, what is best practice, right? Like, is there a best practice at
There were like internal writings of like, here is how this org is doing some parts of software engineering building or experimentation or whatever.
Yeah, one of my favorite things to do was, you know, I was curious, well, my team might have a perspective on testing or user experience, but how does the YouTube team think about it? And are there parallels? Are there things that we could learn from each other? And there certainly were. You know, even looking at how other people think about the world can sometimes lead to collaboration opportunities. We actually worked with the YouTube team to improve their Core Web Vitals at one point, you know, and they were excited to see that they were just more refined metrics and ways of thinking about experience.
I mean, I guess it's just important to point out that this collaborative nature is not a given in any and all large companies. There are some companies, don't want to name names right now, but where organizations don't feel that they're incentivized to work with each other because they might have different goals. And it's not that they hate each other, they're just like focused on themselves and it can feel a lot more, I guess, political in that sense.
We talk a lot about high agency these days. And I think that sometimes when you see those collaborations happen, it's because there are people with enough agency on both sides that they want to make it happen. And they see the mutual value in collaborating because exploring, you know, how to improve the user experience for something like YouTube, it was extremely nuanced, extremely educational, very nuanced, but also took a very long time. And we just felt like the value was there. I was glad that we could make it happen, you know.
Can we talk about your specific career path inside of Google? So, you spent 14 years there, which is a very long tenure, and I'm starting to develop a bias for like, it's nice to have long tenures somewhere at some point in your career. There's a lot of value to it. We were just talking with Simon, the founder of turbopuffer, about this earlier. What level did you get in? How was your career progression? At what point did you become a manager? And how did you think about things like career, compensation, growing?
Yeah. So, I started my Google career back when I was living in the UK. Actually
So, you joined Google UK? Yeah, I joined Google originally and I believe I joined at a level four, like at the, that was one of the mid-level software engineer. Back then, and I was a developer relations engineer. So, a person that's in DevRel, okay, you're a little bit more focused on, you know, the builder side of things. Over the years I kind of got promoted in that role to like level five and level six. I became a manager within DevRel and then
When you were at level six, at the staff level?
Yeah, and then I was leading part of the DevRel team, and at some point, maybe five or six years in, I started to feel like, you know, I loved doing developer relations, but I am very much a builder at heart. I love engineering and I love product. I love all of it, you know, but
I get it.
I was very curious, you know, what it would be like to be on the other side of that, because I'd been an engineer prior to Google. I hadn't been, you know, in an official DevRel position prior to that, and I was interested in going back down that direction. And so over the years I transitioned back into kind of software engineering, and specifically like an engineering manager role. That gave me the flexibility to both do like engineering work but also manage teams.
But you just had a smaller team at that point.
At the start, had a smaller team and then it grew out. I would say the average at one point was probably in the 45s to 50s.
I think that depending on where you are in your leadership or manager journey, you know, success means different things, and not success from a career perspective, but just success for the organization. Because ultimately what you want to get to is a place where ideally the team is almost self-sufficient. And I write about this a little bit in my book Leading Effective Engineering Teams, but you want to get to a point where, you know, your machine, your org is self-sufficient enough that, you know, you just need to occasionally tap the blimp, make sure that things are working. You can course correct if it's not, but that frees you up to then focus on the next important sets of problems that the org needs to, you know, tackle head on. And that allowed me, for example, to really get deep into thinking about, okay, well, model quality is starting to get better. What does that mean for developers? What does that mean for developer tooling? What does it mean for how we think about benchmarks and collaborations with third-party vendors and all of these other things that are part of developer success?
And so I was glad that I had that time, and then I could take those learnings back to the team and work with them to evolve us into this moment where we could, you know, help developers maximize how useful DevTools can be, and DevTools and Chrome can be for agents.
So, do I understand correctly that, you know, you were an individual contributor, you were going up the career ladder, which is somewhat expected at a large company like Google with the right mentors and the right structure. And then when you became a manager and you switched, but you decided to build a bit more, you then focused on, you still have a growing and increasingly large team, I mean 40 people, that's not a small team, but you tried to help the team fix any issues, help them mostly run by themselves so that you would have some time to actually do some individual contributor-like work so you can keep your hands dirty, but also help the team.
Yeah.
So, like, do I understand that you just prioritized to have that time to build? Because of course, when you're a manager, this could easily suck up all of your time.
Absolutely. And I don't want to make small of all the work it takes to get to that point, because a lot of management is trying to work towards that point. You need to build out a team structure, like in some cases you have managers managing
Managers of other teams.
Of other teams. And we had a global team of people. Of course, that comes with navigating time zones and coordination overhead and communication and all those things. And so I think that we were fortunate that we were able to get to a place where the team was largely pretty effective. We were able to create more of the space. And then I take those learnings and I try to help some of my other managers. Like, how do you create the space now for you so that you can also help us on this journey of modernizing for the AI moment? So I went from L6 to L7 to director in my
Is director L8 or L7 as director?
Directors are L8.
Oh, wow. So that's kind of, well, congrats. Every level becomes somewhat harder and harder, as I understand. But did you care too much about the actual levels, or was it more about the work and things just followed?
I think for a very long time it was about the work. But also, as you get to a point where you feel like the organization is in a healthy place, you do start thinking about, okay, well, next level of my career, taking on different kinds of problems.
Yeah, is it more like the next challenge, right?
Yeah, the next challenge. And so I was very much wanting to go for a director kind of promotion for quite a while, and I was working towards that. And I think that, you know, for anyone that's gone through career changes or promotions, you know that you kind of have to be doing the job for a while before you get it. And what kind of got you where you are isn't what's going to get you to that next level, right? It's a different set of challenges. And so I was excited to, you know, get to start experiencing those kinds of challenges and working more across Google, you know, working more with our VP and SVP layers to try figuring out, well, yeah, what is the next couple of years, or what does the next year look like for Android, for Chrome, for our different platform teams as we're going through these kind of revolutionary moments. We're trying to rethink everything.
I did want to ask, because we have a lot of pretty experienced viewers and listeners: what is the difference in becoming a director at Google? Specifically because director is the first executive level. I mean, different companies call it different things, but like, it's the first one, in terms of responsibility, the weight on your shoulders, because it does feel like, in the management chain, that is the biggest jump at a large company like this.
I think that a good way to think about it is, when I was coming up through the ranks, your director was very often your first point of contact, as you said, at the executive level. They would be the ones who would be keeping you on the hook for making sure that any of your annual goals, quarterly goals, any of that was on track. They'd be the ones that you'd be looking to sponsor any large programs, any new projects, things like that. If things were going like way off track and you were, you know, being held accountable, the directors were often the ones that would be having review forums regularly to make sure that that whole ship is actually still steering in the right direction. And so there was an increased feeling of accountability at that level. You have to pay attention to the details. I think that you can't be successful in that role if you're kind of just letting go. And when I say you want ideally to have a self-running org, it's not about letting go entirely at all, but it's about having enough of a system in place where you get the information you need, and any decisions, any blocks that your teams are running into are surfaced quickly to you so you can help unblock them. I think that that's really one of the biggest pieces, like making sure that the business goals get done, and making sure that people who perhaps sometimes don't necessarily understand how to connect the tech that's being done back to the business goals see that through line very clearly.
I remember that, you know, I was doing the director role through my time working on Gemini and Cloud AI. I was responsible for one of our top goals for the year. You're expected to report on that every week or two and be held accountable. So, you need to do what you need to do to make sure that everything happens to keep those numbers and those goals moving in the right direction. So, there's a lot of accountability, I would say, that comes with that level.
Sounds like it's almost like if you're juggling stuff, you're given like two extra balls, which is like now you have both the accountability, communicating upwards with execs, owning the business goals, while doing everything else in terms of running a now probably larger team, being able to deep dive into the details, so keeping yourself up-to-date. So, yeah, I guess it kind of makes sense that there's a trajectory where the longer you work in an organization, the more context you'll have, the more ready you often become.
Absolutely. Absolutely. And I think an interesting anecdote that I think is worth sharing is, and I don't think this was specific to Google, one of the things I found most exciting in the last couple of years was seeing, as model quality has gotten better and harnesses and tools have gotten better, how many people that were directors or VPs or SVPs or any of these levels were actually rolling up their sleeves and trying things out. And that was awesome to see, because every week you could then have conversations with people like, "Hey, what did you build at the weekend? What models are you trying out? Like, what are you running into friction with? What workflows are you using?" And that's not something that was happening before. Execs were very typically, you know, focused on big company problems or big work problems. Yeah, but that's been changing in the last couple of years.
Which is a very nice segue into the next topic, which is how AI, in your observation and experience, is changing software engineering. Right before we started, we were talking about one question that comes up from a lot of people and you're also thinking about, which is this idea of cognitive surrender.
Yeah.
Let's get into that.
Yeah, so there's two pieces here. The first is cognitive debt. So the more that you use AI, it's sort of the erosion of your ability to have good memory and have good understanding of the problems that you're working on. And the natural follow-up to that is cognitive surrender, which is where you, you know, blindly give in to whatever the AI says as your answer. Its answer becomes your answer. And so you start to really let go of critical thinking, and your ability to solve problems just goes to the wayside. I think that that's something we want to avoid because, you know, personally I'm a big fan of the evolution curve we're seeing with harness engineering and loop engineering and software factories and all of these things. I'm very excited about them. But at the same time, I think that we still need to understand enough about how things work so that if something does go wrong, we're actually able to fix it and not just hope and pray that the agent is able to figure things out.
And it was about a year ago where you wrote about the importance of, when you're working with an agent, and this was before they were as capable as today, but when you're working with an agent, read through what it has thought through, and then, you know, when it generates the code, read through that code, make sure you understand. So like you're kind of doing a review. Now we have a lot more powerful agents. Some people now work with multiple agents. What is your thinking on the reading, going at the same pace of the agent? And, you know, there's a friction of, like, it's not so much easier, so, not like you want to let go and have cognitive debt, it's just that they're faster and it's pretty good for the most part.
With many challenges people run into in life, there's a lack of intentionality around wanting to avoid them. And this is one of those places where I see this happen quite a lot. And my thinking on this has changed a little bit. A year ago, you know, maybe you would have, you know, like one thinking message from the agent saying, "Hey, I'm thinking in
the background." And you'd expand it and you would see a trajectory and you'd see the summary of like all the things that are happening.
And it was like in speed where you could follow as well. It's like every few seconds something coming.
Yeah. And now, if you're using Claude Code or Codex, it's very possible that 20 or 30 sub agents have fired. I am not going to click through 30 of those things to read through their trajectories. But I do make sure that I do two things. The first thing I do is I try to make sure that if there is a summary at the very end, here are all the decisions that were made, I will read through that end to end. If there hasn't been, I will prompt for that decision process. And you have to be careful because you don't want a model to kind of BS you about like the decisions that were made cuz sometimes it can just like make things up, right?
It runs, and as we know it's not deliberate necessarily, but it runs out of context window, like there's limitations to these things.
Exactly. And I guess this goes on to the second thing. I'm a really big fan of this idea of mutual amplification. If you are working with an agent, a coding agent, there are a lot of things that you can do to make sure that the agent is getting better every day and you as an engineer are getting better every day. There are simple things that can play into that. Things like even within the session or within this project, can you log your learnings from the session? Can you log any decisions that were made, any friction that you ran into, anything that you think is unique about how you've approached this problem that I should just keep in mind? That's intentionality. That's like I want to understand how things work behind the scenes. And as long as you have that curiosity and that thirst for at least being just a little bit curious, I think that, you know, you can work with models in a way where you're still preserving a little bit of your cognitive understanding about how things work.
One new building block that's coming up in gen engineering is this idea of loop engineering, and Peter Steinberg wrote about it, Boris Cherny wrote about it, about running loops. A lot of us are trying to figure out what loops exactly are. You also wrote a post about loops. What do you think loops are? How should we think of them? Or is it just something that is useful for a few people? Where are you at with that?
A good way to think about... So, loops are part of this journey we are on to effectively create software factories, or, you know, some people like...
A software factory meaning like a thing where you, you know, you give some instructions, you're like in a factory, like I would like to produce a car, and then there's a fully automated factory and the car comes out.
So, instead of focusing on purely the prompting towards getting an outcome, you're building a system that can do the prompting and generate the outcome, do the testing and verification for you. And it's effectively the next step of, you know, every phase of software evolution is just like a rising tide of abstractions. This is the next abstraction. And it comes with a lot of nuance because I think that, you know, if you tell someone, "Yeah, create a system that will just do all of your work for you," anyone that's been in the industry for a while is going to have obvious questions like, "What about quality? How are you making sure things aren't going off the rails?" And so, I think that you have to be very intentional with, okay, well, what are the parts of this where you're keeping the human in the loop? Are you having your system flag to you that, hey, there are changes that were touched that actually, you know, are hitting a pretty critical part of the system, and you probably do want human review on this. But simply just having your loops build everything without having some guardrails around the blast radius, without having guardrails around how you think about quality, I think is a recipe for disaster.
I hear the analogy of software factory a lot of places. And of course, dark factory as well. Dark factory meaning it's a fully automated factory, as lights are turned off because the robots don't need to see and you save energy and money. But one thing that I keep thinking that is off on this analogy is like, okay, in a factory, you produce a thing. It could be a car, it could be a screw, it could be something. It's there, and it's done. But with software, specifically SaaS and most software that we do, it's not done. Like, when it's finished, we release it to production. And that's where it crashes, the bugs come out. So, I wonder if this whole idea of like, okay, we'll have a factory that produces the software, and it does all the testing, if in production it's not connected to how it's running and having that feedback... You see what I mean? Like, it's a different type of factory that we're talking about.
You hit the nail exactly on sort of the next phase of that. If you can have a system that can sort of decide what needs to get built, how to verify, how to test, and all of those things, there's nothing stopping you from then connecting that up to your telemetry, up to your other systems, up to user feedback, up to any other signals that can help build out the product. You can connect it up to the product backlog, and you can potentially see a world where you then have this system that has access to all of these different signals for how the product can be improved, to a point where maybe it even could get proactive.
Yeah, and we're seeing there's so many examples that you can plug it up. For example, if you're using Sentry, Sentry has automations where like if an error fires in Sentry that is not new, you could have a hook that kicks off your favorite coding agent where it one-shots a fix and it puts it in your review. Now, of course, you could take it further and you could allow it to automatically do it, which sounds like a bad idea today, but you could do it. And I wonder, when we're talking about loops, is this for example a loop that we say, and maybe loop is just not a good word for it. Maybe it's... I think I heard workflow. I heard, like, or if I say feedback loop, okay, like that might be a better word. Is it just a wording thing where like we're a little bit confused with the...
Yeah, I mean, I think that given how fast things are moving, we are very likely to see new terminology sprout out every month. And some of them will be good fits and some of them will continue to require some refinement. So, I could totally see workflow being a better fit than loop, but from a visual perspective, I personally do see it as a loop. Workflow also works.
But then, can you give me examples of loops that you've used or you have seen people on your team or people in this industry use?
Yeah, so I was just mentioning being able to connect multiple signals up to, you know, your software factory or your real...
From production, be it logs or errors or all those things?
I have one app where I allow people to submit issues to it if they run into any problems. Yeah, like a bug report. And historically, I would, you know, manually go through every one whenever I had time and then make a call in terms of like, okay, well, I only have time to address so-and-so and so. I can't go through the full backlog. You can now connect up so many other sources of data. You can connect up your Google Analytics. You can connect up, you know, if you're deploying to a certain hosting provider, there are all kinds of logs that you might get from those sessions as well. You can connect it up to that. And then you can end up with a system where it's able to make decisions and prioritization and then of course do the implementation based on not just one dimension of feedback. So, for example, if in my product it's noticing that there is a particular view that is really, really slow, but it now knows that that's happening for users in India, but that I'm getting a lot of traffic from people in India, it can influence the priority of how much I care about that. Now, how much does priority matter these days when an agent can go through your whole backlog and implement everything? I think it still depends if you care about having to go in and manually do some work to like take a look: okay, well, you said you improved performance. What did you actually change? How much do I have to manually test this thing on these kinds of devices myself? Because I can tell you to go and, you know, do some emulated testing. I'm sure it'll help, but for me it's just about being able to make more refined product decisions without having to sift through all the different signals myself.
I do see more and more people experimenting, trying to put these things in place, again, like from the one-shotting, the bug fix. There's really no excuse to not act on errors, on logs. Open source projects, popular ones, now have things like when people submit an issue, there's a bot that tries to reproduce it. All of these things. So, I see them as loops. One question that does come up there is, okay, well, this is a lot of stuff that software engineers used to... then we didn't have all the time for it, but we did a lot of it. And what this means for the future of the profession: Ryan Dahl, the creator of Node.js, wrote that, I quote him, "This has been said a thousand times before, but allow me to add my own voice. The era of humans writing code is over. Disturbing for those of us who identify as software engineers, but no less true. That's not to say software engineers don't have work to do, but writing syntax directly is not it." And a lot of our time spent... I remember when I interviewed people at Uber, I would tell them like, well, we're going to spend at least 50% writing code, so we're testing you on writing code. This is kind of vanishing.
Yeah.
What do you see replacing it? And what do you see the essence of software engineers, builders, AI engineers, however you call them, being?
Yeah. I always go back to what is alpha. So, my definition of alpha, alpha meaning, the definition of alpha is advantage, right? So, what is the current thing that models are not very good at doing? Alpha is going to decay in some way with every model release or every series of model releases. So, it's going to change over time. So, for software engineers, we very often say that your alpha is in taste, in terms of: are we building the right thing? Where are we putting our energy? Is the thing that we are building actually good? And good, you know, sometimes people will say, yeah, but an agent can tell you if it's good. I push back on that. An agent can tell you if a thing looks correct, if it's matching a spec. It doesn't necessarily mean it can tell you what's good. I think that good can mean good from a user experience perspective, could be delightful, could be something that a person will actually want to come back to. And it is still something that is sufficiently nuanced that I think it's going to take time for models to actually catch up to a point where they can replace that fully. We tell people that judgment, verification, all of these other aspects continue to be important. And I do believe that. But even if you follow through and you say, okay, well, maybe a year or two from now, models will catch up on these different aspects, we still need engineers to be answerable for these different systems. And to be...
Accountable, right?
Yes, accountable, answerable. And that's something that doesn't just happen overnight. That happens when you understand a system, people trust you, and you have that expertise. An example I've been telling people this week is back when I worked on Chrome, you know, Chromium is a massive code base. It's one of the largest code bases in the world. And it's sufficiently complex that for every key part of that system, you will have a directory with an owners file. And that owners file is going to contain a small number of people who are effectively accountable for that part of the system. They might not have written all of the code for it, in the same way that, you know, we may not have written all of the code; our agents may have written some of the code. But they're the person that's on the hook for understanding, for gating, for making sure that someone is deciding what ships, what's blocked, what do we defer. And so, I think that that is something that engineers are going to continue to be valuable for. And that's going to help us to make sure we're building stuff that is stable, reliable. People can actually, you know, use it with some confidence.
I do agree with this because I think accountability is some... That's why so many businesses are working. That's why, you know, lawyers always have a job, because the regulation is there, and you can look up all the court cases, and you could understand how the law is interpreted. But they've done this, and they often take some level of accountability. In fact, if they grossly do not do their job, you actually have an option to, for example, take legal action against a firm if they would be proven to actually just ignore what they're doing. And I guess, you know, that's a good example where in software and anywhere where there's value, this will be valuable. Like again, if you're renovating your house, if it's not a big deal, you might do it yourself. If it's a big deal, you just call a professional.
Yeah. Yeah. And I think there's two related notes to this topic. You know, every time that we've made it easier to create software, we've exponentially created more of it. So, the total addressable market for builders...
And this is happening right now. We're seeing it in the stats in iOS app releases, websites, all of that.
Yeah, it's going through the roof. And that's not without nuance. That's not to say, you know, that every single app that's being created has the same value, right? We of course have these conversations about like, yes, if your app is like a prompt away from somebody else copying it, you know, it's a different world that we live in. But that still doesn't change the fact that we have a much larger number of people that can build now. And that's a lot more potential businesses and startups that could potentially thrive. I continue to be very excited about the profession from that aspect. I also think that every point in time in human history when automation or a form of automation has come into the picture, we've automated away certain kinds of jobs and then replaced them with other kinds of jobs. And so, I think a big question for the future is what are those jobs going to be in this new knowledge economy? We may not necessarily have exact, you know, frames for what they're going to look like just yet, but I do think those are going to come.
Yeah. And I wanted to talk to you about your writing, as a fellow writer to a fellow writer. You've been a really prolific writer in terms of books released just in the past few years. You've written the short book Software Engineering: The Soft Parts, a free book, about 50 pages, a really good read. You've written Leading Effective Teams 2 years ago, and last year Vibe Coding. And on top of this, you regularly write long-form on social media, LinkedIn, X, your blog, your newsletter. You write a lot for someone who actually has a full-time job, and I can tell you, when my full-time job often involves writing... How has your workflow changed in writing, especially now that we have AI tools or other tools?
I would say that now that it is very easy for anyone to use an agent to create a body of text, I think it's more important than ever for us to make sure that the ideas we're putting out in the world are actually worth people reading. Because if you're asking somebody to spend 5, 10, 15 minutes reading a thing,
like actually put some effort into it.
My workflow has changed quite a lot in the last couple of years. I've now published, I think, 18 books. I've worked with O'Reilly on many, many titles over the years. I think the agents have helped me the most probably with just being able to reason about the thoughts in my head, and especially try to connect those back to how other people are thinking about related problems.
So, on any given week, we can take Loop Engineering as one example, cuz I was recently putting together a piece on that. When I'm working on a piece these days, I'm always curious like, what are other people thinking that's related to this? And so, I can fire off a ton of, you know, deep research agents to go and check out, you know, Hacker News or Twitter or other places, and just give me a sense of what are things people have tried out? What are things where people have particularly strong opinions either way? What is considered contentious? What are people excited about? And what do they have big questions around?
And that's not to say, "Hey, agents now write the text for me." It's about forming a thesis about what people are struggling with and what educational content could be useful for them around that. And so, agents have been very useful for me in my research.
When it comes to writing itself, this is a very interesting one. I think that a lot of people struggle with, like, how we be using agents and these tools these days. I will very often do two things when I've got an idea for a piece. I will start to write out a personal kind of handwritten version of a thing and I'll also have an agent or different models write out a version of the text as well. And I'll compare sort of, okay, well, I started out with a thesis. How did those other agents reason about this thesis? Did they take it in a very different direction to what I was thinking about? Or did we all kind of converge on the same thing roughly and there's not really a lot of value that they're offering me.
Once I actually feel like I've got the ideas then pulled for my piece, for me readability is important and so even if I've handwritten a thing, I will often put that through a model to try improving readability. And that's a hard thing sometimes because I still feel like models are sometimes not the best at writing human looking, even if you're editing something. Writing human looking text.
I remember the other day I was trying to play around with, hey, can I improve my workflow? And I feel like I easily wasted three hours of time because no matter what I did, one of the models I was using continued to generate text that looked like it had triads and these patterns, you know, of AI written text, and I really struggled with that because it feels like, you know, is it diminishing returns trying to use the tools
But specifically
Yeah.
specifically on loop engineering. So Navi, aka NeetCode, he did a video where he looked at the loop engineering, specifically your post as well, and he was reading it. He was reading the part "Automation: this is the heartbeat," and he kind of read a few paragraphs and he was saying it felt abstract. It felt that there were like terms that an AI would have done, and he said that, well, either this is AI generated or it might be someone's thoughts, but then an AI put it out. And when I also read it, in that specific Loop Engineering piece, like I missed the specifics. Right now we talked about, like, all right, here's the things that you have been doing. Yeah. And I was wondering on, like, how did you write this and how do you feel about that specific piece now?
For that piece specifically, so I started off with a handwritten piece. I did a number of like editorial passes myself. I then handed it off to a model to try improving the readability pass, and that's where I start to struggle because I'm a fan of structured writing.
Yeah. Oh, I know.
You know, I'm a big fan of very structured writing. I am not a... I remember the kind of writing I was writing 10 years ago maybe even, and I remember, you know, you were mentioning people who like to write struggle with time. I would very often write articles in the 15 minutes I had before, you know, my next meeting, and I'd just try to add more paragraphs, and it wasn't as structured or as high quality as I'd like. And so I sometimes struggle with: are people looking for authenticity even when it's not that structured?
So, for example, I might have the version of Loop Engineering I started out with. Even after a few iterations, did it have a good enough line through or thread through the whole thing that made sense? Probably not, or at least that's how I felt at the time. But using an agent to try helping rework that, so it did have a clearer line through it, I felt better about the piece. But somebody else reading it may felt like, okay, well, actually I would have felt better if it wasn't as structured
If it was a bit more raw that you did not feel that good about.
Yeah. Yeah. And you know, you can say the same thing about, you know, typos, about, oh well, hey, is the article following a consistent structure or beats compared to a final piece? So, that's something I struggle with.
Yeah, cuz the final piece, don't get me wrong, like, you know, when we look at views and comments, a lot of people appreciated it. And when you read through, again, I'll link it in the show notes below to read it, it did have a clear top-level structure. To me, it felt that it was maybe more wordy than it needed to be and it just lacked the specifics. But I do see, by the way, other people experimenting with this as well. I talked with Michael Novati, who was criticized for a post which he actually worked a lot with, but it sounded very AI because in the end he did it. And later he wrote another post which looked a lot better, and I asked him what he changed. He said like, "Oh, I'm actually just like tweaking the output a lot more." Because he's also a big fan of dumping his thoughts, getting some help for structure, and actually, you know, like behind the scenes spending a lot of time like, "Will this be worth reading?"
Yeah. And I think that even with access to better tools these days, very often people see an article that I put out, I very likely spend probably at least 3 to 7 days just trying to work through it.
It's kind of like brewing the ideas.
If something comes out very, very quickly, it's probably because I had a surprisingly clear vision for what I wanted to say. But very often it takes a while for something to bake well. And I do end up going through every line of text very often before publishing it, a few times. And I think that for those of us that are working with AI tools, you start to question, well, am I being influenced by the way
Yeah.
that the models... cuz the models are going to, in some cases, I think, provide a homogeneous take on what writing looks like. And you start to feel like, wait, what was my writing style, you know?
Do you feel you're kind of in this middle right now a little bit?
I very much do, and there's a part, you know, you take it to the experimentation phase, and there will be, you know, maybe there are people who will say, "Hey, you should train a custom model on your old writing." I was a different person when I did my old writing, and I had a different set of perspectives, a different set of nuances that I cared about. And so, I would feel a certain way about that person writing, you know, these new pieces, too. So, I think that we're still learning what the best practices around this are.
And I also, you know, I remember the other week, I was just curious, like, if people are checking out these posts using Pangram or GPTZero or any of these things, how does that change my workflow? And I was just getting very frustrated cuz I would literally type out like human sentences or human paragraphs, and I'd paste it into Pangram, and it'd be like, "No, this doesn't look human-written." I'd be like,
Wow. Okay.
What? And that's not to say anything bad about Pangram. That's just to say that I think that there's opportunity for tools to help writers, you know, flag things and then help them understand like, "Okay, well, how do you get back to your human way of writing?"
Yeah. So, as closing, you've just closed down 14 years at Google. You've announced a big decision last year, leaving. Congratulations. I know it must have been, you know, like a big one to decide on. What is next for you? What are you looking at? What are you excited about?
What I am most excited about right now is helping developers and businesses kind of meet this next moment of software engineering changing. I think that there are a lot of open questions. Every single company I talk to has got so many questions about what the future's going to look like, but there's also a lot of opportunity in there to help and to figure it out with them. So, I'm excited about that. I'll be sharing, you know, in the next couple of months what my next thing is, but I'm definitely going to be staying very much in this space, and I'm someone that enjoys working with developers, working with the ecosystem. So, I will very much still be staying in a role that allows me to do that.
So, right as you're exploring your next career step, you know, may that be joining another company in an exciting role, who knows if you'll do something yourself. Clearly, you're thinking a lot about your own career. What advice would you have to someone who has some experience in the industry working as a software engineer or engineering manager, and they might be in a similar shoe where they are like, all right, it's time for me for a change. I want to set myself up for success looking ahead. What skills would you advise that they invest in? What activities, how to think about their network? What do you think will be important in the next couple of years to stay at the front, like the meat of the industry?
What we're very likely to see happen next with engineering careers, as well as product and other roles, is the unbundling of these careers. In a way, we're unbundling, where we begin to see more of these roles converge. So, the engineer that also has product sense, the product person that also has engineering sense or UX sense, the UX person that also cares about product. So, the guidance that I would give people is think beyond just that narrow lens of engineering. There are many people, especially if you are senior, that have already had to think about these different aspects of success. If you are not someone that has had a chance to think about, you know, product or technical evangelism or go-to-market or any of these other aspects that are generally different puzzle pieces of how businesses are successful, think about the non-engineering things. I think there's a lot of value there. And if you can show employers that you are not just a builder, but you are someone that can help them as these roles start to become a little bit fuzzier, I think that you can be successful in these times. Don't just be an engineer.
And so it sounds like it's one of your core values: go back to being curious, learn, and see where you can help beyond just building.
Yeah, be a lifelong learner, be endlessly curious, and I think that there will be roles for you in the future.
Addy, thank you very much.
Thank you so much. This is great. Thank you for having me.
I really enjoyed this chat with Addy and I hope you did as well. I do feel that Addy's career shows that working hard at always going a layer deeper pays off over time. When he was a teenager, Addy already built a web browser from scratch, which is a massive project. He never stopped being curious about web browsers, then developer tools, and today he brings his same curiosity to AI agents.
The concept I liked that Addy mentioned was this idea of cognitive surrender. The more capable AI agents become, the easier it is to let them do what they do. For example, a year back you could still follow along how an agent worked step by step. But today, if you have multiple sub agents in parallel, there's no way you'll keep up with every step. This creates this weird paradox. AI makes it easier than ever to build software, but it also makes it easier than ever to lose understanding of the software that we're building.
Addy had a really good counterpoint to this, which was mutual amplification. You want to amplify your own understanding as the agents get better. So, have the agents record decisions, document what it learned, and explain unusual choices. It's just no longer feasible to understand how or why every single token was generated, but you want to keep understanding the important decisions that any of your agents make.
Finally, I liked our conversation about what happens to software engineers when writing code is a smaller part of a job. Addy's answer is that accountability remains the job. For example, inside of Chromium, specific software engineers own parts of the code base. They obviously have not written every line of code that they own, but they understand the area, they decide what matters and what doesn't, and they are accountable for their part working well. I think it's reasonable to assume that AI will push this kind of accountability to be more visible for any and all software engineers.
Do check out the show notes for related The Pragmatic Engineer deep dives on Google's engineering culture and AI engineering. If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A big thank you if you also leave a rating on the show. Appreciate it and see you in the next one.
Article published
