Charity Majors on Shipping Code You Haven't Read and Why Both AI Camps Are Right
The Pragmatic EngineerCharity Majors, co-founder of Honeycomb and co-author of Observability Engineering, describes herself in this conversation as someone who was skeptical of AI for software development in 2025 and changed her mind based on what she saw. The conversation with the host of The Pragmatic Engineer is built around one question she keeps coming back to. The question isn't whether engineers will ship AI-written code they have never read, but what it would take for them to be comfortable doing so. From there she covers code review, non-determinism, the growing split between AI enthusiasts and the people on call, the unfinished business of DevOps, observability for agentic systems, and what all of this means for managers, directors, and junior engineers.
From Parse to Honeycomb: the tool that made a hard problem easy
Majors' first job was at Linden Lab, building Second Life. She later worked at Parse, the mobile backend service, which Facebook acquired and eventually shut down. She calls it her "first great lesson" that most acquisitions fail. She says she is still grateful for the experience. As a lifelong "startup kid," nobody knew her name until she was leaving Facebook, and that was when investors started offering money.
Honeycomb's origin comes from a specific experience at Facebook. Parse ran on AWS and Ruby on Rails, but after the acquisition the team got access to Facebook's internal tools, including one called Scuba. Parse was growing on a hockey stick curve, with over a million mobile apps hosted by the time she left. Almost every week some app would suddenly hit the top 10 on iTunes and break something. Finding the cause was a needle-in-a-haystack problem. The app spamming the logs might not be the real culprit, since it might be stuck behind the actual cause. Diagnosis used to take hours or weeks and often depended on luck. Once Parse's data went into Scuba, she says, it became "click, click, click, oh, there it is."
When she was about to leave, she planned to become an engineer or engineering manager at a company like Slack or Stripe, and realized she would be far less effective without a tool like that. Her original plan for Honeycomb was modest. All startups fail, so she would spend a year or two writing Go, then open source the result and take it with her wherever she went.
Individual productivity and the obsession with speed
In 2020, before AI tools, the host and Majors both answered a reader's question about whether individual developer productivity can be measured. Asked the same question now, she says measurements are useful as supporting data, "like color in a painting on the wall," but Goodhart's law still applies. Managers need data so their judgment isn't just an offhand opinion shaped by bias, but the numbers are not the whole picture. She also pushes back on the recent emphasis on individual output. Teams, she says, are still what matter.
What encourages her about the AI moment is that it forces everyone to keep asking what good looks like, what productivity means, and what "better" or "great" would look like. She thinks it telling that the industry jumped straight to speed, doing the same thing faster, and she now considers that "a very immature description of what better is."
The host gives an example. Anthropic had just published a podcast with Spotify's engineering leadership describing thousands of changes shipped per day or week with Claude Code. Meanwhile, the host had trouble publishing episodes because Spotify was down. The discussion was all about speed and not about quality or functionality people want. Majors asks whether customers really want the buttons in their apps moving around all the time.
How her view of AI shifted
In a year-end post, Majors wrote that 2025 was for AI what 2010 was for the cloud. In March 2025 she and Fred Hebert gave the closing keynote at SREcon, just after the term "vibe coding" was coined. When they told the audience to try vibe coding, people groaned and laughed. Their pitch was that engineers should learn AI "because you can complain better if you learn it," which she says she meant sincerely. At that point she saw AI as a big feature, bigger than a programming language and comparable to the cloud, but not generational.
Her own turning point was November 2025, when Opus 4.5 was released. Looking back, she has written that the shift was visible earlier, and that it came from the harnesses and tooling more than the models. Around July, people were already saying it was coming faster than expected. The tooling went from something like a shell script that retries to a lot of surrounding infrastructure. In the early months of 2026, she says, it felt like everyone around her was trying AI again and changing their minds.
She doesn't think the earlier skepticism was wrong. It's an extraordinary claim that AI can write code about as well as a median software engineer, within limits, and the industry has heard that promise before. She points to a sticker joking that "we already have a programming language that lets you" do this, with COBOL as the punchline. The host adds no-code and low-code to the list. Skeptics were right those times. Now, they agree, skeptics are wrong.
What would it take to ship code you haven't read?
Majors sees the same pattern playing out with the question of shipping unread code. She thinks arguing about whether or when it will happen is pointless. The useful question is what it would take, and answering that is engineering.
The host's first reaction was that they would never do it. On reflection, they note that trust has always come from indirect signals. Pre-AI, a trusted teammate saying "I vouch for this, I've hammered it" was enough. If a change could be shown to have been tested in some harness, that could work too. Majors suggests a concrete path. Run a human and an AI reviewer in tandem for a few months and compare how much each catches, whether it's about the same, more, or less, while training the AI. Whether that takes five days or five years, she says, the direction is clear.
The host's closing summary picks up her image of a "trust account." If trust is being debited during code creation because no human reads the code, it has to be built up somewhere else: tests, evals, and guardrails.
Rewriting instead of editing: the Phoenix architecture idea
Majors thinks this direction is good for the profession, and she leans on Chad Fowler's writing on the "Phoenix architecture." The host reads Fowler's argument. Immutable infrastructure, stateless services, containers, blue-green deployments, and infrastructure as code all share the premise "never fix a running thing, replace it." AI extends that premise from infrastructure into application code. When rewriting is cheap, editing in place becomes risky, because mutation accumulates entropy and replacement resets it. Majors sums it up: "code is cache."
The analogy is the move from servers as pets, configured and repaired by hand, to disposable machines. For 60-plus years, software has been edited in place. Majors says the economics are what change this. You could generate 10,000 variants of a function faster than you could write it once. That means you'll need a lot of evals and tests, but cheap generation pushes in that direction. The cost of software has always been tied up in maintenance, and we trust code largely because it has been running in production for a long time.
She is careful about how far this goes. Anyone who has done a hard database migration, she says, should be humble about our ability to extract implicit contracts and store them. She believes the industry can go some distance and further than it has now, but she doesn't know how far. Her example is Parse. The team spent about six months writing the original Ruby on Rails API and then two years rewriting it in Go. Go was immature then, so they had to write their own database drivers and other pieces, and the original code had no type safety. In a strangler-fig migration, she says, "you literally find the contracts with your users by breaking them, one after the other." That doesn't seem like the ideal artifact to her. The contracts should be stored somewhere, and there should be reviewable architecture diagrams that generate code to spec.
The host points out that this echoes 1990s ideas like UML-driven code generation at Rational Software, which didn't work out, possibly because generating and reviewing code was still expensive. Both wonder whether those ideas may now be feasible. Majors says that is her hope.
Learning from sysadmins, ops, and QA
Majors started as a sysadmin at 17 at a university and remembers how stressful earlier automation shifts were. Sysadmins agonized that things would never be the same, and everyone adapted. The people who ran updates by hand on every server built the systems that replaced that work and spent their time writing code instead. The host notes that the sysadmin role has largely disappeared, but its practitioners understood operating systems and hardware and moved into software engineering, product management, even tech sales. Majors adds that her generation are still the best debuggers. She's glad not everyone has to learn about CPU and memory, but that knowledge still comes in handy, and she sees parallels for generated code.
She has written that lines of code are not the ideal artifact to review. The tools to do better don't exist yet, but many of the ideas do, and most come from operations and QA, "two domains that software engineering has historically been rather snobbish about." Ops and QA, she says, have always been concerned with what is, while software engineering has been concerned with what should be. She finds it strange how much software engineers believe the world exists in the repo. It doesn't; it's production. Some companies don't even let engineers look at production.
What excites her about AI is that it pushes the discipline where it has needed to go for a long time: "Production is not what happens after development. It is a stage of development." The host recalls the old "don't deploy on Fridays" trend on tech Twitter, when Majors argued that teams should be able to deploy anytime without fear. She restates her position: as soon as you merge, the code should be going out, and you should have to stop the train to keep it from reaching production.
Unbundling code review
The host says nearly everyone they talk to agrees that code generation is cheap and human code review is now the bottleneck, so the industry is building better review tools. Uber, for example, has built tooling to surface the important reviews. The host admits they never liked doing code review unless it was small and with someone they cared about, and they care even less when the author is an AI.
Majors says part of the problem is that code review means many different things in different places, which invites projection. If you say you don't want code review, people hear that you don't want to talk to coworkers or mentor juniors. Code review is overloaded. Some of its functions are good, some could be done better another way, and some are highly cultural.
She cites her former Parse colleague David Poll, now working on pull requests and code review at GitHub. For him, code review is where a team decides whether it wants a change in the product at all. Majors calls that a great discussion and exactly what humans are good at: whether the mental model is coherent, whether to add something, what the architecture and API design should be. Ideally that happens before the code is written. Reading for syntax and bugs isn't evil and can be a teaching opportunity, but it doesn't feel like a great use of anyone's time. The host suggests it is most useful when onboarding someone onto a team without written guidance or lint rules. Majors replies that people fill cracks with their own time when guardrails haven't been built, and that the industry has never thought to extract many of those things from the process.
Her model example is Intercom. A decade ago their CTO said "shipping is your company's heartbeat," and they ship a Ruby monolith hundreds of times a day, which she says is not trivial. She considers them a high-water mark for teams founded before AI, with strong engineering discipline, that have become AI-native. They wrote about AI-validated PRs with a very high bar, effectively putting their most senior engineers' wisdom on every diff. That frees humans from nitpicking, which they aren't good at anyway, so they can discuss whether the change is going in the right direction.
Non-determinism demands more discipline, not less
Majors has written that non-deterministic systems require more engineering discipline. Asked what that means, she starts with tests and evals. When code is treated as a trusted artifact, tests cover only the failures humans can predict, plus regressions added after incidents, which she says is not a high bar. She highlights conformance testing: if I'm not going to read this code, how do I know it performs within the boundaries of the last version I generated? For many workloads, she says, not changing too much matters as much as absolute performance.
She prefers to think less about "AI" and more about deterministic and non-deterministic systems that have to work together. Determinism isn't going anywhere and is enormously valuable. The job is to make AI "kind of boring": give a valuable but erratic tool carved pathways where it acts as a superpower instead of eroding foundations.
The host recalls Martin Fowler saying on the podcast a year earlier that non-determinism is AI's biggest change, and notes that businesses want software that behaves the same way every time. Their example is an open-source resume-scoring tool used in HackerRank's applicant tracking. Run locally with a small model (Gemma is recommended), the same resume scored anywhere from 66 to 99 across about 100 runs, while many companies set the bar at 85. That turns a recruitment aid into a coin flip. Majors says we have to be able to call that bad, because AI is not the right tool for every use case.
Honeycomb's AI norms: you own the loop
Honeycomb has been running a series of conversations on AI norms and values. Majors says she wouldn't have trusted the company's judgment a year ago because it didn't know enough. Now, if a coworker says AI is the wrong tool for a job, she trusts them. Her view is that "you got to get worse before you can get better."
The norms start from the premise that the bar has gone up for everyone, as it always does with powerful new tools. Asked whether it's the bar or the floor that has risen, she says she doesn't know, and that the industry is "wandering in the wilderness." But you can't stop wandering or you'll be left behind. The only viable way to define the bar is better outcomes. Another principle: there is no "human in the loop." You own the loop. "Claude said this" is not an excuse; the work is yours.
On communication, she offers a baseline. You can't send anyone something you haven't read, and if it would take them longer to read than it took you to make, it's probably slop. That is disrespectful, because it pushes the cost onto the recipient. She has also caught herself asking colleagues questions without first trying to find the answer. The host compares this to how new grads learn to respect senior colleagues' time, and argues that attention is now the scarce resource.
Majors distinguishes two uses of AI: a shortcut to avoid thinking, and a tool for thinking more deeply and rigorously. Both have their place, but for core job functions, and especially when asking someone else to review your work, Honeycomb wants the second. She says this isn't absolutist. People writing in a second language or who are neurodivergent may lean on AI, and that is still respectful. Existing standards of quality and respect are enough, she says; they just need to be applied. The "look what this can do" phase, she adds, is over.
The two camps: people on call and people not on call
The host describes two camps, the "AI-pilled" and those who seem to hate AI, with little dialogue between them. Majors says neither side is making it up. Enthusiasts feel the race: competitors are moving faster and leapfrogging, and the industry is "on the inside of an exponential curve," which she calls rare and usually short-lived. That is real, and it's wise to prepare. Ironically, she says, both sides feel like the embattled minority defending the truth.
Often the division comes down to who is on call. The people where the buck stops are seeing melting mental models, slop, and their hard work dissolving, with no end in sight. The host confirms that on-call engineers report more incidents and carelessness, with systems getting "way worse" since AI adoption.
The host shares something heard from inside Meta. There has been a flurry of SEV0s, the highest severity, over roughly two months, in areas including Instagram and WhatsApp where reliability and trust-and-safety staff had been cut. The host says each incident has its own postmortem and the connection isn't direct, but Meta hasn't had it this bad in close to a decade. When the host told this story at a conference, people from other VC-funded and public companies quietly said the same thing was happening to them, and that they'd thought it was just them.
Majors points again to Intercom, which she praises for publishing the unflattering data. By her account, their reliability and code quality declined for 18 months and had only recently started possibly recovering, still below where it was. Her plea is that teams talk about the wins and couple them with the costs. There are incredible gains in rewrites and in automating away toil, and she says nobody she has talked to would give them up. But people who see only the wins assume their on-call coworkers are just afraid of being automated away ("No, dude. You be on call and then see how you feel"). On-call engineers who never hear the costs admitted assume the wins are fake. "Tell the whole story," she says.
Why software is AI's killer app
Majors calls AI "just technology," and argues that software is its killer app. Software is made of logic and language, and so is AI, which means guardrails, checks, and validation can be built in. She doesn't see how to do that elsewhere, citing courts sanctioning lawyers who submitted briefs with hallucinated content. The host adds that software may be the upper bound: if something can't be automated in software, where validation exists and training data compiles, other industries are unlikely to manage it.
The host tests each new model by asking it to write in the style of The Pragmatic Engineer and can always tell it's AI, whereas generated code often looks like something they could have written. Majors agrees AI is much better at code, calling software "a simplified version of language for a purpose." She spent time trying to write more efficiently with AI and stopped, because "writing is thinking on paper" and there's no shortcut for that thinking. She uses AI for structure and feedback, not to generate her writing.
The host predicts that in five years, hiring will come down to conversation, and candidates who invested in their own thinking will stand out against those who outsourced it. Majors says she is excited about "leaning into the parts of being human together." She dislikes spending the day switching between agents and people on Slack because it feels too similar. Honeycomb is fully distributed, and though she loves not leaving the house, she craves in-person connection. Her hope is that people remember "we're in charge of the machines."
DevOps set out to build one feedback loop, and failed
DevOps, Majors explains, existed because devs and ops were split, with code thrown over a wall. She calls that split the "original sin": half the people write the code and the other half understand it, and she argues you can't really understand code you don't operate. The first wave, getting ops people to write code, succeeded. The second, getting software engineers to understand their code in production, was less successful. In her view, 20 years of DevOps was about creating one feedback loop connecting people who write code to that code in production, "and it failed."
She doesn't want to erase the separation between platform and application teams. It is a healthy seam: one side owns the stability and resilience of the infrastructure, the other owns whether every user gets a good experience, and one can be true without the other. The host notes that Anthropic has arrived at the same split, with cloud platform teams and applied AI teams, and that the two groups even disagree about whether software engineers will become obsolete. The platform people say no, while the applied people say maybe. Majors is not surprised. What matters, she says, is that engineers need fast feedback loops, and agents are breaking the idea that the code is the source of truth.
Modern observability: telemetry as a product decision
Majors revisits the "three pillars" of metrics, logs, and traces. She calls metrics and logs system exhaust. Every team runs third-party software it doesn't own, and that output should go somewhere cheap. Your own code is different. It is the crown jewels, and its telemetry should be a product decision, stored once with its connective tissue. She says the value of rich data grows "not linearly, not even exponentially, combinatorially": adding a 30th attribute to a wide event with 29 is worth more than all the others. With non-deterministic software, you know upfront that you can't predict its behavior, so capturing traces is essential.
Asked what this looks like in practice, she says auto-instrumentation has become very good, everyone should use OpenTelemetry, and because models are trained on its common patterns, it's now faster to build with instrumentation than without. She says instrumentation is how you declare intent and then check on it in production. She doesn't blame developers for not closing the DevOps loop, because it used to be prohibitively hard. For each piece of data you had to decide whether it was a metric, log, trace, exception, or profile; whether a metric was a counter or gauge; whether cardinality would blow it up; which log level to use. That could multiply coding time, and after deploying you still had to find the data and build dashboards. Now it can happen inside the development environment. Honeycomb has built features that suggest relevant data about code you've just written, with adjustable verbosity.
On spans, Majors describes a trace as structured logs with extra fields and a span as a piece of that trace's duration. The transaction has been the default building block since the start of the web, but she says that no longer works. Honeycomb has shipped a feature called Timeline on top of spans for agentic workflows. In a chat support scenario like Intercom's, a supervisor agent spawns more agents that call APIs and storage backends, and the whole interaction can span hours. Timeline acts as a "trace of traces" so you can zoom out and see it all. Without it, she says, you end up copying IDs between browser tabs. She looks forward to tests, telemetry, and evals converging.
On agents' limited context windows, she says a lot of traditional telemetry is "trash data" that fills context with noise, when what matters is the relationships in the data. She cites an AI SRE startup's writeup reporting that its agents often bypass typical observability data and go upstream for richer, intact telemetry. Her advice is to give agents the relationships, since those are what help AI make decisions.
Observability Engineering, second edition
Majors says the second edition is a complete rewrite, roughly 600 pages versus the first edition's 250. The first was written from 2019 to 2021 while the definition of observability was shifting, and she recalls finishing it with a feeling of "please take it, I hope that's enough." Now the definition feels more settled, even as everything else changes.
The book has six parts. Majors wrote parts one and six. Part one deals with running deterministic and non-deterministic systems. Parts two and three, by co-authors Liz, Austin, and George, cover instrumenting and understanding code, with parallel tracks with and without AI. Parts four and five contain guest chapters and deep dives, including one on CI/CD, one from ClickHouse on columnar storage, and one on iterative use of observability.
Part six was planned as three chapters for observability teams and grew to about 200 pages on governance and leadership. It opens with an open letter to CTOs arguing that their AI ambitions are blocked by their ability to make sense of their systems. It covers software delivery through systems theory ("if you like Donella Meadows"), quantifying observability's value to finance, and when to treat observability as a cost center versus an investment, which she says follows from the type of software being observed. It also includes a guest chapter from Rick Clark on how staff-plus engineers drive change without authority, a chapter on build versus buy versus open source, and one of her favorites, on vendor partnerships. There she argues that successful transformations depend on insiders with trust and credibility, and that with vendors, trust must be built through reciprocity. The best relationships feel like two teams at the same company. That is rare, and transactional relationships are fine, but she calls these "durable skills" for senior engineers in the AI era.
Leadership: be kind, and be good at business
The host quotes a recent Majors post: the most effective leaders are kind, caring humans and skilled business operators; next come terrible humans who are skilled operators; then everyone else. She wrote it about Twitter/X. She hedges on calling Elon Musk a skilled operator, but says Twitter had 16 years to figure out its business and visibly didn't, while X now runs with a fraction of the staff. She mentions roughly 30 engineers on the core product and about 60 in total versus 1,700 before. Her lesson: if we don't hold ourselves to high standards and efficiency, "someone will come and do it to us." She adds that when money tightened after the 2010s, many companies cancelled DEI programs, which she reads as showing they never believed in them. Her conclusion is that kind people usually outperform sociopaths in the same roles, but only if they are good at business.
On how engineering management changes, her first point is that everyone now gets to be hands-on. Leaders should know what it feels like to get a PR through, and it has never been easier to pick coding back up. Teams are getting smaller, which she thinks can be good if they can own more surface area. She worries that cuts are often driven by CEOs copying other companies or treating AI as magic. She dislikes the anti-management tone in the industry. She agrees that power drifts toward managers and has to be pushed back, and that bureaucracy breeds too many managers. But she considers middle management essential for sense-making and context-giving. She doesn't want engineers just handed Jira tickets, and people can't engage creatively without understanding, which is hard to build and fragile.
For middle managers struggling to find roles, her tactical advice is to go back to being an IC for a while, even if it isn't what they ultimately want, and to get AI experience on their resume. Working somewhere that doesn't build those skills is, she says, a massive career risk. The host adds that engineers with two or three years of AI engineering experience are in heavy demand, while people without any struggle to get the benefit of the doubt. Majors agrees the gap will widen the longer people wait.
For directors who fear how much tech has changed during a decade in management, she draws on her experience as a pianist. Anxiety and excitement feel almost the same physiologically, and the difference is agency. Waiting for the water to reach you means panicking, so run toward the waves. Going back to IC work is widely respected, she says; tell yourself you're excited even if it isn't true yet, and share what you learn.
Junior engineers and AI fatigue
On junior engineers, Majors says the hardest thing about quantifying their value is that nobody knows how to quantify any engineer's value. She is optimistic. Her friend Boris, who runs a new observability startup, talks to high school and college students and says they are "cooking," doing impressive work despite not knowing the software development life cycle. The kids will be okay, she says, and will reach conclusions others wouldn't have, but they need to be hired. The host shares a story from founders who hired an outstanding open-source contributor who turned out to be 17, and suggests internships as a lower-risk way to give juniors a start.
Asked about AI fatigue, Majors asks which kind. Some people mean receiving slop. Some mean the hype and what Cal Newport calls "doom trolling," AI company CEOs repeatedly suggesting catastrophe, which she considers irresponsible because it frightens people, including their families. Unlike earlier technology waves that promised to make life better, this one is framed around fear. She has also stepped back from social media because of AI slop posts.
Her response is to remember that we are in control. Because the frustration is universal, it's a good time for teams to propose experiments to take back control: no AI-generated PR descriptions, no AI on Wednesdays, a week off. As a leader, she says she'd prefer teams not ask permission at all and just report afterward what worked, what didn't, and what they learned, so others can learn too. Top-down permission doesn't work well when nobody knows what works yet. The host recalls the early smartphone era, when the best iOS engineers were often 18- or 19-year-olds, and a 22-year-old could be the staff engineer while a 40-year-old was entry level. In periods of big change, people can become experts quickly, and "no one knows" means nobody will say no.
Two book recommendations
Majors closes with two books. The first is Catastrophe Ethics by Travis Rieder, whom she describes as a bioethicist. It addresses the modern feeling that every choice implicates us in harm, whether cow's milk, almond milk, or soy, while problems feel too large for any decision to matter. As she describes it, Rieder shows that no traditional ethical framework offers a recipe that doesn't lead somewhere absurd. That doesn't mean everything is relative. It means living with integrity requires educating yourself about the world and then listening inward to decide what matters to you. She contrasts that introspection with the performative rage she is exhausted by.
The second is More Everything Forever by Adam Becker, a San Francisco journalist with a philosophy undergraduate degree and a PhD in astrophysics. She says it demolishes the "AI religion" of the singularity, effective altruism, accelerationism, and infinite growth. The one thing we know about exponential growth, she paraphrases, is that it must end, in an S-curve or a crash. She enjoyed Becker's dry humor about grand visions of colonizing star systems and about life-extension enthusiasts. What stuck with her is that underneath it all is humanity's oldest fear, the fear of death, and once you see it, you can't unsee it.
Can you measure an engineer's individual productivity?
There's been this whole push towards individual output, but teams are still what matter. Team output.
You mentioned how you're seeing two camps. There's the AI-pilled folks who get it, and then the people who seems like just hate AI.
See, the problem is that neither side is making it up. They are seeing really scary trends. They're grappling with real hard problems that are getting worse. What would it take for you to be comfortable shipping code without you reading it and understanding it? Cuz that's engineering.
How do you see the role of good skilled engineering managers and engineering directors change?
It's just easier now than it's ever been to pick it back up, fill in the blanks. And it's always been the case that
Why are there firmly two camps within software engineering when it comes to AI? Those hating its effects and those who are AI-pilled. Charity Majors emphasizes that both camps and thinks they are talking alongside one another. Charity is a co-founder and CTO of Honeycomb, previously worked at Facebook and Parse, and is one of my favorite voices in engineering. Today, we discuss what it would take for us engineers to ship code we have never read, and why this is more of a when question, not an if question. Why reliability is quietly getting worse across the industry, and why it will take some time to recover. Career advice in this age of AI. Why middle managers should consider going back to being an IC, and why junior engineers will be okay.
If you want to hear from someone who was skeptical about AI in 2025, but has changed her mind based on the evidence, this episode is for you. In today's episode, Charity will say, spoiler alert, that the question is not if we will stop reading code written by AI, but when. And we should take lessons from ops and QA and how they prove that software that others wrote works in prod.
And she's got a very good point. As any ops engineer or SRE will tell you, that's how software has always been written: by unreliable agents from their point of view. That is, software engineers like me, your colleagues, or you. And let's face it, you probably haven't read all the code in your code base, either.
This is where I need to mention our presenting sponsor, Antithesis. Antithesis verifies software written by unreliable agents. It runs your whole system in a hostile simulation and roots out the bugs for you. It does this by using an approach called deterministic simulation testing, or DST. Antithesis turbocharges testing by running your whole system under aggressive fault injection. Imagine Antithesis' hundreds or thousands of versions of the Mario game running, each instance aggressively trying to break the game with increasingly weird input combinations.
If it finds a breakage, this is where the determinism comes in. Instead of you having to try to reproduce a tricky bug you saw in production, Antithesis can provide you with the perfect deterministic replay of anything it finds every time. With Antithesis, you can specify properties at the whole system level, and Antithesis will actively try to disprove them. So, you can be confident that if your system holds up in Antithesis, it will hold up in production. Head over to antithesis.com/pragmatic to learn more.
Charity, it's so nice to do this in person.
You're in my city. This is amazing.
So, I want to kick off with AI, but before we kick off with AI, I just want to make it kind of clear for people who don't know you that, you know, you're not an AI hater or an AI lover. You actually built a lot of cool stuff pre-AI, right? Starting at... We just saw Linden Lab. Was that your first job?
My first job at Linden Lab, right across the street.
The street. We're just talking about that. So, you were building Second Life?
Yeah, we were building Second Life. Yeah.
And then from there on, one of the big hits was Parse, the developer tool, which was beloved by developers, back end for all, best back end for mobile services. I used to use it. And then what happened? Facebook bought you.
Facebook bought it. Yeah, it was my first great lesson in: most acquisitions fail. Most of them are terrible. This one failed. They shut it down. But ultimately, I'm very grateful to have had the experience because if it wasn't for that... I've always been a startup kid and so nobody knew my name, and it wasn't until I was leaving Facebook that investors were like, "Oh, would you like some money?" And that's how we started Honeycomb.
And then you saw stuff at Facebook, right? It inspired you that
Yeah. Facebook, there was a tool called Scuba. And so we were in a weird position. We were building on AWS, Ruby on Rails, all this stuff. And then we got to use the internal Facebook tools. And Facebook had this tool called Scuba, and it was... We were experiencing hockey stick growth. It was just like we had over a million mobile apps hosted on Parse by the time I left. Yeah. And every single week a new one would break. It would hit the top 10 on iTunes or something out of nowhere. I'd just be like, "Oh..." And these apps... needle in a haystack, you know, and it went from it would take hours or weeks, we'd have to get lucky. We'd finally find... cuz it might be one app that's spamming the logs, but that might not be the reason. They might all be backed up behind the reason, you know.
We started getting our data sets into Scuba and finding them, it just went from being a really hard engineering problem with a lot of luck to just being like poor problem. Click, click, click. Oh, there it is. And it was just mind-blowing. Like, that was a huge problem for our entire existence and then it was solved.
With Scuba. And then when you started Honeycomb, so was this a bit of inspiration that you wanted to build something that feels like Scuba did?
I just, the idea of... I was planning to go be an engineering manager and engineer at Slack or Stripe or something. And I was just like, ooh. I would be so much less powerful as an engineer without this. And so, you know, the grand plan in the beginning, I'm just like, well, all startups fail. So, you know, we'll fail, but I'll go sit in the corner and write Go code for a year or two and then I'll open source it and I can take it with me wherever I go.
And then that's how Honeycomb started.
That's how Honeycomb started.
And we'll get back on Honeycomb, but before we do, now with AI, you know, it's changing everything.
Yep.
But I kind of had a bit of a blast from the past, which is one of the first places we connected was in 2020, so almost 5 years ago or so. Someone submitted a question to both my blog and your blog. And the question was like, "Can you measure individual developer productivity?" Now, I wrote an answer and you wrote an answer. And I wanted to ask you, that was 5 years ago, no AI, no nothing. Today, someone should you a question saying, "Hey, engineer's individual productivity, you know, they're using AI tools and all the stuff?" What would you tell them?
I would tell... God, I don't even remember what I said. I remember that blog post, but
We both agreed, by the way, that you can measure some dimensions and they're not going to give you the full thing, and they will, for example, not tell you how a team is doing if someone is actually a really key part of the team. And that as long as you measure individual things, we both agreed that you need to be in the details to know. And as a good manager or a good team lead, you will know.
You will know, but you have to have data to back it up. It's like color in a painting on the wall. And is it Goodhart's law? Yes, it's Goodhart's law. It's so like never go well, it's this thing that matters, right? You need to actually understand, but you need it to not just be your opinion that was tossed off because you have an opinion about some... you know, we all have biases, we are selective, you know, you need to look at the picture. I also believe that, you know, there's been this whole push towards individual output, but teams are still what matter. Teams output. And honestly, if there's one thing that I am encouraged and excited about with the AI movement, I think it's forcing us all to ask ourselves early and often, "What does good look like? What does good mean? What does productivity mean? What would better look like? What would great look like?" You know, and these questions are hard. I think it's telling that we all jumped so fast to speed.
Yep.
Oh, fast. We can go fast. Let's do same thing faster, you know, just boom. And I've come to feel like that is a very immature description of what better is.
Yeah, just today I saw the Anthropic team posted a podcast with Spotify's head of engineering or VP of engineering, I'm not sure which one it is, in which they talked that with Claude Code they're shipping 4,500 changes per day, per week, I'm not sure which one, but they talked about speed. And I was kind of thinking, like, my experience has been different cuz I struggled to publish... some of my episodes did not go on Spotify cuz it was down.
Yeah.
And yeah, they're talking about speed, but we're not talking about quality. We're not talking about more functionality, better functionality, or just things that people want. And in a comment some people are asking like, "Okay, so what exactly does that mean that they're shipping more frequently?"
Yeah. Do customers really want the buttons on their app to move around all the time? I don't think they do.
Yeah, that's an interesting one.
Thing to measure.
Let's jump back to last year, in 2025. You wrote a blog post right at the end of the year looking back, saying that 2025 for AI was what 2010 was for the cloud. Can we talk about that before we go into this year? Like last year, how was your perspective? Of course, you're working at an observability company, AI will give you lots of business as well, but you said it went mainstream, right, last year?
Yeah. In March 2025, Fred Hebert and I gave a keynote at SREcon. We gave the closing talk. And it's Fred and I standing in front of a... The term vibe coding had just been invented.
Oh, yes.
And we were like, you guys should try vibe coding. Pause. Groans. Audible groans. Just like people laughing, like ha ha ha. Our big pitch was that people should learn AI because you can complain better if you learn it. Which is legit. I mean, I really mean it. But at the time I think I still saw it as a really big feature, or like bigger than a programming language. Like the cloud, but not generational, you know, not changing everything.
And I think that was accurate. For me it was November of 2025 when they released Opus 4.5. But I actually wrote about this recently in a couple blog posts back, about how in retrospect, you could see it coming sooner. And it wasn't actually the models. It was the harnesses. It was all the tooling. And it was people starting to say that around July. They were like, this is coming faster than you think and this is what it's going to look like.
And this was people were saying at the word once you were playing with either Claude Code or maybe Pi or OpenCode. So, the harnesses, you're right.
They were getting better at the tooling. You know, it went from just being kind of a shell script that would try again to, like, they built a lot of stuff around it. And then, you know, the Opus thing kind of... It was a weird time at the beginning of... It's been a weird time every time for a long time, but the early months of this year it felt like everyone around me was just trying it again and changing their minds. Everyone.
Yeah. But I think what we just talked about right before we started recording, that both you and me respect people who do change their mind.
Yes. And I don't think we were wrong to be skeptical the first time. It's a pretty extraordinary claim that AI is going to write code about as well as the median software engineer can, and, you know, for limited bounds of that.
Well, especially cuz if we look back at the history of software engineering, this claim has happened again and again.
Yes.
You know, neural nets should have been doing something magical.
There's a sticker in your pack that says, "We already have a programming language that lets you..." COBOL is the punchline. So, I don't think we were wrong to be skeptical.
And also don't forget no-code and low-code.
Oh, yeah.
I mean, we know it turned out to be a joke, but the promise was the same. And we were skeptical, and we were right. And now we're skeptical again.
And we're wrong.
And we're wrong.
What I was saying in that piece, though, was I think we were right to be skeptical that time. But now, I see the same thing playing out with: would you be willing to ship some code you didn't read? There's no point in arguing about if it will happen or when it will happen. Talk about what it would take.
What would it take for you to be comfortable shipping code without you reading it and understanding it? Cuz that is
That's engineering.
Kind of, it goes back to, like, you know, my gut reflex would have been saying, "Oh, no, I would not do that because I've been used to that." However, you're right, you know, if I could have a way to, for example, see the change, I could tell that this was tested in a harness or something. Same way where, for example, pre-AI, if I had a team member who said, "I vouch for this and I've hammered it," and I trust that person. So, you're right. There's these things which of course would never... but I thought that AI or something can do anything like that, but if it could, and that's engineering, right?
For example, if you and the AI would both do it in tandem for a few months, and you would get to: how much are they catching? How much am I catching? Is it about the same? Is it more? Is it less? And you're training it and it's getting better. Whether it takes 5 days or 5 years or whatever, I think it's pretty clear that directionally that's where we're going. And the other thing that I would say is this is good for us. If you spend much time with the Phoenix architecture stuff that Chad Fowler has been writing about
You have been quoting, yeah.
I've been quoting liberally. I should probably let you get to it in your own order, but I just feel like anyone who's ever done a painful rewrite should be on board with us.
Yeah, and here's a quote from Chad Fowler: "Immutable infrastructure, stateless services, containers, blue-green deployments, infrastructure as code. These ideas all share a common premise: never fix a running thing, replace it. AI pushes this premise beyond infrastructure and into application code itself. When rewriting is cheap, editing in place becomes risky. Mutation accumulates entropy, replacement resets it."
Yes, code is cash.
This is a very interesting idea because Chad compared, and you've also of course shared this, that when we look at how infrastructure changed before, you know, like specifically a server, you need to configure it. I think we call it like pet
Pets versus servers.
Having pets versus
Yeah.
And at some point we stopped configuring individually, we stopped fixing individual machines, we just throw it away and have a new thing. And with code, the history of the profession, you know, 60-plus years or maybe a bit longer, is that we edit code. And are you thinking this might
Because of the economics of it. I mean, if you think about it, you could generate 10,000 variants of a function faster than you could write it once. And so when you start thinking about it that way, it's like, well, okay, we're going to need a lot of evals, we're going to need a lot of tests.
But the generation is so cheap that it really, I think, forces us in that direction. And I think that the expensiveness of writing code and maintaining code and the expense of software has always been bound up in its maintenance. And those lines of code, the reason that we trust something is because we've been using it, because we know, like there's this deep thing about production, it's like, well, it's trusted, we know.
And I know that as well as anyone. And I will also say this: anyone who's ever done a hard database migration should have some real humility about our ability to extrapolate those contracts, store them. Like, I am not one of the people who's like, this is good, we're going to generate our code. I don't know how much code. I believe that we can go some distance in that direction and it will be good for us. I don't know how far we can go. I believe we can go farther than we are now.
I just, man, the last project I did at Parse, we had spent like six months writing the original Ruby on Rails API.
Yep.
Spent two years rewriting it in GoLang.
Wow.
Yeah, it was, it was...
And was it two years because new stuff kept being added that you needed to pull?
And also GoLang was a pretty immature language at the time. We had to write, you know, the [ __ ] DB drivers and all the other bunch of things. And also just, like, when you're writing in Ruby and [ __ ] DB and JavaScript and everything, there's no type safety and it's just painful.
You know, in the strangler figs that they do, you rebuild the architecture outside the architecture and you literally find the contracts with your users by breaking them, one after the other. Like, that just does not seem like the ideal artifact. We should be able to store them somewhere. We should be able to have architecture diagrams that we can review and discuss that generate that code to spec.
This is very interesting because some of these ideas, they've been around decades ago. Specifically, you know, if we had Grady Booch as a third person sitting here, the idea of like, "Hey, we can have architecture diagrams that translate to code." UML started there. I think Grady would disagree that he never wanted it to go there, but Rational Software back in the '90s, they said, "Hey, you'll define UML, it generates code, it will be beautiful." Now, it wasn't beautiful because, I guess, some complexity, and it turns out generating code was still expensive, and reviewing it. But I wonder if some of these ideas now might be just feasible.
That's my hope. That's my hope. I mean, I'm just barely old enough that my first job, I was like 17, university, I was a system admin. I remember when, you know, I wasn't really aware of what was going on, I was just a kid, but yeah, I remember how stressful it was and how people were agonizing about how we'll never be able to get that information back, and everyone adapted just fine.
And I think I got the systems that, you know, they built the systems that replaced them, but not as in replaced them and worked them out of a job. They built the systems and they spent their time writing code instead of, like, running updates by hand on every server in the closet.
And I guess this is an interesting one because clearly this sysadmin role and profession, it doesn't exist today. Let's just say it has been eliminated. However, the people who were sysadmins, they did understand the operating systems, they understood hardware.
Yes.
They were in a really good position to adapt, and a lot of them just became either software engineers or product managers. I know someone who became a tech salesperson.
Yeah. Yeah.
So, it's almost like...
And I will hold that our generation of engineers are still the best debuggers. I'm glad that people don't all have to learn about CPU and memory and all this stuff, but there's value in knowing that stuff. It comes in handy. I think there's some analogies there to the generative code stuff.
Also, you know, you took a bunch of inspiration in your recent writing from, well, sysadmins, but also QA. And you wrote something interesting. You said, "The lines of code are not the ideal artifact to review." And I'll quote a little bit from you: "The tools to do this don't exist yet, but many of the ideas do exist. Most come from operations and QA, two domains that software engineering has historically been rather snobbish about." Should we revisit our relationship to QA and ops? I feel we always put ourselves as software engineers here, and ops and QA somewhere. And maybe it's time to eat some humble pie?
Ops equals toil, right? Yeah, I think it's time. I mean, ops and QA have always been more concerned with what is. Software engineering has always been much more concerned with how should it be.
Mhm. So, ops and QA have always been more concerned about validating, about correctness, about does it work as expected?
Does it work to start with?
Yeah. I mean, it's always weird to me just how much software engineers really seem to believe that the world exists in the repo. It doesn't. It's production, you know? The code has part of the information. Some of it is very necessary, we need that, but like, I know some software engineers who... well, I don't know. Okay, some places don't even let software engineers look at production. Just, like, how?
I know a lot of people are very upset about AI, but the things that get me very excited, genuinely excited, about AI are that it is pushing the discipline in directions we have desperately needed to go for a very long time. Production is not what happens after development. It is a stage of development.
And you've been saying this consistently pre-AI. I'm just going to say this for those who don't know, because I remember we also bonded a little bit over this. There was this thing trending on Twitter when it was still Twitter, and it was tech Twitter, everyone was there who mattered. And there was a trend going: it's Friday, don't deploy. There was maybe a hashtag even, I'm not sure what it was, don't deploy Friday or something like that. And the point was, it was well-meaning. It said, look, when you deploy, often there's an outage, and on the weekend we don't want that. So every Friday it went viral, saying don't deploy on Fridays.
And you came in and you said, "You know what? You should be able to deploy anytime without fear, because you should be able to just know," you know, however that might be, CI/CD. And then on top of this you were like, "No, you should actually just not even have a user acceptance testing environment, a UAT. You should just deploy to production and test in production, right?"
As soon as you merge, it should be going out. You should have to stop the train to make your code not go into production as soon as you've merged. Absolutely.
And one more interesting thing is you had a long train of thought about AI and what it could be. One thing you said is our brains are not built for validation. Almost everyone I talked to, including Andreas Heimburg, said that, look, it's very clear that code generation is cheap. We are generating more code, and the bottleneck for human engineers is code review. And everyone is trying to figure out how do we make code review easier, how do we build nicer tools. Uber has built amazing tools to try to surface important code reviews. But everyone is pushing like, all right, let's do more code review.
As an engineer, I'll be honest, I never liked doing a code review. When there's very little to do and it's with someone I care about, I'll entertain it.
It's more of a coaching opportunity then, right?
But as soon as there's an AI, it's kind of like, I don't know, I don't really care. I'm just being honest here. Do you care when...
I've never... So, one of the problems is that I think code review means so many things to so many people in so many places. And so there's a lot of projection going on. If you say that you don't want code review, you're saying you don't want to talk to your coworkers, you don't want to mentor juniors, you don't want to, you know, which is not true. We just bundled so many things into this, you know, it's like...
Usually overloaded.
Usually overloaded. And some of those things are really good. Some of those things could be done better in other ways, you know? Some of those things are very cultural, very specific. My friend David Poll, who I worked with at Parse, he's now working at GitHub on pull requests...
Amazing.
...and code review. The Parse mafia. Yeah, exactly. He's like, "To me, the code review is when we decide, do we want this in our product or not?" I'm like, "Well, that's a great discussion. That is what humans are good at. We should talk about: is this mental model coherent? Should we add this? Should we not?" Like, love that. Architectures, you know?
But the code is not necessarily a great artifact for all of those. So, should we be talking to people? Yes. Is the code review the right form factor? Maybe, but I think the emotional reaction is when people are getting to that. The validation, in my book, is at the very bottom of the list.
I'd like to stay here a little bit more. Can you break out the parts? Because it feels to me code review is overloaded, but the parts of code review are the things that you have seen are good things, and maybe we don't need to ask code review, and the things that have just never been that good, and maybe we just need to throw them away.
Yeah, I mean, I think "do we want this in our product" is great. Ideally you'd talk about that before you write the code for it, but you know, whatever. And, you know, is this API design, you know, those are great conversations. Reading for syntax and bugs and that sort of thing, it's not evil, but it's a teaching opportunity if that's the best teaching opportunity you have, and I guess some folks at some point maybe need them, but it doesn't feel like a great use of anyone's time.
It feels the only time where it's useful is if someone joins a team, and initially...
Yeah.
...give a little bit of feedback.
Yeah.
Especially when nothing is written down, there's no guidance, there's no linting rules that would give you that.
Well, see, that's again, yes, we can fill in the cracks if we haven't built the guardrails. We can fill in the cracks all kinds of ways with our own time, but there are so many things, I think, that we never think to extract out of the process of building and validating software. So we rely on us.
So, I am a huge fan of Intercom, you know, Fin, their engineering org, and I have been forever. I noticed their CTO a decade ago had this saying, "Shipping is your company's heartbeat." And I love that. They ship a Ruby monolith in like 10, 15 minutes, hundreds of times a day. That is not trivial. It's not a trivial thing to do, right? So they're kind of a high watermark in my mind right now for teams that were founded pre-AI, have a lot of engineering discipline, and have become AI native.
And they wrote a great post about how they do PRs that are AI validated, and the bar for them is very high. It's like they have all the wisdom of their most senior engineers looking at every single diff, and that is fantastic. Which means that you don't have to worry about remembering and looking and nitpicking and all the things that we're not good at anyway, and they can talk about: is this the direction we're going? Is this the right path?
You've also written that non-deterministic systems require more engineering discipline, not less. So, what is the thing about these non-deterministic systems, or specifically AI, right? We're talking about AI, let's just name it. We see that AI does amplify both discipline and lack of discipline. Why do we need more? And when you say discipline, what specifics are we talking about?
Well, I mean, tests and evals for one thing, right? If we're treating the code like a trusted artifact and we're trying to predict everything with our human brains, then we're writing the tests for what we can predict might break, you know? And then anytime the system breaks, we try and write a test for that, but that is not an especially high bar.
And so I think the sort of behavioral tests, or the... I don't remember the word. It starts with C, but the QA folks have these tweet tests where it captures...
Smoke tests. Yeah, there's so many. There can be performance tests. There can be load tests. There can be just kind of fuzz testing as well.
Something that's like, okay, if I'm not going to read this code, how do I know it's going to perform within the boundaries of the last code that I generated? That is conformance testing.
Conformance testing.
Just as important for lots of workloads as absolute performance is it just not changing too much. And so I think we're going to need the trust to go somewhere, right? If you're debiting from this trust account in the creation of the code, it has to get built up somewhere else.
And I feel like one of the things that I'm really excited about in the coming months is, I actually really like thinking about it less as AI and more as deterministic and non-deterministic systems that have to play nicely together, because determinism is not going anywhere. It's incredibly valuable. And we have to learn to make AI kind of boring, you know? It's a non-deterministic tool, which means that it is all over the place, but it's so valuable, but it's all over the place. So we have to learn how to give it carved pathways, places where we kind of corral it, where we use it in the way that it's a superpower and not in a way that erodes our foundations.
This is interesting because Martin Fowler, who was on the podcast a year ago, talked about how the biggest change with AI is non-determinism. And when we think back in the history of software, it's always been deterministic. Same for neural nets, but most of us software engineers didn't really touch too much of it, because it just wasn't that useful for us. But we've been used to that: when we programmed it, it just happened the same way. Unit tests were easy because you just run them once. You don't really run them twice, because why would you?
And I wonder if we need to just realize how big of a deal this change is, and that any business that employs us, they want software that works the same way. We just had a recent post on Hacker News. There's this ATS, applicant tracking system, scoring system that HackerRank outsourced, where it scores your resume. And you can run it locally. It's open source. You use a local model; I think they recommend Gemma, Google's small model. And when you run it like 100 times, it will score the same resume anywhere from like 66 points to 99 points. And typically, most companies have 85 set as the bar. And you're like, "Hang on. So what they were advertising as a tool to help your recruitment, we just proved that it's just a coin flip." That's bad.
Yeah, and we have to be able to say that it's bad. AI is not the right tool for every use case, you know? And I think every company is going through this in microcosm. Something I was saying to folks just earlier today: we've been doing this series of conversations on our AI norms and values. And it was like a year ago, I don't trust us. Like a year ago, if we were like, yes, we should use AI, and now we should, we didn't know we...
didn't know enough. We've gone on such a journey over the past year and we know so much more now, but like if one of my co-workers is like, AI is the wrong tool for this job, I'm like, I trust you. You know, you got to get worse before you can get better.
So tell me about where you are right now inside of Honeycomb, how you're thinking about AI, how you're thinking about how to think about AI, and then what values you came up with that work right now for you.
Yeah, it starts with just acknowledging that the bar has gone up for all of us. That's what happens when we get powerful new tools.
Has the bar gone up or has the, you know, the floor gone up?
That is a great question. Maybe yes, and maybe, yeah, I don't know. We're definitely in a sort of wandering in the wilderness phase. But you can't not wander or you will be left behind. You know, we acknowledge that the bar is going up for all of us and that the only viable way to define that bar is better outcomes. And asking ourselves, like, is this good? Is this better? What does good look like? Another thing we point out is just there is no human in the loop. You own the loop. The loop is yours. The loop is mine. It would not exist if it was not for me. So I am the owner, right? There's no "oh, Claude said this," so no, no, no, it's your work. You own it.
Charity just talked about owning the loop. Owning the loop also means controlling what every agent inside of that loop is allowed to do, which brings us to our season sponsor WorkOS. Today agents are increasingly able to act on their own and the old auth model was never designed for that. Who is this agent? What's it allowed to touch? On whose behalf? You really don't want to get answers to these questions wrong. WorkOS is built exactly to solve this problem. WorkOS's fine-grained authorization FGA, designed for how agents actually operate, plus SSO and SCIM, and not just user auth with agents bolted on after. The fastest-growing AI companies, Anthropic, OpenAI, Cursor, Perplexity, already trust WorkOS.
Check it out at workos.com. I also want to talk about Buildkite, the CI orchestration platform trusted by Cursor, OpenAI, Anthropic, Nvidia, Uber, Canva, and more. Charity talked about owning the loop, but here's a challenge. Thanks to AI, your agents are writing a lot more code. To trust this code, every change that an agent makes still has to be built, tested, and proven safe before it ships. So, obviously, you need CI more than ever. But, when agents are pushing 5, 10, or 50 times the commit volume to your pipelines, faster CI runners won't be enough to keep up with it. Shaving 30 seconds off a single build is meaningless when the queue is 100-plus jobs deep. What you really want is a CI system that gets faster as the volume grows, and CI that offers instant parallelization to give you unlimited concurrency, and to intelligently route changes at runtime. This is what Buildkite does, and why global software leaders at every level continue to rely on it. The same architecture that absorbed the scale of Shopify and Uber a decade ago now runs about 1.4 billion job minutes a week across Cursor, Meta, Reddit, and Snowflake. While the rest of the CI world are cracking under the weight or re-architecting their platform, Buildkite continues to reliably grow. Agents run on your infrastructure or on Buildkite's. Any cloud, any chip, your secrets, your scale. Every artifact and log is captured, so when something fails, either you or your agents have immediate insight for why. As you're engineering the context you give to your agents, think about how you'll verify what they hand back. If your system is buckling under the increased volume, head to buildkite.com/pragmatic. 30-day all-access trial, no credit card, and an actual human engineer on standby. His name's Ola and he's very helpful. And with this, let's get back to Charity and communication norms with AI.
I think there was this frenzy of, "Oh my god, I could do this. Oh my god, it's so cool." And I know you have also become very weary of this slop. I just don't even read it anymore. As soon as I can tell, as soon as I recognize this might have been AI, it's like trash.
Here's a baseline. You cannot send anyone something you haven't read. And in fact, if it would take them longer to read it than it took you to make it, it's probably slop. That's really disrespectful, actually. And I think like just like asking someone to, like, you're asking... Anytime I give you something, I'm asking for your time and attention. And if I'm giving you something that I don't even know what's in it and I'm putting it on you, it costs you instead of me, that is not good.
I also think that even before that, it's like I've noticed as I start working on these norms and values, I'm noticing myself as I start to ask someone a question without trying to look up the answer. Ooh, I shouldn't do that. Or if I'm giving someone something that I kind of generated and I'm like, "Ooh, you know." A part of it is just self-awareness.
It's interesting because everything you talked about reminds me of when a new joiner would join a team, a junior engineer, a new grad. Either they had emotional intelligence or they picked up on really quickly that, for example, you go and ask a senior of their time once you put in a little bit of work and you start to respect their time as well. And obviously, it doesn't start like that. We don't want them, but there's this balance. And I almost feel it's the same thing. We're like, "Look, respect your colleagues, respect fellow humans. If you are communicating with them, make sure that you're not wasting their attention," cuz now I guess attention is where we're kind of running low. Like we have all of these, like a lot of people have a bunch of agents doing, but the point is that's kind of the currency. And as long as you respect that, it doesn't matter. Like I think we're not talking about don't use AI for this or that. Like you use it as much as you want to make yourself more efficient. Just don't degrade, cuz it really degrades those personal skills, right?
You can use AI as a shortcut to help you not have to think too much. And you can use AI to help you think more deeply and more rigorously. And both of those use cases have their place. But when it comes to your core job function, we primarily want the second one, right? And especially if you're involving someone else and you're asking them to review or...
You know, and this is not absolutist. Like there are people who English is a second language and they use it. People who are neurodivergent and use it, and that is, again, that is still being respectful, you know? So that's not, like you said, it's not no AI, but it's like make reasonable asks of each other and, you know, we don't need to reinvent a new bar for quality or respect because we have great bars already for quality and respect. We just need to apply them. For a while there, I think that there was a bit of oh my god, this is so cool. Do you see what this cool thing can do? And I think we're all just like so over it.
The reason I really respected you came from the sys, you know, the sys dev background. You're also very well in SRE. These are all folks who have been pretty skeptical of AI. And you mentioned how you're seeing two camps. Two very clear camps. There's kind of like the AI-pilled folks who get it, and then the people who seem like they just hate AI. And you said that you're not seeing these two camps have any sort of way to go between, any feedback. Can we talk about what you're seeing and, you know, where you see some of these camps forming?
See, the problem is that neither side is making it up. Like they are seeing really scary trends. They're grappling with real hard problems that are getting worse. You know, and on the enthusiast side, they're acutely conscious that it's a bit of a race and that we need to push ourselves out of our comfort zone, and they see other companies moving faster, catching up, leapfrogging. They're really worried about, you know, we're falling behind.
Right. And the first thing, I don't want to make it sound like false equivalence because
well, there are elements of this that are true. I think every company is more one or more the other. But they're not wrong. They're not wrong. We've never seen technological change this fast. We're on the inside of an exponential curve, which is very rare and it never usually lasts that long, but it's still happening, you know? Things are happening that shock us and we would be wise to prepare for them. So like that's real. That's real. And these folks are usually, at most companies, usually they are the small minority and they are constantly feeling outmanned.
One of the things that's ironic though is that both of these sides feel like they are the tiny minority and they're outmanned and they're being suppressed and they are standing up for what is truth and valor in the face of the big AI folks or the big skeptics. But the other side... So this often starts to come down to the group that is on call and the group that is not.
Ooh, yep.
Because the people who the buck stops with, they are seeing melting mental models, they're seeing slop, they're seeing all their hard work just dissolve, and they don't see any end in sight and they're
Just to be clear, we're saying that the people who are on call for a lot of these systems are seeing more incidents, they're seeing carelessness being caused by it. They're actually seeing that since that group started to use more AI, our systems are getting way worse.
Way worse.
Yeah. And that's very real and not making it up.
No, no, no. Actually, I was just talking to someone inside of Meta. There's been this big drama where people have been resigning.
post.
So, not just my post, since then, I haven't written about this since, I'm not sure when this podcast comes out, I might have not talked about it. Inside of Meta they track SEV zeros, which is the highest severity. Well, you remember SEV zeros. There has been a flurry of SEV zeros, so many of them. And you cannot hide. Like, you know Meta. Like this is black or white. And the past about 2 months, it's been crazy. And just so it happens, it's happening inside of Instagram, it's happening inside of WhatsApp, where the trust and safety, basically the reliability folks have been axed, removed. So, it's impossible to deny the connection as well. Of course, it's not a direct one, and each one has a postmortem, but Meta has not had it this bad for closer to a decade.
Yeah. Move fast and break things.
You put two plus two together. And when I told this story at a conference, people came up to me and they said, "I'm so glad you talked about this cuz my company," different company, often VC funded or publicly traded, "the same thing is happening." People are like whispering to me, like, we are not Meta, but the same thing is happening.
Same thing is happening.
And you know what they all told me? They told me, "I thought it's just us." Or I thought it's us and then my buddy who works at this other company. And suddenly it was like, "Oh, it's all of us."
No, it's all of us. Yeah. No, it's a real thing. And the Intercom folks, you know, what I love about them is they published the real gnarly stuff, right?
Yeah, they don't color it out.
Color it out. And they showed that for 18 months, reliability and code quality went down. And it had just started to possibly be going back up.
But it's still not there where it was. And they're honest about that.
it. Finally.
So what? This is the thing. Like stop spitting in my... and telling me that, you know, like it's just... This is my thing. It's like we need to hear the wins. We need to hear about what's possible. We need to hear what's exciting. But you got to couple it with the costs. You got to couple it with: is it worth it? You got to couple it with: what are we doing? What is happening? And I feel like part of the reason that both of these sides are getting so frustrated is because they're not connecting it all. And so the people who are seeing really incredible... There are some really incredible things happening in software right now. Like with rewrites and with, you know, automating away real toil and [ __ ]. Like not a single person that I've talked to would give it up. It's amazing. Like, I get so excited. Nobody wants to take it away.
But half of the people are seeing the wins. And they're not connecting it to the cost, which makes them think that their coworkers are just [ __ ] nuts who are just like, "Oh, they're so going to lose their jobs. They're just afraid of getting automated out of existence. They're just blah blah blah blah blah." Like, "No, dude. You be on call and then see how you feel."
You know, and there's a mirror effect kind of happening where the folks who are on call, who are responsible for this stuff, they don't actually believe that these wins are real. They think they're all cooked because they're not hearing the quiet parts said out loud that, "Yeah, we're seeing this win, but this is what it cost. We're still cleaning this up. We're still..." And so that's my beg to everyone who loves Gergely's podcast and listens to this: tell the whole story. Talk about the costs. We're all in it together.
Yeah, cuz you're right. Like this technology is not going anywhere. It will make really big positive change at a bunch of places. It's here.
But it's not magic.
It's not magic. And I think this is what you said in the "make AI boring" and another great article of yours. What you said is AI is just technology.
It's just technology.
And you were arguing that let's just realize it's technology, it's a tool, and let's learn to use it well. Now, one other thing you said, which is very interesting, is software will be the killer app with AI.
Yeah, I think
Which is very unique. Let's talk a little bit about that.
Software is made of logic and language. AI is made of logic and language. And because of that, we can bake in guardrails. We can bake in checks. We can bake in validation that we... I don't know how we do that in other parts of our lives or other applications. And so it totally makes sense to me that software is what AI is best at. I mean, you see like in the courts they're starting to get lawsuits for... The court is suing lawyers who are submitting briefs that have hallucinated crap in them. How do you check for that? You know, with the same... We have structured data. We have, you know, a whole... And I just don't know how you account for that in the same way.
It might also mean that whatever will work outside of the software industry for AI, it will be a subset of what will work in the second... Basically, if we can do something with AI, if we can automate a process or something, you might be able to do it in other industries, but maybe not. But if we cannot do it, good luck. You will not be able to do it because we have the domain where you can validate stuff. We have incredible training data on code that compiles, right?
Yes. Yes.
Like in a bunch of places you might have training data, like with magazines you might have low-quality magazines or whatnot. See what I mean.
I mean, back to your point about humans, like their determinism, they like things to happen the same way.
And it's very interesting because as I think of it, you know, one of my businesses is writing. I write a newsletter that is, I like to think it's good and it's worth reading.
It is.
And I would have said, if you asked me, "What is AI really good at?" Now, obviously it's good at coding, but before that it was good at writing. My mind was blown that it can actually control the language. When all the newer models come out, I do this
test where I say like, "All right, like, you know, write an article in the style of the Pragmatic Engineer." And every single time I can tell it's AI-generated because it's repetitive, it has the same... So, my point is AI is actually not as good at writing prose; it's a lot better at writing code.
Way better at writing
When I ask it to write code, often I'm like, "Yeah, this is something I could have written." Whereas when I ask it to write words, I'm like, "I would have never written this." And it has training data on me. So, who knows? This might prove that software is the best fit.
I think it is. Software is a simplified version of language for a purpose. Yeah, you know, at first everybody was trying to come up with ways to be more efficient and write with AI and everything. And I sunk a lot of cycles into that. And I have decided not to sink anymore because writing is thinking on paper. And there's no shortcut for doing that thinking. Anything that I write is not content, you know? It's not content where it's just like, "Well, generate me a couple thousand words." Which I'm not shaming anyone who generates content, but that's not what I'm trying to do. I'm trying to think through hard and interesting problems and share them with people. And I don't think AI is the appropriate tool to use for that. I use it for structure. I'll be like, "Hey, read this and give me feedback," and stuff. But
You know, so I think we should not forget that as we improve our skills, our capability, our experience, our thoughts, we do become more valuable.
And I have this idea, and this might be a flawed idea, but I think it'll be correct that, you know, 5 years from now, how will people be hired? Now, of course, we know the tools will be better and all that, but in the end, I think it'll be like this. Someone sitting here, and I'm going to be interviewing with you. I'm going to be trying to get into your company, probably Honeycomb, right? And we will be having a conversation, and you will judge me based on how I respond, and the more I have spent thinking and bettering myself, the more valuable I will be to you because you will have all these candidates, and some of them will have outsourced other things to AI, and they will have a blank because that thing is off. Guess who you will want to work with, right?
I am so excited about leaning into the parts of being human together.
I don't like the feeling of chatting all day back and forth between agents and people on Slack. Like, it feels way too similar. It's just gross. Honeycomb is a fully distributed company, which was never... We always wanted to have a hybrid model, but the office has not come back, and I feel all kinds of ways about this because I love not leaving the house, but at the same time, I crave this more full... Like, I'm so glad you're here. It's so nice to see you.
We were just talking how it is different. We've done a podcast remote, and it was a decent one, but this is more enjoyable.
Yes. And so, part of what I hope we do is just remember that we're in charge of the machines. They serve us, and this is still what matters.
Now, I want to pull back to something different. I'd like to talk a bit more about ops and DevOps and then give one of your spicy takes. So, now that we have AI, we can actually just, you know, bad-mouth some of the other things or just be real. Let's talk about DevOps. Can we go back a little bit in time? You were there. Why was it created? And in the end, there was this massive DevOps movement in the 2010s. Do you think it succeeded? Do you think it failed?
So, before DevOps, we needed a DevOps because there was devs and ops. And there was the proverbial wall that code got thrown over, right?
And ops were the people who were in charge of the IT. They deployed, managed the servers. They set the Linux version.
Linux
Yeah.
pluggable storage models and everything. That was always a bad idea because it's split brain. Half of you are writing the code and the other half are understanding it. I would argue that you can't really understand the code you write unless you're operating it. So, you know, the DevOps movement did a lot of good trying to knit back together that sort of original sin. And you know, around the time that I was a sysadmin, there was this big push: all right, ops people, learn to code. And great, I'm glad that happened. Everyone who works with computers should be writing code.
I feel like the wave after that was a little less successful, which is like, okay, software engineers, time to learn to understand your code in production. But I also think that in my mind, 20 years of DevOps was really about one thing. Trying to create one feedback loop that connected people writing code to that code in production. And it failed. I mean, it failed. To this day, they're done by... they're two different domains, you know? There are some people who... I mean, it's
And I'll show you this diagram that you drew. We're now out of the agency, so we'll put it on the screen so viewers can see it. But this is your... I think it's a really nice drawing of how there's no feedback loop. Like the ops people, or oftentimes we call them platform teams, they manage the infra layer. Engineers deploy there.
And so, to be clear, I think that's actually good and fine and healthy. I think that there are separations of concerns where you can't expect anyone to do everything. And the nice separation of concern is: do I own and am I responsible for the stability of the things that you put code on? Or am I responsible for the code that I put on the thing, right? That is a nice seam because you want the infrastructure to be stable, to protect itself, to be resilient, all these things. And you want your code to be oriented towards: does every single user get a good experience? You could have one of those things be true and the other not be true. They are decouplable.
And actually, this is like even the most modern companies. I often refer to Anthropic as this company which operates in a very different way to most companies. They're very successful despite doing a lot of different things. However, internally, they have platform teams. They have the cloud platform teams. And then they have applied AI, which is more of the kind of the feature teams, the integration. And the two, I talked to both of them. They just have a very different outlook. They have a very different view on even basic stuff like will software engineers be obsolete? The people on the platform team were like, "No, we're working really hard." And on the applied side, they're like, "Well, maybe it will happen."
Does not surprise me one tiny iota.
No, but it's also... so this company Anthropic did start from a blank page. They arrived at the same place.
Yeah. No, I think it's the right separation of concern. And I'm not trying to erase it, but I think that to be a good engineer, you need fast feedback loops. And this is part and parcel with the whole, oh, the source of truth is the code. If that's where you live, if you live in the land of how it should theoretically work, no. And I think that with agents, they're breaking that. Right? They're breaking that and they're forcing...
Another thing on the observability trap is a lot of people, if you say like what is observability, they'll be like, ah, well there's three pillars. There's metrics, logs, and traces. We talked about this last time. Metrics and logs I would say are system exhaust. They're the exhaust pipe. And they're never going away because every team runs a ton of third-party software. They didn't write it. They don't own it. They just have to run it and it's outputting [ __ ]
Yeah. And you want
And you just got to put it somewhere.
it. You see what... And then you do stuff with it.
Yeah, and you know, you should put it somewhere cheap. There's a ton of it. It's not super high value, but you definitely need it, right? And you can't do anything about it. You just take it and put it somewhere.
Then there's your code. There's your crown jewels, the code that makes you a company. And for that code, your telemetry should be a product decision. It should be you store it once with all the connective tissue because the value of rich data goes up not linearly, not even exponentially, combinatorially. If you have a wide event or a trace with 29 bits of data and you add a 30th, that 30th is more valuable than all the others... Like it is just so powerful. And with non-deterministic software, you know right up front you can't predict what it's going to do. You have to care. Like that is a product decision to capture that trace.
So let's talk specifically about modern observability, like companies that are, you know, either building AI-heavy code or just complicated code that they're generating. You know, in the old world, again, I'm just being, you know, observability 101, back in the day, the way I would have written the code is you write the code and you think like, hmm, something funny might be going on here. Let me do a log or an info or a warn. And then maybe if we're putting it in production, I realize like, "Okay, well, I guess it's crashing and we don't have any logs there. So, I guess it's some other part. Let me put in a tool that will log everything and I'll have a bunch of stuff." Now, this is the very simplest way of thinking. In kind of a modern business where I'm like, I know this is high-value stuff, what are ways that I can go about that's actually maybe a bit more practical? Because I just want to use a super basic one.
Auto-instrumentation has gotten so good in recent years. If you're using OpenTelemetry, and everyone should be using OpenTelemetry, all of the common patterns, all of the models are trained on them. So, it is literally faster and easier to build with instrumentation than not to.
And with instrumentation, do you just... once I have the code and a compile step or an extra step, it just adds it to the right lines?
This is what's important, right? It's part of just developer intent. Right? This is how you declare your intent and that's how you check up on your intent in production. It's honestly gotten so much simpler. And you know, I don't fault developers or anyone else for not kind of closing that loop with DevOps because the fact is it was prohibitively hard and time-consuming and difficult because, you know, you're an old-school software engineer and you sit down, write some code, you're like, "Ah, here I should instrument it and look at it in production." So, you're like, "Okay, I've got a bit of data and I want to do something with it."
All right, is it a metric? A log? A trace? An exception? An error? A profiling? You know, it's just like, okay, if it's a metric, is it a counter? Is it a gauge? You know, it's just like all the way down. And then, well, what type of data is it? Is it going to have high cardinality? Is it going to be a...
Cardinality, you have to worry about that.
Blow it up. Yeah, you just... like if it's a log line, which log level do I do, do I append it to a... Like, you could double, triple, quadruple the amount of time that you spent writing the code trying to instrument it and then you still wouldn't be done. Like you deploy it and then it's like, "Okay, I know the name of the thing that I added, but how do I find it? How do I display it? How do I create a dashboard?" It's just like, that was prohibitively... That was really hard.
But now we can bring all of this to you right in your development environment. It is easier and faster to instrument with telemetry than without it. And you don't have to leave your development environment to go and get it, you know? You could have the agent... Like we've built some really cool [ __ ] at Honeycomb where it'll just be like, "Oh hey, that thing that you wrote, you know? Maybe you want to look at this." And you can control how verbose it is, you can, you know, but it's right there. And that's how it should be. It should be part of your development loop.
Can we talk about what spans are? Because I'll quote Eric Reda, who was on the Roll LinkedIn: "The basic idea of observability for applications is don't use logs or metrics, just put it all in spans." What are spans?
Spans are bits of a trace. I mean a trace is just a structured log with some fancy fields, right? And so the span is a subset of the trace that makes up the entire duration. And I don't know if you've followed any of this, but the default building block has been the transaction for as long as the web has been around. Yeah. That doesn't work anymore.
Specifically with AI.
Yeah. We just shipped something called Timeline that sits on top of spans. So, you know, if you run something like Intercom, you've got a chat thing and a customer's like, "I'm coming
conversation going on.
Yeah, customer's like, "I'm complaining." You're like, "Okay." So you spin up a supervisor agent that spins up more agents and each of them calls APIs, each of them calls storage backends and stuff, and then they return, and then the customer browser... that could span hours, right? And you need to be able to zoom out and visualize the whole thing. It's super cool.
And so this is a new primitive that you came up with for these use cases where there's a conversation or an LLM is involved, and you have like a
meta trace.
Oh, okay. Yeah, so I guess this
A trace of traces.
So we need these new building blocks to actually just, you know, be able to work with.
Yeah.
Interesting. So I guess this is something to keep in mind for any engineer who's building on top of LLMs, who is an AI engineer now, as we know.
It's either that or you've just got all these tabs open with traces and you're just copy-pasting IDs from one to the next.
Yeah. Or if you're a large enough company you might have built your own in-house tool, but we know that's... it's doable but it's painful.
It's doable, it's painful. I'm really looking forward to seeing over the next few months or year or whatever just the marriage of tests and telemetry and evals from a telemetry perspective.
With AI agents being around, a lot of them are now very useful to connect to observability stores. You can go on and do stuff. However, one question that comes up is, well, agents have a finite context window and with observability you can really easily overload that. What are approaches you've seen of agents using either Honeycomb or some other data sources to make them productive? Have you seen some patterns?
There's a lot of trash data out there. And a lot of traditional telemetry data, metrics, logs, traces, what have you, tends to fill up your context window with crap, when the most important part of the data is, again, the relationships between the... So, in fact, one of the AI SRE startups posted this great piece a couple months ago about how they see the agents that they deploy in the wild bypass the observability data most of the time, and they go upstream to find richer, intact telemetry data. So, that's what I would say. Either you give your agents the... But it's the relationships that matter, right? Because that's what actually helps the AI make decisions.
And when it comes to observability, I cannot not mention your book, Observability Engineering, and you have a second edition. Can you tell me why you felt the need to write it and what's new in it?
Oh, man. The whole thing is new. So, O'Reilly, anytime a book is considered successful and if the topic is still relevant, they'll ask if you want to write a second edition. So, it's not really... But I was really excited to write it. The first book, I don't want to say I wasn't proud of it. You're not... like your children and your books are not supposed
to like to say anything bad about them, you know, cuz it's fine. They're, you know, but it was written 2019 to 2021. The definition of observability meant one thing when we started and another by the time we ended, and there was at no point where I was like, oh, this book is great. Let's ship it. It was just like, oh god, I can't do this anymore. Just like take it. Please take it and I hope that's enough. Now, it feels like the definition of observability is more subtle. It's everything else in the world that's like changing and crazy and all. So, I think it's a good book. I hope it can help a bunch of folks.
It's got six parts. So, the first part is, and I wrote parts one and six, first part is just kind of like grappling with what does it mean to run deterministic and non-deterministic systems, you know? And then, you know, my co-authors Liz and Austin and George, the part two and three is how do you instrument your code and how do you understand it? And there are parallel tracks for doing this with or without AI. And a couple of great guest columns from Jeremy. And then parts four and five are, we have a whole lineup of guest authors and use cases and deep dives. And
Interesting. So that came out mobile.
We've got some great ones on CI/CD. ClickHouse did one on columnar storage. Some really, really stellar things. There's a chapter from Kasha Finn on how they use it iteratively to like do observability.
Oh, so this is a brand new book. It's not a, a lot of second editions are like, oh, we added like, you know, two chapters.
It is an entire rewrite and it's twice as long. The first one was 250 pages. This one is 600 pages.
Okay. So I'm interested. I'm going to get this book.
And the part six is my baby, and it was originally supposed to be three chapters for observability engineering teams, and it turned into, it's a third of the book. It's 200 pages. But it's topics for observability governance for leaders. And it starts with an open letter to CTOs telling them why all their big AI goals are blocked behind their ability to make sense of their system. You know, and then we talk about, you know, software delivery for, no buzzwords, just systems theory, right? If you like Donella Meadows stuff, then you will like it. And then stuff, and then there's a chapter on how to quantify the impact of observability for your finance. How to treat observability as an investment versus a cost center, and when you should use observability as a cost center and when you should treat it like an investment, cuz it inherits the type of software that you're observing, you know?
And there's a great guest chapter from Rick Clark on staff plus, principal, distinguished engineers who are trying to drive massive change without authority. How do you do that? And how is observability vital to that? And then there's a chapter on build versus buy versus open source and
It sounds to me that anyone who is inside or wants to be inside a platform engineering team, whether you're an engineer or a leader,
in charge of
you probably want to read this book.
And at the end there's a chapter that is possibly one of my favorites, which is, it's called the art and science of vendor partnerships. And it's just talking about how we can't build all the software that we need. And great vendor partnerships are ones where you have influence over their road map and they trust you to do these things, and like talking about how most transformations fail. The ones that succeed succeed because someone on the inside has trust and credibility. People believe when you say something it is true. You know, it cuts through bureaucracy like a hot knife through butter. When it comes to partnering with, you know, the sales org of another company, you do not have trust and credibility. You work to build trust through reciprocity. You learn just how much you can trust them over time, right?
But the best vendor relationships are the ones where you genuinely feel like their successes are your successes, your successes are their successes. You're happy to see each other because each of you are delighted cuz you know you're getting something from them. It feels like you are two different teams working at the same big company. That is rare. Doesn't usually happen. And that's fine. Most vendor relationships are ones where you shake hands, you exchange money and services, and that's fine. But I think in an era of AI, these are durable skills. These are durable skills for very senior engineers who care about impact.
Senior engineers and also engineering leaders and anyone who wants to become an engineering leader, cuz I guess, like, I mean, both of us have been in engineering leadership, like you've been in much higher positions than I have, but I think it's fair to say that the way for you to get to that CTO role, that head of engineering, that director of engineering, is to do the work for six months to a year to a year and a half, and to do so you need to know these things. I feel Observability Engineering might be underselling this book, I'll be honest, the title. But I'm also going to get it and I'll probably think of ways to share a bit more, but thank you for writing and thanks to all your co-authors.
But speaking of leadership, I'd love to talk about a little bit of engineering leadership, cuz there's a lot of things that are changing. But I loved one of your very recent takes on leadership and I'm going to quote you: "The most effective leaders are kind, caring humans and skilled business operators. The second most effective leaders are terrible humans and skilled business operators. And after that comes everyone else. There are plenty of good kind humans who are sloppy operators and bad at business, because being good at business is very hard." And you said this in relation to what happened at Twitter/X, referring to Elon as a terrible human but a skilled business operator.
Yeah, I don't know that I would call him a skilled business operator, but my point was that Twitter had 16 years to figure it out. And everyone could see that they were not figuring it out. And whatever else he has
Figuring out the business specifically.
Yeah, building products, you know, reaching folks. And you could argue that X has gotten better or worse, but you can't argue that he is running it with 20% as many people.
Yep. And it's working.
And it's working. And some of that, you know, 30 engineers on the core product. And another 30, and the like 60 engineers; there were 1,700 before. You know, and you could argue, and I think it would be true, that it's some of the work that those engineers did that... But like, this is the point. If we don't do it ourselves, meaning hold ourselves to a high standard, build with efficiency, constantly be like trying to get better, if we don't do it ourselves, someone will come and do it to us.
This is what you also said; you closed saying, if we want to remain in leadership, if we want to set the culture and the tone and take the ethical stance that we believe in, we first have to win at the business. And I think this is, like, especially now that there's so many changes happening in technology, there's the whirlwinds, business will go up and down, I guess the reminder that you want to keep your eyes on the prize, especially if you're a leader,
The 2010s, there was so much money sloshing around in Silicon Valley, and times started to get tough, and all of these companies canceled their DEI programs and blah blah blah. Yeah, they never believed in that. They were just trying to buy people off. You know, and that is very telling to me. And I have taken a lot of lessons away from that, which is just that it's not enough to be a good person. I believe that people who are kind and care about people can, and usually do, do better than sociopaths in the same roles, but only if they're good at business.
Learn the business, stay close to it.
You got to.
With AI, now that coding has become cheap, now that engineers are running agents, how do you see the role of good, skilled engineering managers and engineering directors change? What has changed?
Well, the first thing that's changed is, I think, everyone has to, gets to be hands-on.
Specifically to generate some code, to ship to production to some extent.
To get a PR through, you know, you should know what it feels like. It's just easier now than it's ever been to pick it back up, to fill in the blanks, you know? And it's always been the case that leaders were better if they had a hand in it. And now it's just, there's no excuse not to. Teams are getting smaller in general.
I think this should be a good thing if we can figure out how to own more surface area. That should be a good thing. I worry that what is happening is it's being done by CEOs who are like, "Oh, well, this other company is doing it, or it's magic, or we're going to do layoffs." Or you like it, and then I really dislike the anti-management tone. So like, no argument that power tends to drift towards managers over time and needs to get pushed back into yours. No argument there's a tendency to have too many managers, you know, the bureaucracy kind of like generates a sort of, you know, it's easier to say yes than it is to say no. And so these things happen, so they need to be pushed back. Number two time.
But I believe that management and middle management is deeply essential. And I look forward to seeing how that works out for them not having any. But like, the role of middle management in my view is sense making and context giving, cuz like, I don't believe in a world where engineers are just given tasks. Here's your Jira, go do the things. Hey, AI can do that. I want people who understand what we're trying to do. Understand how we're trying to do it, or who are there to help us figure out how we're going to do it. And you can't engage emotionally, creatively, collaboratively without understanding. And that understanding is incredibly difficult to build, and it's fragile and never lasts very long.
For those of us listening who are middle managers, it's been a tough few years, because what they're seeing is there's a push to have fewer of them. A lot of their colleagues, if they're in unlucky places, they were made redundant, and many of them have struggled to get similar positions. We're talking director positions. We're talking head of engineering, senior engineering manager. That role is disappearing faster than ever. I think directors might still be there. For folks who are in this position and they do like middle management, they do believe they're good at it, what do you think tactics could be to give them a bit more career options?
Tactically, I would say go back to be an IC for a while, even if you know it's not what you want to do.
Mhm.
If you're at all capable. If you're not capable of it, then I would try to work it. You've got to get AI in your resume. You just have to. And this is a huge career risk. If you're working somewhere where you're not getting these skills, that is a massive risk. I would do whatever I could.
And this is very interesting that you're saying get AI in your career. Because I remember about a year, year and a half ago, I started to pay attention to like, okay, this is happening. And I remember a year ago I wrote an article about how to become an AI engineer. And I talked with engineers who just, at their work, they started to do AI, and now they're AI engineers. Next thing I'm hearing right now is the people who have like two to three years of AI engineering experience are so in demand. I'm doing research on the job market, and they're like, this is the best job market ever. However, you know, the people who are like, okay, I have none, but I want to get it, and let's say they're out of a job, they're struggling because no one's giving them the benefit of the doubt.
It is really hard, and I'm not saying it's right, but it's how it is.
And I guess the reason we're ringing this alarm bell is we know this change has not been as fast, so do it now because later
Now. The next time you go out for a job interview, anyone, you're going to be asked, and you're going to be filtered out if you don't have it. And the delta between those who are just getting started and those who have been doing it, it was here for a little while. It was very easy to get started. Now it's here. But it's opening up, and the longer it goes, the harder it will be to catch up. You just got to get some.
Let's talk about directors.
Yeah, directors are usually the ones who have been in management for like 10 years, usually. And there's a real feeling of fear often of like, God, tech has changed
a lot in 10 years. And
And this is where I would say your body, like the way we experience anxiety and the way we experience excitement is physiologically almost the same. Like, I used to play piano, right? And before a performance, I'd be like, "I'm excited. I'm so excited to do this." You know, cuz I'm like trembling and sweat. But like, the difference is agency. If you sit back and wait for the water to come to you, you're just going to be freaking out. But if you run towards the waves, just like run towards it. Try it.
You know, if you have a job now and you're a director and you're afraid of it, it's always seen as kind of noble when managers want to go back to being ICs, I think. It's very well respected. Own it. Run towards the waves. Own it. Be part of the wave, the frontier of people who are like, "I'm so ex-" Just tell yourself. Doesn't have to be true. "I'm so excited to be an IC again. It's never been easier to go back and try. I'm going to do it and I'm going to talk about my experience and tell everyone else about it." Just, you got to own it. Don't wait.
And then let's talk about junior engineers. Obviously, it's a harder time to get started as a junior, but how do you think about the value that they bring?
The hardest thing about quantifying the value of junior engineers is that we don't know how to quantify the value of any engineer. So it's all vibes. You know, it's so interesting, because I feel like we're over here doing all this hand-wringing about will juniors be okay? Will they ever learn the basics? But like, my friend Boris, who has a new observability startup, and he talks to these high school, college kids all the time, he's like, they are cooking. They don't know what the software development life cycle is, but they are just like, they're doing so much cool [ __ ]. I believe that the kids are going to be okay. We just have to hire them. We just have to give them a shot. They're going to come up with a lot of the conclusions and the ways and the hows that are going to be things that we wouldn't have thought of. But we just have to hire them. We just have to be willing to give them a shot.
This week and a half I've talked with a bunch of founders of young startups, and they've been telling me stories of this open source contributor who was outstanding, so they want to hire him or her. Turns out it was a 17-year-old kid. They still hired them, and now they have to tell me, like, oh my gosh, the things they do. So I think when you're saying the kids are going to be fine, just give them a chance and give them a shot, even if it's an internship.
Yes, totally.
I feel more companies should, cuz an internship is lower risk, lower duration.
Yeah.
And even if that person doesn't work out, with an internship under their belt
Yeah.
so much better for everyone.
Totally. Totally.
One question that came up when I asked, you're going to be on the show, what I should ask, they said AI fatigue. Someone asked, like, can you please ask Charity, as an engineer, if I'm starting to get just really, really drained of this? Have you had this? Do you see people having it, and what is a good way to, you know, just
deal with it? We know it's here, we know it's here to stay, but still.
I mean, my follow-up question would be like, which variety of AI fatigue?
Okay, tell us about the varieties.
You know, 'cause for some people, when they say AI fatigue, they're talking about receiving slop. Some people are talking about all the hype and the — have you heard the phrase or the term doom trolling?
No.
Cal Newport is, I think, his name. He's a computer — he's an AI researcher professor on the East Coast, and it's his term for what the CEO of Anthropic and OpenAI keep doing about, oh my god, this might be the end of blah blah blah. And he's like, it's just doom trolling, and they need to stop it because they're stressing everyone the [ __ ] out.
Yeah.
And stop, because it's just not responsible. You know, so yeah, I think there's a lot of fatigue around that. I think that a lot of people, their family members are afraid. You know, it's just — always before in the history of technology it's been something cool or fun, or this will be the iPhone, it'll make your life better, and now it's just like fear. It's pretty crappy. So there's that.
There's the fatigue of like, I found myself being off social media because I'm just so tired of all of the AI slop posts. It's just like, I'm not interested. There are a lot of different varieties here, and yes, we are all feeling it. So I guess I would repeat my call for us to remember that we are in control. We are in charge. I think the universal nature of the frustration means that this is a great time to propose experiments where we take back control. Maybe you and your team agree we don't actually want any more AI-generated PR descriptions. None of us use AI on Wednesdays. Maybe we take a week, you know, just take control back. Try something. Propose something.
I guess because change is so big, experimenting has never been easier. And I guess most businesses, most directors, most leaders would welcome teams saying, you know, we're going to try out — 'cause your answer would probably be, I mean, you're in this position, your answer, I guess, will be sure.
Better yet, don't even tell me. Come and tell me what worked afterwards.
Yeah, and what didn't and what you learned.
And then other teams can learn from that, right? I think sometimes people are waiting for top-down permission, but we don't know what permission to give until it works. It's so much better when it's bottom-up, when people just try. Take control of your time and your calendar.
And I guess maybe we just forgot that there have been major changes in the industry. I remember the iPhone change, and I remember when the iPhone came out, iPhone and Android, the smartphones. The people who were the most kick-ass iOS engineers, you know who they were? They were typically 18- or 19-year-old kids who went into this and they tried it out. Guess what? Two years later, they were the domain experts. The staff engineer was a 22-year-old, and then the entry-level engineer was a 40-year-old. And again, not always, but my point is, when there's such big change, you can actually become an expert by
Very little time.
By you taking
Just taking charge?
Taking charge. And also no one's really going to tell you no, because no one knows.
Exactly. Exactly. There's some liberty there.
So as closing, just to go back to a little bit of being human and slowing down, what are one or two books that gave you something?
Ooh. I really got a lot out of Catastrophe Ethics. I haven't seen it mentioned many places, and I think real philosophy nerds would be like, "That's kind of a pop book, you know?" And I think the people who are not real philosophy books are like, "That's kind of a lot of philosophy." But you know, he's a bioethicist, I think. Travis Rieder, R-I-E-D-E-R. Catastrophe Ethics. And he talks about how the puzzle of modern life is that it feels like in everything we're implicated. Every choice we make — are you going to use milk? Well, you know, the cows were tortured. Are you going to use almond milk? Well, water is a problem. Well, soy milk, like, all hormones. And it's just like, whatever you do, you are hurting someone, and it feels like the problems are so large that none of our decisions really matter.
And that tension, like what — and then he kind of walks through traditional ethical frameworks like utilitarianism and stuff, and just shows how there is no recipe anyone can follow that doesn't lead you to some really stupid — and he's like, "This is just no gods, no masters." Which doesn't mean that everything's relative. What it means is that the way to live an ethical life of integrity is you need to educate yourself about the world. You know, you need to know things, right? And then listen inside, you know, and where are you drawn? What suffering really speaks to you, or what cause do you — you know, 'cause no one can tell you what matters. You have to decide what matters. And so that introspection, it's so at odds with this sort of performative rage, you know, which I'm just so exhausted by.
All right, so that's one. Number two, this is a book that I've recommended a couple times, but I'm just going to keep recommending it 'cause it's so good. It's by Adam Becker, and it's called More Everything Forever. He is a journalist based in San Francisco. He has a philosophy undergrad and a PhD in astrophysics, and he just demolishes all of the AI religion, the singularity and the effective altruism and accelerationism, and the whole, what if we could have infinite growth foreverism. And he's like, the heat death of the universe. You guys, literally the only thing we know about exponential growth is that it must end. It must end in an S-curve or in a crash. It must end. And he's got this dry sense of humor.
And there are a couple times where he's just describing some of the very real things. He's just like, "Why do Oxford ethicists want this?" He's talking about taking over star systems and stuff. And it's just ridiculous. And he also gets into — he talks about all these people who are working so hard on life extension. And he's like, "These are a bunch of sad little boys who miss their daddy." And I was just like, "Oh my god." The oldest fear of humanity is a fear of death. And you just see it. You can't unsee it. So, yeah, those are my two. They're both so good.
Charity, thank you so much. This will finally be able to happen.
Finally. It's a good time.
I always really, really enjoy talking with Charity. I hope you also liked it. I appreciated how Charity talks about the trust account. If we are debiting trust from the creation of code because AI wrote it and no human read it, then that trust needs to be refilled somewhere else. Testing 5,000 guardrails are all ways to add more trust that we lost by using AI. I also appreciated how she talked with empathy about both AI camps.
The enthusiasts or AI-pilled folks are seeing the practical wins, while those operating production systems see the slop. Neither side is wrong, but they should talk to each other more. So, if you see wins with AI, share with the broader team. But also talk about it when it creates more work, reduces reliability, or when it degrades quality. And for those of us feeling anxious about all of this change, especially directors and managers, I'll leave you with Charity's advice. Anxiety and excitement are psychologically almost the same, but the difference between them is agency.
So, instead of waiting for change to come to you, take charge however you can and make changes yourself. Do check out the show notes below for related The Pragmatic Engineer deepdives on how AI is changing software engineering, and for another discussion with Charity on observability. And I can very much recommend her book Observability Engineering, second edition. If you enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. Special thank you if you also leave a rating on the show. Thanks, and see you in the next one.
Article published
