Charity Majors on Shipping Code You Haven't Read and Why Both AI Camps Are Right

Open on YouTube ↗
Overview

Charity Majors, co-founder of Honeycomb and co-author of Observability Engineering, describes herself in this conversation as someone who was skeptical of AI for software development in 2025 and changed her mind based on what she saw. The conversation with the host of The Pragmatic Engineer is built around one question she keeps coming back to. The question isn't whether engineers will ship AI-written code they have never read, but what it would take for them to be comfortable doing so. From there she covers code review, non-determinism, the growing split between AI enthusiasts and the people on call, the unfinished business of DevOps, observability for agentic systems, and what all of this means for managers, directors, and junior engineers.

32 min read

From Parse to Honeycomb: the tool that made a hard problem easy

Majors' first job was at Linden Lab, building Second Life. She later worked at Parse, the mobile backend service, which Facebook acquired and eventually shut down. She calls it her "first great lesson" that most acquisitions fail. She says she is still grateful for the experience. As a lifelong "startup kid," nobody knew her name until she was leaving Facebook, and that was when investors started offering money.

Honeycomb's origin comes from a specific experience at Facebook. Parse ran on AWS and Ruby on Rails, but after the acquisition the team got access to Facebook's internal tools, including one called Scuba. Parse was growing on a hockey stick curve, with over a million mobile apps hosted by the time she left. Almost every week some app would suddenly hit the top 10 on iTunes and break something. Finding the cause was a needle-in-a-haystack problem. The app spamming the logs might not be the real culprit, since it might be stuck behind the actual cause. Diagnosis used to take hours or weeks and often depended on luck. Once Parse's data went into Scuba, she says, it became "click, click, click, oh, there it is."

When she was about to leave, she planned to become an engineer or engineering manager at a company like Slack or Stripe, and realized she would be far less effective without a tool like that. Her original plan for Honeycomb was modest. All startups fail, so she would spend a year or two writing Go, then open source the result and take it with her wherever she went.

Individual productivity and the obsession with speed

In 2020, before AI tools, the host and Majors both answered a reader's question about whether individual developer productivity can be measured. Asked the same question now, she says measurements are useful as supporting data, "like color in a painting on the wall," but Goodhart's law still applies. Managers need data so their judgment isn't just an offhand opinion shaped by bias, but the numbers are not the whole picture. She also pushes back on the recent emphasis on individual output. Teams, she says, are still what matter.

What encourages her about the AI moment is that it forces everyone to keep asking what good looks like, what productivity means, and what "better" or "great" would look like. She thinks it telling that the industry jumped straight to speed, doing the same thing faster, and she now considers that "a very immature description of what better is."

The host gives an example. Anthropic had just published a podcast with Spotify's engineering leadership describing thousands of changes shipped per day or week with Claude Code. Meanwhile, the host had trouble publishing episodes because Spotify was down. The discussion was all about speed and not about quality or functionality people want. Majors asks whether customers really want the buttons in their apps moving around all the time.

How her view of AI shifted

In a year-end post, Majors wrote that 2025 was for AI what 2010 was for the cloud. In March 2025 she and Fred Hebert gave the closing keynote at SREcon, just after the term "vibe coding" was coined. When they told the audience to try vibe coding, people groaned and laughed. Their pitch was that engineers should learn AI "because you can complain better if you learn it," which she says she meant sincerely. At that point she saw AI as a big feature, bigger than a programming language and comparable to the cloud, but not generational.

Her own turning point was November 2025, when Opus 4.5 was released. Looking back, she has written that the shift was visible earlier, and that it came from the harnesses and tooling more than the models. Around July, people were already saying it was coming faster than expected. The tooling went from something like a shell script that retries to a lot of surrounding infrastructure. In the early months of 2026, she says, it felt like everyone around her was trying AI again and changing their minds.

She doesn't think the earlier skepticism was wrong. It's an extraordinary claim that AI can write code about as well as a median software engineer, within limits, and the industry has heard that promise before. She points to a sticker joking that "we already have a programming language that lets you" do this, with COBOL as the punchline. The host adds no-code and low-code to the list. Skeptics were right those times. Now, they agree, skeptics are wrong.

What would it take to ship code you haven't read?

Majors sees the same pattern playing out with the question of shipping unread code. She thinks arguing about whether or when it will happen is pointless. The useful question is what it would take, and answering that is engineering.

The host's first reaction was that they would never do it. On reflection, they note that trust has always come from indirect signals. Pre-AI, a trusted teammate saying "I vouch for this, I've hammered it" was enough. If a change could be shown to have been tested in some harness, that could work too. Majors suggests a concrete path. Run a human and an AI reviewer in tandem for a few months and compare how much each catches, whether it's about the same, more, or less, while training the AI. Whether that takes five days or five years, she says, the direction is clear.

The host's closing summary picks up her image of a "trust account." If trust is being debited during code creation because no human reads the code, it has to be built up somewhere else: tests, evals, and guardrails.

Rewriting instead of editing: the Phoenix architecture idea

Majors thinks this direction is good for the profession, and she leans on Chad Fowler's writing on the "Phoenix architecture." The host reads Fowler's argument. Immutable infrastructure, stateless services, containers, blue-green deployments, and infrastructure as code all share the premise "never fix a running thing, replace it." AI extends that premise from infrastructure into application code. When rewriting is cheap, editing in place becomes risky, because mutation accumulates entropy and replacement resets it. Majors sums it up: "code is cache."

The analogy is the move from servers as pets, configured and repaired by hand, to disposable machines. For 60-plus years, software has been edited in place. Majors says the economics are what change this. You could generate 10,000 variants of a function faster than you could write it once. That means you'll need a lot of evals and tests, but cheap generation pushes in that direction. The cost of software has always been tied up in maintenance, and we trust code largely because it has been running in production for a long time.

She is careful about how far this goes. Anyone who has done a hard database migration, she says, should be humble about our ability to extract implicit contracts and store them. She believes the industry can go some distance and further than it has now, but she doesn't know how far. Her example is Parse. The team spent about six months writing the original Ruby on Rails API and then two years rewriting it in Go. Go was immature then, so they had to write their own database drivers and other pieces, and the original code had no type safety. In a strangler-fig migration, she says, "you literally find the contracts with your users by breaking them, one after the other." That doesn't seem like the ideal artifact to her. The contracts should be stored somewhere, and there should be reviewable architecture diagrams that generate code to spec.

The host points out that this echoes 1990s ideas like UML-driven code generation at Rational Software, which didn't work out, possibly because generating and reviewing code was still expensive. Both wonder whether those ideas may now be feasible. Majors says that is her hope.

Learning from sysadmins, ops, and QA

Majors started as a sysadmin at 17 at a university and remembers how stressful earlier automation shifts were. Sysadmins agonized that things would never be the same, and everyone adapted. The people who ran updates by hand on every server built the systems that replaced that work and spent their time writing code instead. The host notes that the sysadmin role has largely disappeared, but its practitioners understood operating systems and hardware and moved into software engineering, product management, even tech sales. Majors adds that her generation are still the best debuggers. She's glad not everyone has to learn about CPU and memory, but that knowledge still comes in handy, and she sees parallels for generated code.

She has written that lines of code are not the ideal artifact to review. The tools to do better don't exist yet, but many of the ideas do, and most come from operations and QA, "two domains that software engineering has historically been rather snobbish about." Ops and QA, she says, have always been concerned with what is, while software engineering has been concerned with what should be. She finds it strange how much software engineers believe the world exists in the repo. It doesn't; it's production. Some companies don't even let engineers look at production.

What excites her about AI is that it pushes the discipline where it has needed to go for a long time: "Production is not what happens after development. It is a stage of development." The host recalls the old "don't deploy on Fridays" trend on tech Twitter, when Majors argued that teams should be able to deploy anytime without fear. She restates her position: as soon as you merge, the code should be going out, and you should have to stop the train to keep it from reaching production.

Unbundling code review

The host says nearly everyone they talk to agrees that code generation is cheap and human code review is now the bottleneck, so the industry is building better review tools. Uber, for example, has built tooling to surface the important reviews. The host admits they never liked doing code review unless it was small and with someone they cared about, and they care even less when the author is an AI.

Majors says part of the problem is that code review means many different things in different places, which invites projection. If you say you don't want code review, people hear that you don't want to talk to coworkers or mentor juniors. Code review is overloaded. Some of its functions are good, some could be done better another way, and some are highly cultural.

She cites her former Parse colleague David Poll, now working on pull requests and code review at GitHub. For him, code review is where a team decides whether it wants a change in the product at all. Majors calls that a great discussion and exactly what humans are good at: whether the mental model is coherent, whether to add something, what the architecture and API design should be. Ideally that happens before the code is written. Reading for syntax and bugs isn't evil and can be a teaching opportunity, but it doesn't feel like a great use of anyone's time. The host suggests it is most useful when onboarding someone onto a team without written guidance or lint rules. Majors replies that people fill cracks with their own time when guardrails haven't been built, and that the industry has never thought to extract many of those things from the process.

Her model example is Intercom. A decade ago their CTO said "shipping is your company's heartbeat," and they ship a Ruby monolith hundreds of times a day, which she says is not trivial. She considers them a high-water mark for teams founded before AI, with strong engineering discipline, that have become AI-native. They wrote about AI-validated PRs with a very high bar, effectively putting their most senior engineers' wisdom on every diff. That frees humans from nitpicking, which they aren't good at anyway, so they can discuss whether the change is going in the right direction.

Non-determinism demands more discipline, not less

Majors has written that non-deterministic systems require more engineering discipline. Asked what that means, she starts with tests and evals. When code is treated as a trusted artifact, tests cover only the failures humans can predict, plus regressions added after incidents, which she says is not a high bar. She highlights conformance testing: if I'm not going to read this code, how do I know it performs within the boundaries of the last version I generated? For many workloads, she says, not changing too much matters as much as absolute performance.

She prefers to think less about "AI" and more about deterministic and non-deterministic systems that have to work together. Determinism isn't going anywhere and is enormously valuable. The job is to make AI "kind of boring": give a valuable but erratic tool carved pathways where it acts as a superpower instead of eroding foundations.

The host recalls Martin Fowler saying on the podcast a year earlier that non-determinism is AI's biggest change, and notes that businesses want software that behaves the same way every time. Their example is an open-source resume-scoring tool used in HackerRank's applicant tracking. Run locally with a small model (Gemma is recommended), the same resume scored anywhere from 66 to 99 across about 100 runs, while many companies set the bar at 85. That turns a recruitment aid into a coin flip. Majors says we have to be able to call that bad, because AI is not the right tool for every use case.

Honeycomb's AI norms: you own the loop

Honeycomb has been running a series of conversations on AI norms and values. Majors says she wouldn't have trusted the company's judgment a year ago because it didn't know enough. Now, if a coworker says AI is the wrong tool for a job, she trusts them. Her view is that "you got to get worse before you can get better."

The norms start from the premise that the bar has gone up for everyone, as it always does with powerful new tools. Asked whether it's the bar or the floor that has risen, she says she doesn't know, and that the industry is "wandering in the wilderness." But you can't stop wandering or you'll be left behind. The only viable way to define the bar is better outcomes. Another principle: there is no "human in the loop." You own the loop. "Claude said this" is not an excuse; the work is yours.

On communication, she offers a baseline. You can't send anyone something you haven't read, and if it would take them longer to read than it took you to make, it's probably slop. That is disrespectful, because it pushes the cost onto the recipient. She has also caught herself asking colleagues questions without first trying to find the answer. The host compares this to how new grads learn to respect senior colleagues' time, and argues that attention is now the scarce resource.

Majors distinguishes two uses of AI: a shortcut to avoid thinking, and a tool for thinking more deeply and rigorously. Both have their place, but for core job functions, and especially when asking someone else to review your work, Honeycomb wants the second. She says this isn't absolutist. People writing in a second language or who are neurodivergent may lean on AI, and that is still respectful. Existing standards of quality and respect are enough, she says; they just need to be applied. The "look what this can do" phase, she adds, is over.

The two camps: people on call and people not on call

The host describes two camps, the "AI-pilled" and those who seem to hate AI, with little dialogue between them. Majors says neither side is making it up. Enthusiasts feel the race: competitors are moving faster and leapfrogging, and the industry is "on the inside of an exponential curve," which she calls rare and usually short-lived. That is real, and it's wise to prepare. Ironically, she says, both sides feel like the embattled minority defending the truth.

Often the division comes down to who is on call. The people where the buck stops are seeing melting mental models, slop, and their hard work dissolving, with no end in sight. The host confirms that on-call engineers report more incidents and carelessness, with systems getting "way worse" since AI adoption.

The host shares something heard from inside Meta. There has been a flurry of SEV0s, the highest severity, over roughly two months, in areas including Instagram and WhatsApp where reliability and trust-and-safety staff had been cut. The host says each incident has its own postmortem and the connection isn't direct, but Meta hasn't had it this bad in close to a decade. When the host told this story at a conference, people from other VC-funded and public companies quietly said the same thing was happening to them, and that they'd thought it was just them.

Majors points again to Intercom, which she praises for publishing the unflattering data. By her account, their reliability and code quality declined for 18 months and had only recently started possibly recovering, still below where it was. Her plea is that teams talk about the wins and couple them with the costs. There are incredible gains in rewrites and in automating away toil, and she says nobody she has talked to would give them up. But people who see only the wins assume their on-call coworkers are just afraid of being automated away ("No, dude. You be on call and then see how you feel"). On-call engineers who never hear the costs admitted assume the wins are fake. "Tell the whole story," she says.

Why software is AI's killer app

Majors calls AI "just technology," and argues that software is its killer app. Software is made of logic and language, and so is AI, which means guardrails, checks, and validation can be built in. She doesn't see how to do that elsewhere, citing courts sanctioning lawyers who submitted briefs with hallucinated content. The host adds that software may be the upper bound: if something can't be automated in software, where validation exists and training data compiles, other industries are unlikely to manage it.

The host tests each new model by asking it to write in the style of The Pragmatic Engineer and can always tell it's AI, whereas generated code often looks like something they could have written. Majors agrees AI is much better at code, calling software "a simplified version of language for a purpose." She spent time trying to write more efficiently with AI and stopped, because "writing is thinking on paper" and there's no shortcut for that thinking. She uses AI for structure and feedback, not to generate her writing.

The host predicts that in five years, hiring will come down to conversation, and candidates who invested in their own thinking will stand out against those who outsourced it. Majors says she is excited about "leaning into the parts of being human together." She dislikes spending the day switching between agents and people on Slack because it feels too similar. Honeycomb is fully distributed, and though she loves not leaving the house, she craves in-person connection. Her hope is that people remember "we're in charge of the machines."

DevOps set out to build one feedback loop, and failed

DevOps, Majors explains, existed because devs and ops were split, with code thrown over a wall. She calls that split the "original sin": half the people write the code and the other half understand it, and she argues you can't really understand code you don't operate. The first wave, getting ops people to write code, succeeded. The second, getting software engineers to understand their code in production, was less successful. In her view, 20 years of DevOps was about creating one feedback loop connecting people who write code to that code in production, "and it failed."

She doesn't want to erase the separation between platform and application teams. It is a healthy seam: one side owns the stability and resilience of the infrastructure, the other owns whether every user gets a good experience, and one can be true without the other. The host notes that Anthropic has arrived at the same split, with cloud platform teams and applied AI teams, and that the two groups even disagree about whether software engineers will become obsolete. The platform people say no, while the applied people say maybe. Majors is not surprised. What matters, she says, is that engineers need fast feedback loops, and agents are breaking the idea that the code is the source of truth.

Modern observability: telemetry as a product decision

Majors revisits the "three pillars" of metrics, logs, and traces. She calls metrics and logs system exhaust. Every team runs third-party software it doesn't own, and that output should go somewhere cheap. Your own code is different. It is the crown jewels, and its telemetry should be a product decision, stored once with its connective tissue. She says the value of rich data grows "not linearly, not even exponentially, combinatorially": adding a 30th attribute to a wide event with 29 is worth more than all the others. With non-deterministic software, you know upfront that you can't predict its behavior, so capturing traces is essential.

Asked what this looks like in practice, she says auto-instrumentation has become very good, everyone should use OpenTelemetry, and because models are trained on its common patterns, it's now faster to build with instrumentation than without. She says instrumentation is how you declare intent and then check on it in production. She doesn't blame developers for not closing the DevOps loop, because it used to be prohibitively hard. For each piece of data you had to decide whether it was a metric, log, trace, exception, or profile; whether a metric was a counter or gauge; whether cardinality would blow it up; which log level to use. That could multiply coding time, and after deploying you still had to find the data and build dashboards. Now it can happen inside the development environment. Honeycomb has built features that suggest relevant data about code you've just written, with adjustable verbosity.

On spans, Majors describes a trace as structured logs with extra fields and a span as a piece of that trace's duration. The transaction has been the default building block since the start of the web, but she says that no longer works. Honeycomb has shipped a feature called Timeline on top of spans for agentic workflows. In a chat support scenario like Intercom's, a supervisor agent spawns more agents that call APIs and storage backends, and the whole interaction can span hours. Timeline acts as a "trace of traces" so you can zoom out and see it all. Without it, she says, you end up copying IDs between browser tabs. She looks forward to tests, telemetry, and evals converging.

On agents' limited context windows, she says a lot of traditional telemetry is "trash data" that fills context with noise, when what matters is the relationships in the data. She cites an AI SRE startup's writeup reporting that its agents often bypass typical observability data and go upstream for richer, intact telemetry. Her advice is to give agents the relationships, since those are what help AI make decisions.

Observability Engineering, second edition

Majors says the second edition is a complete rewrite, roughly 600 pages versus the first edition's 250. The first was written from 2019 to 2021 while the definition of observability was shifting, and she recalls finishing it with a feeling of "please take it, I hope that's enough." Now the definition feels more settled, even as everything else changes.

The book has six parts. Majors wrote parts one and six. Part one deals with running deterministic and non-deterministic systems. Parts two and three, by co-authors Liz, Austin, and George, cover instrumenting and understanding code, with parallel tracks with and without AI. Parts four and five contain guest chapters and deep dives, including one on CI/CD, one from ClickHouse on columnar storage, and one on iterative use of observability.

Part six was planned as three chapters for observability teams and grew to about 200 pages on governance and leadership. It opens with an open letter to CTOs arguing that their AI ambitions are blocked by their ability to make sense of their systems. It covers software delivery through systems theory ("if you like Donella Meadows"), quantifying observability's value to finance, and when to treat observability as a cost center versus an investment, which she says follows from the type of software being observed. It also includes a guest chapter from Rick Clark on how staff-plus engineers drive change without authority, a chapter on build versus buy versus open source, and one of her favorites, on vendor partnerships. There she argues that successful transformations depend on insiders with trust and credibility, and that with vendors, trust must be built through reciprocity. The best relationships feel like two teams at the same company. That is rare, and transactional relationships are fine, but she calls these "durable skills" for senior engineers in the AI era.

Leadership: be kind, and be good at business

The host quotes a recent Majors post: the most effective leaders are kind, caring humans and skilled business operators; next come terrible humans who are skilled operators; then everyone else. She wrote it about Twitter/X. She hedges on calling Elon Musk a skilled operator, but says Twitter had 16 years to figure out its business and visibly didn't, while X now runs with a fraction of the staff. She mentions roughly 30 engineers on the core product and about 60 in total versus 1,700 before. Her lesson: if we don't hold ourselves to high standards and efficiency, "someone will come and do it to us." She adds that when money tightened after the 2010s, many companies cancelled DEI programs, which she reads as showing they never believed in them. Her conclusion is that kind people usually outperform sociopaths in the same roles, but only if they are good at business.

On how engineering management changes, her first point is that everyone now gets to be hands-on. Leaders should know what it feels like to get a PR through, and it has never been easier to pick coding back up. Teams are getting smaller, which she thinks can be good if they can own more surface area. She worries that cuts are often driven by CEOs copying other companies or treating AI as magic. She dislikes the anti-management tone in the industry. She agrees that power drifts toward managers and has to be pushed back, and that bureaucracy breeds too many managers. But she considers middle management essential for sense-making and context-giving. She doesn't want engineers just handed Jira tickets, and people can't engage creatively without understanding, which is hard to build and fragile.

For middle managers struggling to find roles, her tactical advice is to go back to being an IC for a while, even if it isn't what they ultimately want, and to get AI experience on their resume. Working somewhere that doesn't build those skills is, she says, a massive career risk. The host adds that engineers with two or three years of AI engineering experience are in heavy demand, while people without any struggle to get the benefit of the doubt. Majors agrees the gap will widen the longer people wait.

For directors who fear how much tech has changed during a decade in management, she draws on her experience as a pianist. Anxiety and excitement feel almost the same physiologically, and the difference is agency. Waiting for the water to reach you means panicking, so run toward the waves. Going back to IC work is widely respected, she says; tell yourself you're excited even if it isn't true yet, and share what you learn.

Junior engineers and AI fatigue

On junior engineers, Majors says the hardest thing about quantifying their value is that nobody knows how to quantify any engineer's value. She is optimistic. Her friend Boris, who runs a new observability startup, talks to high school and college students and says they are "cooking," doing impressive work despite not knowing the software development life cycle. The kids will be okay, she says, and will reach conclusions others wouldn't have, but they need to be hired. The host shares a story from founders who hired an outstanding open-source contributor who turned out to be 17, and suggests internships as a lower-risk way to give juniors a start.

Asked about AI fatigue, Majors asks which kind. Some people mean receiving slop. Some mean the hype and what Cal Newport calls "doom trolling," AI company CEOs repeatedly suggesting catastrophe, which she considers irresponsible because it frightens people, including their families. Unlike earlier technology waves that promised to make life better, this one is framed around fear. She has also stepped back from social media because of AI slop posts.

Her response is to remember that we are in control. Because the frustration is universal, it's a good time for teams to propose experiments to take back control: no AI-generated PR descriptions, no AI on Wednesdays, a week off. As a leader, she says she'd prefer teams not ask permission at all and just report afterward what worked, what didn't, and what they learned, so others can learn too. Top-down permission doesn't work well when nobody knows what works yet. The host recalls the early smartphone era, when the best iOS engineers were often 18- or 19-year-olds, and a 22-year-old could be the staff engineer while a 40-year-old was entry level. In periods of big change, people can become experts quickly, and "no one knows" means nobody will say no.

Two book recommendations

Majors closes with two books. The first is Catastrophe Ethics by Travis Rieder, whom she describes as a bioethicist. It addresses the modern feeling that every choice implicates us in harm, whether cow's milk, almond milk, or soy, while problems feel too large for any decision to matter. As she describes it, Rieder shows that no traditional ethical framework offers a recipe that doesn't lead somewhere absurd. That doesn't mean everything is relative. It means living with integrity requires educating yourself about the world and then listening inward to decide what matters to you. She contrasts that introspection with the performative rage she is exhausted by.

The second is More Everything Forever by Adam Becker, a San Francisco journalist with a philosophy undergraduate degree and a PhD in astrophysics. She says it demolishes the "AI religion" of the singularity, effective altruism, accelerationism, and infinite growth. The one thing we know about exponential growth, she paraphrases, is that it must end, in an S-curve or a crash. She enjoyed Becker's dry humor about grand visions of colonizing star systems and about life-extension enthusiasts. What stuck with her is that underneath it all is humanity's oldest fear, the fear of death, and once you see it, you can't unsee it.