Nicole Forsgren on why AI speeds up coding but not shipping

Open on YouTube ↗
Overview

At The Pragmatic Summit, Nicole Forsgren, co-author of Accelerate and author of the new book Frictionless, discussed a tension that runs through the book: AI helps developers write software faster than ever, yet delivery often stays slow. Forsgren now works at Google on how people build software, including what they called "agent experience." Their answer centered on bottlenecks that sit outside the coding loop, the changed nature of cognitive load when working with agents, and the need for measurement that goes beyond counting output.

22 min read

What Forsgren is working on now

Forsgren described their current work as a continuation of what they have always done: figuring out how to improve the way people build software. That now includes making agents smarter and better so that humans can work with them more effectively. Borrowing a line from Laura Tacho earlier in the day, Forsgren joked that if we can't call it developer experience, we can call it agent experience "and then it's all going to work."

A large part of the job is still measuring hard things like productivity. Forsgren acknowledged, echoing an earlier session, that productivity measurement was already weak and has become "extra bad." They argued that the alternative is worse. Developers often dislike productivity metrics because they can feel like an attack, but without any signal, decisions come down to a director's or VP's gut feel. Some kind of signal, they said, is better than vibes.

Inner loop acceleration, outer loop overload

The host pointed to a claim at the start of Frictionless: AI is producing software faster than ever, yet shipping is still slow. Forsgren's explanation was that the industry focused generative AI on the coding inner loop because that is where the results are visible and where "all of the dopamine hit comes." The systems downstream of coding, such as security review, launch processes, deployment, and code review, were already known to be imperfect. They worked well enough, often because a handful of people managed them. AI, in Forsgren's words, "threw gas on the fire."

The result, as Forsgren described it, is that teams are now chasing constraints and bottlenecks much more visibly than before. In the short term more work did get out the door. Now both technical systems and human processes are being overwhelmed.

Where the bottlenecks show up: review and release

Asked for concrete examples, Forsgren said the answer depends on where an organization is in its process, but code review comes up often. Humans were already a bottleneck in review, and AI has made it worse in a specific way. Some companies had automated review for fairly straightforward changes. They have removed that automation when AI is involved because they worry about the verifiability or reliability of AI-generated code. The review burden has shifted back onto people.

The second area Forsgren sees in conversations with companies is deployment and release. For many engineers this is a black box, but it is often run by humans: selecting the right candidate build, verifying it, working out cherry-picks, rebundling, and sending it out. When one or two people, or a small group, make those decisions and do shared sense-making, the process does not scale to the volume of changes AI now produces.

Organizations still organize: process debt and onboarding

The host raised a scenario from the book, which is based on a true story. A new hire using AI tools produces a first contribution, and it then sits for two or three weeks because code review didn't flag it and the new hire lacks database access. The host asked how common this is, given that many people in the audience work at startups where going from code to production is quick.

Forsgren's reply was that "organizations are still going to organize." Review processes where one person must sign off, and all the structure built to make work more uniform, are often exactly what slows teams down. Coding is speeding up, and agents are starting to help with review and other tasks. According to Forsgren, though, many companies have only recently started to think about applying AI to the human, business-process side of delivery. Until they do, that side will keep slowing things down.

Onboarding illustrates the point. Waiting two weeks for database access was historically tolerable: "not great," but fine most of the time. Now a new hire can commit code on day one, and most companies are not structured for that. Etsy famously had engineers commit code on their first day, but Etsy expected it; other companies do not. Forsgren described one or two cases involving an intern who, because of policies and supply-chain problems, went about two weeks without a laptop and worked on a loaner. The intern had committed a lot of code before the laptop arrived. For one particularly secure piece of work, nobody could figure out how to make it fit, because the source didn't match where the systems expected it to come from. Forsgren's broader point was that AI is putting a spotlight on friction that used to be acceptable and now really slows teams down.

The DevEx framework: flow, cognitive load, feedback loops

The host asked Forsgren to explain the DevEx framework, which predates the AI wave. Forsgren described three pieces that support one another: flow state, cognitive load, and feedback loops. (Forsgren briefly forgot the third and turned to Laura Tacho in the audience for help.)

Feedback loops matter for flow. Waiting 20 minutes, or a week, for an answer or a review breaks concentration, and that also makes cognitive load harder to manage. Forsgren defined cognitive load as the work the brain has to do. Some of it is inherent: difficult tasks take brainpower. Easy things should not. Returning to a codebase after time away carries high cognitive load, while staying in it lets much of that work come "for free." An arcane 100-step process may be straightforward, but it still takes a lot of effort. Forsgren repeated a theme from other talks, that what's good for humans is good for systems. Well-structured code, good documentation, and clearly defined APIs help people and, by implication, agents.

Forsgren also cited Gloria Mark's research on focus, which they summarized as finding that humans max out at about three to four hours a day of truly deep work. They said it makes them laugh when executives demand eight hours of intense work: "Not with humans." The question now is how to use those few hours well when working with AI. For many people, deep work means blocking the calendar and focusing on one thing. Many AI models are highly interruptive, and Forsgren said individuals and organizations need to rethink how they manage cognitive load now that the nature of the work has changed.

Why faster feedback can be exhausting

The host described something many in the room have experienced. Agents such as Claude Code return results very quickly, yet working with them is tiring. Fast feedback loops were long an ideal, and now that they have arrived, cognitive load seems to go up. Is that good or bad?

Forsgren said it is simply different. Fast feedback used to help because, for example, a quick answer about a library let you continue with only a brief pause. Now feedback can arrive so fast that developers must rebuild their mental model "dozens of times in like a 30-minute period." Feedback that outpaces a person's ability to keep up, or tools that inject completions before the person is ready, can be disruptive. Forsgren said they sometimes turn the tool off to write for a while and then let it review. They stressed that this is largely an open question. People are starting to study it, but the environment six months earlier was very different from the current one.

The host added an example from a conversation with Mitchell Hashimoto, founder of HashiCorp. Hashimoto keeps an agent running alongside him but has turned off all notifications. He kicks off a task and checks back only when he is ready, regardless of when it finished. The host suggested that people may be discovering their own working styles.

Flow depends on people, not just tools

The host noted that Frictionless argues flow state depends on more than tooling. It also depends on psychological safety, project ownership, how technical decisions are made, and autonomy. If AI tooling makes flow easier to reach, might people still struggle because those other elements are missing?

"Tech is easy, people are hard," Forsgren replied. Getting into flow requires understanding what you are doing, having clear direction and goals, and knowing the purpose of the work, not just having a well-scoped feature, so that you can make informed decisions. It also requires the psychological safety to take a risk or ask a teammate for help.

Forsgren referred to Kent Beck calling AI tools "genies" and said working with a handful of genies is not the same as working with a handful of friends. The energy and the conversations differ, and the tools "just agree with us constantly." Forsgren joked that the AI always tells them how smart they are when they know what they said was dumb. They said they could not think of a paper or book they had written entirely alone. They typically draft most of it and then need someone to point out holes, what they're missing, and what makes sense in their head but not to others. In Forsgren's view, AI tools and agents are not there yet: sometimes they guess well, and sometimes they go in a completely orthogonal direction.

The deleted chapters of Frictionless

The host asked Forsgren to tell the story of the book's early drafts. Forsgren explained that as a former software engineer turned researcher, they wrote the way researchers write: lots of detail and background, with the point arriving around page 105. Early on, they produced several chapters on how to do research as a non-researcher, covering how to write good survey questions and how to talk to people to understand them. It was detailed and easy to follow, and it was also, in Forsgren's words, "100 pages that no one needs to read ever."

Forsgren threw it out and asked Abby to co-write the book, partly to be told when they went down a rabbit hole. The material was eventually turned into workbooks at the back of the book, with tables to fill in and checklists, to make it practical and actionable. Forsgren described this as their personal problem with flow: getting into flow and writing for the wrong audience or at the wrong altitude. The same happens when coding with agents. They sometimes give an agent a broad instruction that should have been more detailed and only discover this an hour in.

Is "wasted" effort actually necessary?

The host suggested a takeaway: that effort was not wasted. The time spent writing the wrong thing produced learning that led to a better result, something a "one-shot" approach would not have delivered. With agents able to generate endlessly, might humans need to relearn the value of such effort?

Forsgren agreed. Being able to clearly articulate a problem, thesis, or idea does not happen without trying. They then raised what they called an open question that interests them. When they did more coding, they had a feel for the system from working in it constantly. They didn't know the whole system, which was huge, but could whiteboard the core reasonably well. Now code changes so fast that it is unclear how people will build mental models of their systems. Forsgren said the goal is not only to reduce cognitive load or improve flow, but to help people understand their systems. As a visual person, when AI tools first came out they constantly asked for Mermaid diagrams because they needed the tool to "whiteboard with me." Forsgren expects this to look different for each person, but said our brains simply work better when we take that time.

What to measure: "it depends," and why

The host turned to the question engineering leaders face from CEOs and boards: we are paying a lot for AI tools, so how do we measure the return? Forsgren said this is literally their job, and the answer is always "it depends," joking that everyone could go into consulting.

Forsgren said it depends on the question being asked. When someone asks whether they are more productive, Forsgren asks what they mean by productive and what it would look like, noting that just as code has smells, "productivity smells" exist too. If the answer is lines of code or PRs, Forsgren asks what that tells them. Does it get a feature to customers faster? If yes, is it the right feature, and do we know that? Which part of the end-to-end process is being amplified? Do they also want more ideas, more code, more reviews, more of everything? Usually the answer is no.

Forsgren said measurement is still evolving and pointed to the SPACE framework as useful. They walked through it:

  • Satisfaction: how satisfied people are with the tool or process.
  • Performance: an outcome, such as quality.
  • Activity: anything you can count.
  • Collaboration and communication: between people, which Forsgren said is evolving, or between systems.
  • Efficiency and flow: being in flow, or simply the time it takes to move through the system.

Forsgren said they have heard many people at the event talk about velocity. Speed can be good, they said, but the question is what guardrails you want around it for quality and satisfaction. If you brute-force speed, "something's going to break."

Risk-based decisions instead of "all fast" or "all slow"

In response to a question from another session about sacrificing quality for speed, Forsgren said some teams do this, but intentionally. They don't describe it as sacrificing quality. They describe it as a risk-based decision. Forsgren gave the example of teams running rapid experiments. If they can ship an experiment in an hour to a very small percentage of users, they accept the risk of latency problems or crashes for that small group, then roll it back and have their answer. Forsgren argued that the right metrics support these deliberate trade-offs instead of defaulting to all fast or all slow. They still see teams that are "all slow," some in security, because they want to pump the brakes. Forsgren called that understandable but said those teams are now overwhelmed by how much there is to do.

Security and regulation in an agent world

The host described a conversation from a small event the day before. Non-developers are getting access to Claude Code and becoming very productive, and at one large publicly traded company a business developer built a useful sales tool and accidentally made it available to the whole world. It was caught in time. The host also relayed that David Cramer of Sentry said the annual security training developers tend to yawn through will need to become far more interactive and engaging, and everyone in the business will need to take it. The host concluded that it is a good time to be in security.

Forsgren agreed and said definitions of security are shifting: what counts as secure, which signals matter, and what levels of security are important. They pointed to regulations in certain countries requiring at least two people to review code before deployment. Over the past decade or two, improvements were made so that passing a set of automated checks and tests could count as one reviewer. Forsgren asked what that means now that agents are involved. They said the industry will need creative, meaningfully consistent solutions, and will need to educate not only the industry but also regulators.

Adoption and engagement as starting metrics

The host asked what a VP of engineering rolling out tools like Claude Code or Codex should measure tactically, while respecting developer privacy and avoiding junk data. Forsgren again said it depends, but that they tend to start with adoption, while admitting they don't like adoption metrics.

Their reasoning is that developers are "a gloriously cranky bunch" who won't use awful tools unless they have no alternative. Forsgren recalled a company insisting its developers had to use a particular CI/CD system; Forsgren bet $20 the developers were quietly spinning up Jenkins, and they were. Adoption therefore gives an early signal about satisfaction. It also matters because people who don't engage with a tool can't learn its capabilities or "kick the sides." They might love it at first and discover weaknesses later, or dislike it and never return.

Engagement comes next: how much people use the tool and for which tasks. Forsgren noted that earlier studies found the tools are used often for fairly straightforward work, and that it helps to watch how people use them. Beyond that, the metrics depend on the goal. Does the VP want usage, or speed? If speed, is it the inner coding loop or features end to end? The latter requires a much more holistic view of the whole system, especially in what Forsgren called "some magical agentic future" where agents are self-driving, which they said is "another metrics rant."

Explicit permission and executive sponsorship

The host mentioned that Rajeev Rajan, Atlassian's CISO and the next speaker, had told everyone at Atlassian they had his explicit permission to spend 10% of their time experimenting with AI systems. Is this top-down approach useful?

Forsgren said it is important in general. It is essentially communications and change management, "the really old-school stuff." They argued it matters especially now because there is so much fear, risk, and uncertainty around AI tools: will I be fired for using them, and what if I make a mistake? Across a handful of companies, Forsgren said they have seen explicit executive sponsorship make a big difference, not only in adoption but in people trying new things and feeling safe to fail within guardrails. They noted that some places have long given a prize to whoever takes down production. Without going that far, such attitudes help pressure-test the systems teams work in.

Supporting yourself through the change

The host asked why Frictionless ends with a chapter titled "Support yourself through challenging work." Forsgren said the final section covers supporting organizations, teams, and yourself through change. Many of the engineering leaders they interviewed for the book said supporting themselves was as important as giving their teams formal executive support for new tools. Any change, whether a new DevEx initiative or an AI rollout, involves a great deal that is unknown, and these are hard problems.

Forsgren recommended having a few people to talk to, what they attributed to "Rose Whitley" as having "your own board of directors." Such people let you safely say what is happening, for example: I have an exec review, I need an opinion, and I only understand half of this, so can you talk it through with me? Forsgren connected this to burnout, citing Christine Maslach's research. Burnout is not just overwork, which Forsgren described as simply getting tired. A critical component is misalignment with your values. Forsgren said that talking things through helped them, and others reported the same, to understand whether their values aligned with the work. Often they found they did, which relieved some of the pressure.

Seeing the system: a frictionless organization in two to three years

Finally, the host asked what a largely frictionless organization might look like in two to three years, and where attendees should start this week. Beyond reading the book, Forsgren noted that the workbooks are free online.

As a self-described metrics person, Forsgren framed the answer around data through a chain of conditions. Suppose there is a future where agents can self-drive and self-improve and organizations run better; Forsgren called this "maybe" true and "a stretch." For that to be true, agents must be able to see and understand the system and improve what needs fixing. For that, humans must be able to see and understand the system and act on it. And for that, the system must be visible, especially when teams move very fast. Right now humans act as a stopgap. People talk to one another and carry tacit knowledge, such as knowing that a problem in one area is usually about the build. Agents won't be able to do that, Forsgren said, "or if they are, like, we probably don't want that."

Forsgren's recommendation was to find easy, cheap ways to surface the signals that inform decisions. The instrumentation doesn't need to be heavyweight, and agents can probably help build it. The steps are to identify the touchpoints that matter, decide which signals you want, make them cheap to collect, and make sense of them, while accepting that they will change. The front end of software development, from idea and design through coding, has already been "smooshed" because teams can prototype so quickly. Forsgren fully expects parts of the outer loop to collapse too as more efficient approaches emerge. In the interim, they said, it helps to know where the quality gates and signals are and to watch where they shift, or whether they disappear, as the process collapses.

The host closed by returning to the personal board of directors: finding peers, ideally at other companies, forming a group chat, and comparing notes. With so much change, the host said, the only certainty is that things will keep changing, and peers in similar industries are likely to share a similar "it depends." Forsgren agreed and said this has been among the most helpful things for them. It lets them test whether an idea makes sense, whether what they see matches what others see, and, when it doesn't, whether things are actually different or people are just using different words. Beyond direct conversations, they keep a back channel with a handful of people they know, respect, and feel safe with, where they can drop a question at any time.