Data vs. Hype: What the Numbers Say About How Organizations Actually Win with AI

Open on YouTube ↗
Overview

In this keynote at The Pragmatic Summit, Laura Tacho asks how to separate the real effects of AI on software organizations from the hype. Tacho shares new industry benchmarks and case studies, then describes what successful organizations have in common. The position is consistent throughout. AI and agents are expanding what is possible and deserve genuine wonder, but AI doesn't fix organizational problems on its own. Companies that want results have to point it at real system-level problems and measure what happens. Tacho frames the whole talk with an analogy to the space race.

20 min read

Wonder and skepticism: the space race analogy

Tacho opens with two recent conversations: one with the CTO and co-founder of a small startup, and one with a principal engineering lead at a large, highly regulated bank. In both, the participants spent a long time trading ideas about what they were building, how AI had "brought back joy, the joy of coding," and how they couldn't build fast enough. Tacho links that feeling to a Carl Sagan quote that comes back throughout the talk: "Somewhere, something incredible is waiting to be known."

The age of space exploration, Tacho argues, had the same mix of awe and doubt. People asked why so much money was going to the moon when there were plenty of problems on Earth. Space exploration was also global, economic and political, and it was no silver bullet for humanity's problems. Even so, the moon landing was a defining moment that changed what people believed was possible.

AI, in Tacho's telling, carries the same two sides. On the optimistic side are claims of 100% productivity boosts, all code being written by AI, and perhaps the first one-person billion-dollar startup run by "a person with an idea and an army of agents." On the skeptical side are corporate doubts about economic impact given cost and environmental footprint, plus studies suggesting AI speeds work up in some circumstances and slows it down in others. Because "it's hard to know what's real," Tacho proposes keeping the sense of wonder while staying grounded, and "beating the hype by looking at data."

New benchmarks: adoption is high, time savings are flat

Tacho presents what are described as brand-new industry benchmarks, pulled just before the talk. The sample covers 121,000 developers at more than 450 companies, with data from November through February 1, 2026. Tacho notes that the figures aren't very surprising, because many haven't moved much since the previous quarter.

Adoption is close to universal. About 92.6% of developers use an AI coding assistant at least once a month, and about 75% use one at least once a week. Tacho adds a caveat about definitions. Most developers take "AI coding assistant" to mean tools such as Cursor, Codex, Copilot or Claude, and not necessarily ChatGPT, but the question is somewhat open-ended.

On time savings, developers self-report an average of about 4.08 hours saved per week due to AI tools. Tacho stresses that time savings is not the only measure of productivity impact, but calls it an important leading indicator. The figure is close to the Q2 2025 number, and the Q4 2025 number was roughly 3.6 to 3.7 hours. It has hovered around four hours for several quarters. Tacho links this to reports, including one from Google, of roughly a 10% productivity increase, and says the time-savings figure sits around that same 10% mark.

AI-authored code is climbing quickly

One metric is moving fast. Tacho's team calls it "AI-authored code": code written by AI that gets merged upstream or into a customer-facing environment without significant human intervention. In a sample of about 42,600 developers over the same period, AI-authored code is at 26.9% industry-wide, up from 22% the previous quarter. Tacho calls that a significant quarter-over-quarter change. Among developers who use AI daily, the share is above 30%. Almost a third of their code passes code review and reaches customers after being written by AI.

Onboarding time has been cut in half

Tacho calls onboarding one of their favorite applications of AI. They had suspected AI would help connect new people with information sooner, and then checked the data quarter by quarter. From Q1 2024 to Q4 2025, onboarding time fell by half. The metric is time to a developer's 10th pull request, which Tacho describes as a milestone the industry has largely agreed on. Plotted against rising AI usage, it "makes a really pretty graph."

Tacho says the effect isn't limited to new hires. There is also evidence that it helps engineers switching projects and non-engineers coming into a codebase. Tacho cites a separate study by Brian Houck at Microsoft, co-author of the SPACE framework for developer productivity. In Microsoft's context, performance at time-to-10th-PR stayed with an engineer through their first two years of tenure. Tacho's point is that faster onboarding is not only an early gain; it appears to persist for at least two years. Tacho reads this as evidence that AI can connect developers with their codebases, reduce cognitive load and speed up ramp-up.

Why averages mislead: AI as a multiplier

Right after presenting these benchmarks, Tacho warns against over-reading them. "Averages are just math." When the extremes move further apart, the average can stay the same. An average does not describe a typical experience or predict what will happen at any given company. The one thing Tacho says is common is that "there is no typical experience with AI," because every company has its own problems and culture.

Tacho returns to the space theme to explain this. After the Big Bang, the distance between objects keeps growing. For many organizations, AI has felt like a similar burst of energy that keeps pushing things apart. Organizational performance is multi-dimensional, and AI acts as an accelerator and multiplier, sending organizations toward different extremes depending on where they started.

Quality is the clearest example Tacho gives. In a sample of more than 67,000 developers from the same November–February window, some organizations are seeing twice as many customer-facing incidents, while others are seeing 50% fewer. Tacho explains the gap by the health of the underlying system. Organizations with healthy systems had those systems amplified by AI: they move faster with fewer incidents, better maintainability and more confidence in changes. Organizations that were already dysfunctional are now "dysfunctional faster."

High adoption, low transformation

Economic results are just as uneven. Tacho cites an MIT study published in July 2025, "The GenAI Divide," based on a survey of 152 organizations. The study describes steep drop-offs between piloting AI, putting it into production, and tying it to profit. Its conclusion, as Tacho relays it, is that the industry currently has high adoption but low transformation. Tacho notes that DORA's research also puts adoption around 90%.

Tacho's explanation is that transformation is uncomfortable. Organizations that gave up on cloud transformations and agile transformations are now giving up on AI transformations as well. Changing results requires looking at the whole organization, finding its problems and deciding to change something. Adoption is not impact, and using a tool does not advance an organization by itself. Tacho calls this an organizational problem that needs organizational change management. That runs against the hype, which suggested "experiment with AI and then something happens and then we profit."

According to Tacho, the tools were deployed mainly into individual coding tasks, and the MIT study found a very low ceiling on productivity gains when AI is applied only to a developer sitting at a desk. Tacho's conclusion: organizational results require thinking about AI at the organizational level, not at the level of individual coding tasks.

Agents expand the universe, and the hype along with it

Tacho then turns to agents and agentic workflows, which they describe as expanding the universe. The possibilities grow, and so do the promises and the hype. Tacho compares today's wilder experiments to the old prediction that everyone would be living on the moon with Jetsons-style flying cars. Examples include Gas Town (Tacho finds it "infinitely interesting" but jokingly advises against using it because "it is unhinged"), OpenClaw (also known under names like Moltbot and Clawdbot) and Ralph loops. Tacho enjoys this experimentation, but draws a line between building a nail-polish color-matching app while sitting at the nail salon and a multinational bank changing its revenue because of AI.

At a retreat with Martin Fowler and Kent Beck, discussed later in the talk, much of the effort went into connecting AI use to profit and P&L. Tacho says the group ended up asking about the value of innovation itself: was going to the moon worth it even though nobody lives there? Tacho argues that it was, while acknowledging this gets murky in a business context, where things have to be done under economic constraints and not only for the good of humankind.

Tacho's answer is that the point of going to the moon wasn't for everyone to live there. It was to improve life on Earth. Tacho lists sunglasses, space blankets, barcodes and quartz watches as benefits that came back from the space age. By the same logic, agents expand what can be built, how and for whom. Not every enterprise will build with Gas Town every day, and Tacho says that's fine, because experimentation pushes the boundary and suggests new ways to solve problems.

Agentic workflow usage and Codex figures

Tacho shares more first-time data on agentic usage, from a smaller sample: about 3,000 developers at six companies. Tacho cautions that these companies are ahead of the curve, since few organizations already instrument their agentic workflows with good telemetry. In this group, about 80% of developers use agentic workflows at least weekly, and more than 50% use them daily.

Tacho also presents figures about Codex, described as received the day before the talk. The Codex desktop app launched on February 2 and has passed a million downloads, with 60% user growth in the last week alone. GPT-5.3 Codex launched the previous Thursday. The system processes trillions of tokens per week. Inside OpenAI, 95% of developers use Codex to ship, and developers who use Codex ship about 60% more PRs per week than those using other AI tools. Tacho calls this "a data point, not the only data point," and presents it as a sign of the very high ceiling agentic tools may have.

Case study: Haven Headache and Migraine Center

To bring the discussion back to a non-AI company, Tacho highlights Haven Headache and Migraine Center, a San Francisco startup that set out to see whether headaches can be treated over Zoom. Tacho says it turns out they can.

In healthcare, Tacho explains, Haven's development team must carefully distinguish between using agents for durable code and for disposable code. As a small disruptor, Haven uses agentic workflows to rapidly prototype new patient workflows. For a patient portal, the team takes Linear and Figma artifacts, turns them into a PRD output as JSON, and then runs Ralph loops on it. According to Tacho, the output isn't "disposable AI slop." It consists of high-quality prototypes with excellent documentation and tests, produced much faster and at higher quality than the team would have managed by hand.

Haven is also training a HIPAA-compliant model on hundreds of thousands of symptom logs. Patients report symptoms by text message, and the model routes each message to the right next step, such as a medication refill or a follow-up appointment. Tacho reports that Haven has three times the industry average in customer satisfaction for this kind of healthcare tool, along with meaningful clinical outcomes: patients have fewer headache days per month, and their headaches are less severe.

Enterprise examples: manufacturing, Cisco, JPMorgan Chase

Tacho gives several enterprise examples. An unnamed enterprise manufacturing company uses agents purely for internal developer purposes; it used Copilot and Claude to build a developer portal that speeds up onboarding. At Cisco, 18,000 engineers use Codex daily for complex migrations and code review, and Tacho reports a 50% reduction in the time code review takes.

Tacho also recommends a paper from JPMorgan Chase describing its multi-agent framework for annotation, MAFA. Tacho describes it as building "a whole business of agents," a true multi-agent workflow similar to Gas Town in which each agent has a specific job. Some agents annotate interactions (what the intent was, whether it was an FAQ, and so on). Another set of agents reranks, calibrates and validates that output. Because multiple agents may disagree, consensus algorithms are needed to settle their outputs. Tacho predicts that consensus among agents will be a major problem to solve in 2026.

The retreat's conclusion: AI does not fix system problems

Tacho then describes the Future of Software Development retreat, hosted by Martin Fowler and Thoughtworks to mark the 25th anniversary of the Agile Manifesto. Gergely Orosz and several others at the summit also attended. The group spent a day and a half in the mountains talking almost entirely about agents: how to use them responsibly, ethically and sustainably, and how to use them in organizations. Steve Yegge was there, and there was plenty of hands-on experimentation, including work on Gas Town.

Despite all that, Tacho says the conclusion was that AI does not solve organizational systems problems. It can only help when it is applied to the system problem, and that requires first acknowledging the problem exists. In an informal conversation between sessions, Tacho, Kent Beck and Steve summarized their view roughly as follows. Organizations are constrained by human and systems-level problems, and they remain skeptical of any technology's promise to improve organizational performance unless those constraints are addressed first.

The risk, in Tacho's words, is that if those problems aren't addressed, "we will just take them to space with us." Going to the moon doesn't make pollution, garbage and traffic go away. So the question Tacho poses is not how to colonize Mars, but how to get real organizational impact from agents and AI.

What winning organizations do, part 1: set goals and measure

At the retreat, the group also discussed patterns shared by organizations that are succeeding with AI. Tacho presents three.

The first is having goals and measuring progress against them. Tacho rejects "spray and pray," meaning giving every developer a license and hoping for the best, and says there is a lot of evidence that it doesn't work. Winning organizations point AI experimentation at a concrete problem, set a goal and check whether they're reaching it. Tacho quotes Spock: "Insufficient facts always invite danger."

Tacho recognizes that measurement is hard, since developer productivity and engineering excellence were already difficult problems and AI now sits on top of them. To help, Tacho points to the AI Measurement Framework, co-authored with Abi Noda, CEO of DX, which complements their Core 4 framework. The goal is to track not only usage, adoption and utilization, but also whether AI is changing speed, developer experience, quality and innovation ratio. Cost is the last piece. Some organizations may be getting a good deal for now, but as tool prices keep rising, they need to ask whether the investment is still the right one.

What winning organizations do, part 2: developer experience matters more than ever

The second pattern is that developer experience matters more than ever. Tacho offers "very unconventional advice": call anything you were going to pitch to leadership as developer experience "agent experience" instead, and you'll get funding for it. Tacho admits it's funny but says it works. Fast feedback loops, clearly defined services, strong documentation, fast CI, and solid testing and quality practices are things engineers have asked for for decades, often getting turned down. These turn out to be exactly what makes agentic workflows succeed. Tacho finds it disheartening that organizations wouldn't spend this money for human engineers but will for "robot engineers," but urges the audience to take advantage of the opportunity.

Tacho supports this with data. Compared with the roughly four hours of weekly time savings, other developer experience factors are larger. AI time savings won't compensate for bad meeting culture, frequent interruptions, unplanned work or outages. The same goes for build and test wait time and dev environment toil. Taken together, Tacho says, coding speed-ups alone won't get organizations very far. The bigger gains come from pointing AI at those problems: using it to reduce meeting frequency, cut CI wait time or reduce dev environment toil. Winning organizations, Tacho says, "are putting DevEx at the center of their universe" and treating AI as a tool for fixing system-level problems.

These organizations also work at the organizational level. Outcomes like revenue, P&L and time to market require applying AI to workflows that span entire value streams, not leaving it to individual developers at their desks. Tacho goes back to the MIT study. The organizational barriers to AI adoption it found were not technical, and not about the models or the tools around them. They were change management; lack of executive sponsorship, as when executives push AI but have never opened Windsurf, Claude Code or Codex themselves; poor user experience; and unclear expectations about AI.

AI readiness models

For organizations that recognize these barriers, Tacho recommends two resources. The first is the DORA AI Capabilities Model, which Tacho describes as an AI readiness model built on a large amount of data from organizations DORA studies. DORA does much more than the four key metrics, and the model identifies correlations between organizational practices and good AI outcomes. One example Tacho gives: organizations with a clear, communicated AI stance do better than those without one. The model is at dora.dev, a new paper came out the previous month, and Tacho notes that Nathan, who leads DORA at Google Cloud, was at the summit.

The second is a Thoughtworks framework, which Tacho describes as a different take on the same idea, available among the white papers on thoughtworks.com. Tacho calls both well-researched, industry-backed readiness models, useful either for convincing leadership or for an internal audit of whether an organization is ready to benefit from its experimentation.

What winning organizations do, part 3: experiment on real customer problems

The third pattern is experimenting by solving real customer problems. Aiming for Mars is great, Tacho says, but it isn't sustainable for a whole organization to experiment that way: it costs too much, distracts from the core business and doesn't serve customers. Tacho encourages continuing to experiment, while focusing experimentation tightly on real customer problems, because that is where organizational results come from.

Closing: stay grounded, stay skeptical, stay human

Tacho closes by returning to "somewhere, something incredible is waiting to be known." There is a great deal of possibility in what can be built, how and for whom, and agents are accelerating it; Tacho says we are clearly in an age of exploration. The final request to the audience is to balance wonder and ambition (aiming for Mars and the moon colony) with the recognition that the problems on Earth still have to be solved, in the reality we live in. Tacho's closing words: "stay grounded, stay skeptical, stay human, most of all, stay pragmatic."