Data vs. Hype: What the Numbers Say About How Organizations Actually Win with AI
The Pragmatic EngineerIn this keynote at The Pragmatic Summit, Laura Tacho asks how to separate the real effects of AI on software organizations from the hype. Tacho shares new industry benchmarks and case studies, then describes what successful organizations have in common. The position is consistent throughout. AI and agents are expanding what is possible and deserve genuine wonder, but AI doesn't fix organizational problems on its own. Companies that want results have to point it at real system-level problems and measure what happens. Tacho frames the whole talk with an analogy to the space race.
Wonder and skepticism: the space race analogy
Tacho opens with two recent conversations: one with the CTO and co-founder of a small startup, and one with a principal engineering lead at a large, highly regulated bank. In both, the participants spent a long time trading ideas about what they were building, how AI had "brought back joy, the joy of coding," and how they couldn't build fast enough. Tacho links that feeling to a Carl Sagan quote that comes back throughout the talk: "Somewhere, something incredible is waiting to be known."
The age of space exploration, Tacho argues, had the same mix of awe and doubt. People asked why so much money was going to the moon when there were plenty of problems on Earth. Space exploration was also global, economic and political, and it was no silver bullet for humanity's problems. Even so, the moon landing was a defining moment that changed what people believed was possible.
AI, in Tacho's telling, carries the same two sides. On the optimistic side are claims of 100% productivity boosts, all code being written by AI, and perhaps the first one-person billion-dollar startup run by "a person with an idea and an army of agents." On the skeptical side are corporate doubts about economic impact given cost and environmental footprint, plus studies suggesting AI speeds work up in some circumstances and slows it down in others. Because "it's hard to know what's real," Tacho proposes keeping the sense of wonder while staying grounded, and "beating the hype by looking at data."
New benchmarks: adoption is high, time savings are flat
Tacho presents what are described as brand-new industry benchmarks, pulled just before the talk. The sample covers 121,000 developers at more than 450 companies, with data from November through February 1, 2026. Tacho notes that the figures aren't very surprising, because many haven't moved much since the previous quarter.
Adoption is close to universal. About 92.6% of developers use an AI coding assistant at least once a month, and about 75% use one at least once a week. Tacho adds a caveat about definitions. Most developers take "AI coding assistant" to mean tools such as Cursor, Codex, Copilot or Claude, and not necessarily ChatGPT, but the question is somewhat open-ended.
On time savings, developers self-report an average of about 4.08 hours saved per week due to AI tools. Tacho stresses that time savings is not the only measure of productivity impact, but calls it an important leading indicator. The figure is close to the Q2 2025 number, and the Q4 2025 number was roughly 3.6 to 3.7 hours. It has hovered around four hours for several quarters. Tacho links this to reports, including one from Google, of roughly a 10% productivity increase, and says the time-savings figure sits around that same 10% mark.
AI-authored code is climbing quickly
One metric is moving fast. Tacho's team calls it "AI-authored code": code written by AI that gets merged upstream or into a customer-facing environment without significant human intervention. In a sample of about 42,600 developers over the same period, AI-authored code is at 26.9% industry-wide, up from 22% the previous quarter. Tacho calls that a significant quarter-over-quarter change. Among developers who use AI daily, the share is above 30%. Almost a third of their code passes code review and reaches customers after being written by AI.
Onboarding time has been cut in half
Tacho calls onboarding one of their favorite applications of AI. They had suspected AI would help connect new people with information sooner, and then checked the data quarter by quarter. From Q1 2024 to Q4 2025, onboarding time fell by half. The metric is time to a developer's 10th pull request, which Tacho describes as a milestone the industry has largely agreed on. Plotted against rising AI usage, it "makes a really pretty graph."
Tacho says the effect isn't limited to new hires. There is also evidence that it helps engineers switching projects and non-engineers coming into a codebase. Tacho cites a separate study by Brian Houck at Microsoft, co-author of the SPACE framework for developer productivity. In Microsoft's context, performance at time-to-10th-PR stayed with an engineer through their first two years of tenure. Tacho's point is that faster onboarding is not only an early gain; it appears to persist for at least two years. Tacho reads this as evidence that AI can connect developers with their codebases, reduce cognitive load and speed up ramp-up.
Why averages mislead: AI as a multiplier
Right after presenting these benchmarks, Tacho warns against over-reading them. "Averages are just math." When the extremes move further apart, the average can stay the same. An average does not describe a typical experience or predict what will happen at any given company. The one thing Tacho says is common is that "there is no typical experience with AI," because every company has its own problems and culture.
Tacho returns to the space theme to explain this. After the Big Bang, the distance between objects keeps growing. For many organizations, AI has felt like a similar burst of energy that keeps pushing things apart. Organizational performance is multi-dimensional, and AI acts as an accelerator and multiplier, sending organizations toward different extremes depending on where they started.
Quality is the clearest example Tacho gives. In a sample of more than 67,000 developers from the same November–February window, some organizations are seeing twice as many customer-facing incidents, while others are seeing 50% fewer. Tacho explains the gap by the health of the underlying system. Organizations with healthy systems had those systems amplified by AI: they move faster with fewer incidents, better maintainability and more confidence in changes. Organizations that were already dysfunctional are now "dysfunctional faster."
High adoption, low transformation
Economic results are just as uneven. Tacho cites an MIT study published in July 2025, "The GenAI Divide," based on a survey of 152 organizations. The study describes steep drop-offs between piloting AI, putting it into production, and tying it to profit. Its conclusion, as Tacho relays it, is that the industry currently has high adoption but low transformation. Tacho notes that DORA's research also puts adoption around 90%.
Tacho's explanation is that transformation is uncomfortable. Organizations that gave up on cloud transformations and agile transformations are now giving up on AI transformations as well. Changing results requires looking at the whole organization, finding its problems and deciding to change something. Adoption is not impact, and using a tool does not advance an organization by itself. Tacho calls this an organizational problem that needs organizational change management. That runs against the hype, which suggested "experiment with AI and then something happens and then we profit."
According to Tacho, the tools were deployed mainly into individual coding tasks, and the MIT study found a very low ceiling on productivity gains when AI is applied only to a developer sitting at a desk. Tacho's conclusion: organizational results require thinking about AI at the organizational level, not at the level of individual coding tasks.
Agents expand the universe, and the hype along with it
Tacho then turns to agents and agentic workflows, which they describe as expanding the universe. The possibilities grow, and so do the promises and the hype. Tacho compares today's wilder experiments to the old prediction that everyone would be living on the moon with Jetsons-style flying cars. Examples include Gas Town (Tacho finds it "infinitely interesting" but jokingly advises against using it because "it is unhinged"), OpenClaw (also known under names like Moltbot and Clawdbot) and Ralph loops. Tacho enjoys this experimentation, but draws a line between building a nail-polish color-matching app while sitting at the nail salon and a multinational bank changing its revenue because of AI.
At a retreat with Martin Fowler and Kent Beck, discussed later in the talk, much of the effort went into connecting AI use to profit and P&L. Tacho says the group ended up asking about the value of innovation itself: was going to the moon worth it even though nobody lives there? Tacho argues that it was, while acknowledging this gets murky in a business context, where things have to be done under economic constraints and not only for the good of humankind.
Tacho's answer is that the point of going to the moon wasn't for everyone to live there. It was to improve life on Earth. Tacho lists sunglasses, space blankets, barcodes and quartz watches as benefits that came back from the space age. By the same logic, agents expand what can be built, how and for whom. Not every enterprise will build with Gas Town every day, and Tacho says that's fine, because experimentation pushes the boundary and suggests new ways to solve problems.
Agentic workflow usage and Codex figures
Tacho shares more first-time data on agentic usage, from a smaller sample: about 3,000 developers at six companies. Tacho cautions that these companies are ahead of the curve, since few organizations already instrument their agentic workflows with good telemetry. In this group, about 80% of developers use agentic workflows at least weekly, and more than 50% use them daily.
Tacho also presents figures about Codex, described as received the day before the talk. The Codex desktop app launched on February 2 and has passed a million downloads, with 60% user growth in the last week alone. GPT-5.3 Codex launched the previous Thursday. The system processes trillions of tokens per week. Inside OpenAI, 95% of developers use Codex to ship, and developers who use Codex ship about 60% more PRs per week than those using other AI tools. Tacho calls this "a data point, not the only data point," and presents it as a sign of the very high ceiling agentic tools may have.
Case study: Haven Headache and Migraine Center
To bring the discussion back to a non-AI company, Tacho highlights Haven Headache and Migraine Center, a San Francisco startup that set out to see whether headaches can be treated over Zoom. Tacho says it turns out they can.
In healthcare, Tacho explains, Haven's development team must carefully distinguish between using agents for durable code and for disposable code. As a small disruptor, Haven uses agentic workflows to rapidly prototype new patient workflows. For a patient portal, the team takes Linear and Figma artifacts, turns them into a PRD output as JSON, and then runs Ralph loops on it. According to Tacho, the output isn't "disposable AI slop." It consists of high-quality prototypes with excellent documentation and tests, produced much faster and at higher quality than the team would have managed by hand.
Haven is also training a HIPAA-compliant model on hundreds of thousands of symptom logs. Patients report symptoms by text message, and the model routes each message to the right next step, such as a medication refill or a follow-up appointment. Tacho reports that Haven has three times the industry average in customer satisfaction for this kind of healthcare tool, along with meaningful clinical outcomes: patients have fewer headache days per month, and their headaches are less severe.
Enterprise examples: manufacturing, Cisco, JPMorgan Chase
Tacho gives several enterprise examples. An unnamed enterprise manufacturing company uses agents purely for internal developer purposes; it used Copilot and Claude to build a developer portal that speeds up onboarding. At Cisco, 18,000 engineers use Codex daily for complex migrations and code review, and Tacho reports a 50% reduction in the time code review takes.
Tacho also recommends a paper from JPMorgan Chase describing its multi-agent framework for annotation, MAFA. Tacho describes it as building "a whole business of agents," a true multi-agent workflow similar to Gas Town in which each agent has a specific job. Some agents annotate interactions (what the intent was, whether it was an FAQ, and so on). Another set of agents reranks, calibrates and validates that output. Because multiple agents may disagree, consensus algorithms are needed to settle their outputs. Tacho predicts that consensus among agents will be a major problem to solve in 2026.
The retreat's conclusion: AI does not fix system problems
Tacho then describes the Future of Software Development retreat, hosted by Martin Fowler and Thoughtworks to mark the 25th anniversary of the Agile Manifesto. Gergely Orosz and several others at the summit also attended. The group spent a day and a half in the mountains talking almost entirely about agents: how to use them responsibly, ethically and sustainably, and how to use them in organizations. Steve Yegge was there, and there was plenty of hands-on experimentation, including work on Gas Town.
Despite all that, Tacho says the conclusion was that AI does not solve organizational systems problems. It can only help when it is applied to the system problem, and that requires first acknowledging the problem exists. In an informal conversation between sessions, Tacho, Kent Beck and Steve summarized their view roughly as follows. Organizations are constrained by human and systems-level problems, and they remain skeptical of any technology's promise to improve organizational performance unless those constraints are addressed first.
The risk, in Tacho's words, is that if those problems aren't addressed, "we will just take them to space with us." Going to the moon doesn't make pollution, garbage and traffic go away. So the question Tacho poses is not how to colonize Mars, but how to get real organizational impact from agents and AI.
What winning organizations do, part 1: set goals and measure
At the retreat, the group also discussed patterns shared by organizations that are succeeding with AI. Tacho presents three.
The first is having goals and measuring progress against them. Tacho rejects "spray and pray," meaning giving every developer a license and hoping for the best, and says there is a lot of evidence that it doesn't work. Winning organizations point AI experimentation at a concrete problem, set a goal and check whether they're reaching it. Tacho quotes Spock: "Insufficient facts always invite danger."
Tacho recognizes that measurement is hard, since developer productivity and engineering excellence were already difficult problems and AI now sits on top of them. To help, Tacho points to the AI Measurement Framework, co-authored with Abi Noda, CEO of DX, which complements their Core 4 framework. The goal is to track not only usage, adoption and utilization, but also whether AI is changing speed, developer experience, quality and innovation ratio. Cost is the last piece. Some organizations may be getting a good deal for now, but as tool prices keep rising, they need to ask whether the investment is still the right one.
What winning organizations do, part 2: developer experience matters more than ever
The second pattern is that developer experience matters more than ever. Tacho offers "very unconventional advice": call anything you were going to pitch to leadership as developer experience "agent experience" instead, and you'll get funding for it. Tacho admits it's funny but says it works. Fast feedback loops, clearly defined services, strong documentation, fast CI, and solid testing and quality practices are things engineers have asked for for decades, often getting turned down. These turn out to be exactly what makes agentic workflows succeed. Tacho finds it disheartening that organizations wouldn't spend this money for human engineers but will for "robot engineers," but urges the audience to take advantage of the opportunity.
Tacho supports this with data. Compared with the roughly four hours of weekly time savings, other developer experience factors are larger. AI time savings won't compensate for bad meeting culture, frequent interruptions, unplanned work or outages. The same goes for build and test wait time and dev environment toil. Taken together, Tacho says, coding speed-ups alone won't get organizations very far. The bigger gains come from pointing AI at those problems: using it to reduce meeting frequency, cut CI wait time or reduce dev environment toil. Winning organizations, Tacho says, "are putting DevEx at the center of their universe" and treating AI as a tool for fixing system-level problems.
These organizations also work at the organizational level. Outcomes like revenue, P&L and time to market require applying AI to workflows that span entire value streams, not leaving it to individual developers at their desks. Tacho goes back to the MIT study. The organizational barriers to AI adoption it found were not technical, and not about the models or the tools around them. They were change management; lack of executive sponsorship, as when executives push AI but have never opened Windsurf, Claude Code or Codex themselves; poor user experience; and unclear expectations about AI.
AI readiness models
For organizations that recognize these barriers, Tacho recommends two resources. The first is the DORA AI Capabilities Model, which Tacho describes as an AI readiness model built on a large amount of data from organizations DORA studies. DORA does much more than the four key metrics, and the model identifies correlations between organizational practices and good AI outcomes. One example Tacho gives: organizations with a clear, communicated AI stance do better than those without one. The model is at dora.dev, a new paper came out the previous month, and Tacho notes that Nathan, who leads DORA at Google Cloud, was at the summit.
The second is a Thoughtworks framework, which Tacho describes as a different take on the same idea, available among the white papers on thoughtworks.com. Tacho calls both well-researched, industry-backed readiness models, useful either for convincing leadership or for an internal audit of whether an organization is ready to benefit from its experimentation.
What winning organizations do, part 3: experiment on real customer problems
The third pattern is experimenting by solving real customer problems. Aiming for Mars is great, Tacho says, but it isn't sustainable for a whole organization to experiment that way: it costs too much, distracts from the core business and doesn't serve customers. Tacho encourages continuing to experiment, while focusing experimentation tightly on real customer problems, because that is where organizational results come from.
Closing: stay grounded, stay skeptical, stay human
Tacho closes by returning to "somewhere, something incredible is waiting to be known." There is a great deal of possibility in what can be built, how and for whom, and agents are accelerating it; Tacho says we are clearly in an age of exploration. The final request to the audience is to balance wonder and ambition (aiming for Mars and the moon colony) with the recognition that the problems on Earth still have to be solved, in the reality we live in. Tacho's closing words: "stay grounded, stay skeptical, stay human, most of all, stay pragmatic."
Today I wanted to have a really pragmatic and down-to-earth conversation about AI, what is actually happening in our organizations, what you can expect to happen and how agents are changing the game. And I thought in order to have this really pragmatic down-to-earth conversation, I wanted to take us to space.
I do see a lot of parallels between the age of exploration and the space race and the age of AI. So last week I was talking with a CTO co-founder of a small startup. I was also talking with a principal engineering lead at a very big bank, highly regulated, and we sat for a solid 15 minutes talking about all the cool stuff that we were building, how it's brought back joy, the joy of coding, and we just had all of these ideas and it seems like we couldn't build fast enough. There is so much to learn and so much to build. And it reminds me of this really lovely quote from Carl Sagan that I love: that somewhere something incredible is waiting to be known. And I think that this quote really captures what a lot of us are feeling about AI and about the experimentation and just the possibility that is out there.
This is the same feeling that we had with the age of space exploration, going to the moon, going to Mars. But it didn't come without skepticism. Why spend all of this money experimenting and going to the moon when we had lots of problems to solve here on Earth?
We had a lot of wonder, but we also had a lot of skepticism because space exploration wasn't just about science. It was also global. It was economic. It was political. Space wasn't a silver bullet to solve all of the problems that we had with humanity. But we also can't deny that when a man landed on the moon that it was a very pivotal defining moment for all of humankind and had a sense of wonder and had the world in awe. It was about redefining what was possible.
And similarly, we have a lot of wonder and a lot of optimism and a lot of promise about AI. We can talk about productivity boosts and all of the hype around, you know, 100% productivity boost and all of our code being written by AI. We have the promise of perhaps the first single-person billion-dollar startup with a person with an idea and an army of agents. There's a lot of optimism out there. Similarly, there's a lot of skepticism. There's a lot of skepticism in the corporate world about the real economic impact of AI given how expensive it is, given the environmental impact. There's also a lot of skepticism in a lot of different studies about the real productivity impact. In certain circumstances, it can really accelerate and in other circumstances it can actually slow us down and get in our way. It's hard to know what's real.
But as technology changes, it is really good and fine to have that sense of wonder of exploring the universe while also realizing that we have problems here on Earth to solve. We have to learn how to balance that sense of wonder and curiosity with the acknowledgment that we are living in reality and we need to keep our feet firmly planted on this earth. We need to understand how these experiments are actually going to apply to everyday companies. How are we actually going to improve the world around us? We need to keep the sense of wonder while also balancing it with pragmatism and beating the hype by looking at data.
And so that's what I want to do right now. I'm going to share some brand new AI industry benchmarks with you. This is new data that no one has ever seen ever before in the world. Okay, it's coming right now to you. I just pulled these down. I guess the static team has seen them because they saw the preview of my slides, but aside from them, no one has ever seen it. This is though not really surprising because a lot of these numbers have not changed very much from the last quarter.
So, what we're looking at here is a sample of 121,000 developers at over 450 companies. This data was pulled from November through February 1st, 2026. I really just did this. We're sitting around 92.6% of developers are using an AI coding assistant at least once a month to get their work done. And about 75% of developers are using an AI coding assistant at least once a week. When I say AI coding assistant, most developers define that as Cursor, Codex, Copilot, you know, Claude, not ChatGPT necessarily. But it is a bit open-ended, so keep that in mind.
When it comes to time savings, time savings is not the only measure of productivity impact, but it is an important signal. It's a good leading indicator. We're sitting around 4.08 self-reported hours saved due to AI tool usage per week per developer. This is not all too different from the number that came in Q2 of 2025. And then the number for Q4 of 2025 was about 3.6 or 3.7. So this is kind of hovering around the 4 hour mark. And there's been a few articles, for example, from Google in the last year citing about a 10% productivity increase. And if we look at it in terms of time savings, we're kind of hovering around that 10% mark. It hasn't changed dramatically over the last few quarters.
What is changing and what is moving up very quickly is the amount of code getting merged upstream or in a customer-facing environment that was written by AI that was merged without significant human intervention. We call that AI-authored code. And in a sample of around 42,600 developers from that same time frame, November 1st to February 1st, 2026, we're at about 26.9% industrywide for all of these developers. That's how much code is hitting production that was AI-authored. This is moving up from 22% in the last quarter, which is actually a pretty significant change quarter over quarter. And we can see that daily users of AI have crested over that 30% mark. So almost a third of their code is being written by AI that is actually being merged, passing through code review and getting into a customer-facing environment.
One of my favorite use cases for applying AI is to onboarding. And I had a bit of a hunch that AI was going to be a great tool for onboarding, helping connect people with information earlier and sooner. And I have all of this data and I thought, let me look at this quarter over quarter. And in fact, if we look at Q1 of 2024 all the way over here on the left side and fast forward to Q4 of 2025, we have a half reduction in onboarding time. This is looking at the time to 10th PR. So by the time a developer hits their 10th PR, that's a pretty important onboarding milestone that the industry has mostly aligned on in terms of onboarding, and that has been cut in half now. And when we correlate that with the uptick of AI usage, it makes a really pretty graph. AI is fantastic for onboarding.
And this is not just brand new hires to your company. We've also seen plenty of evidence that this is for engineers who are moving projects or even non-engineers coming onboarding into projects. What's really important about this number is that there was a separate study done by Brian Houck at Microsoft. He's the co-author of the SPACE framework of developer productivity, and they found in Microsoft's context that the time to 10th PR, actually that performance sticks with an engineer for their first two years of tenure. So if you onboard faster, that productivity gain isn't just onboarding, it actually sticks with them for at least two years after they have started at the company. So this is a very important and significant trend that we're seeing here with using AI to connect developers, reduce cognitive load, and get them onboarded more quickly into their code bases.
One thing that's really important for me to call out, although I have just shared with you industry benchmarks, averages are just math. And as the poles move further away from each other, the average stays the same. Average does not mean typical. It does not mean what is going to happen to you. And it doesn't mean what a common experience is. One thing that is absolutely true, one thing that is common is that there is no typical experience with AI. There is no typical experience with AI. It is extremely different in every single company because every company has their own problems and their own culture.
This uneven impact can take us back to space for just a minute. So we can go back to the origins of the universe. We had the big bang and there was this massive release of energy, and as this energy released, the space and time in between objects grows bigger, right? Things are moving apart. And for a lot of us, the emergence of AI and AI coming into our organizations and in the industry has felt a lot like this big bang. We've had this explosive release of energy in the center of our world and things keep moving apart. Organizational performance is multi-dimensional and these organizations are just going off into different extremes based on what they were doing before. AI is an accelerator. It's a multiplier and it is moving organizations off in different directions.
The best example I can share with you of this is quality. Okay, so in this case, this is not every organization, but some organizations are facing twice as many customer-facing incidents, and this is from a sample of over 67,000 developers from Q1. So that same time frame of November to February. So just looking in that time frame, organizations are experiencing twice as many customer-facing incidents. At the same time, at the same time companies are also experiencing 50% fewer incidents. So some companies have used AI, they have a really healthy system, it has amplified that system, they are seeing fewer incidents, they're moving faster, they are accelerating with higher quality, higher code maintainability, higher change confidence. On the other side though, blasting off into the other part of the universe, we have organizations who were dysfunctional already. No, they're more dysfunctional. They're dysfunctional and dysfunctional faster. Okay.
Similarly to this uneven impact, organizations are seeing really uneven results economically from using AI. There are a lot of steep drop-offs when it comes to using AI in a pilot context to production and then actually trying to tie it to profit. This is from an MIT study that was published in July of 2025 called the GenAI Divide. And what the study concluded, they did a survey of 152 organizations, was that right now where we are in the industry is that we have really high adoption, right? That 92.6 number. DORA also does its own research. We're hovering around that 90% adoption number. High adoption but actually low transformation, because as it turns out transformation is really uncomfortable, and organizations that were ready to give up on the cloud transformation, on the agile transformation, are also giving up on their AI transformations. It is really, really difficult to look at your whole organization and look at the problems and think we got to change something about this, and that is what organizations need to do in order to actually see change to their bottom line.
All of this to say, back to my previous point, we have 92.6 percent adoption among developers in our industry, but adoption doesn't mean impact. Using the tool doesn't mean that it's going to actually advance your organization or do anything. It is an organizational problem that needs organizational change management. But that's not really what we were promised with all of the hype, which was like, hey, experiment with AI and then something happens and then we profit.
What happens though is that these tools were primarily deployed into individual coding tasks. And what this MIT study found in this high adoption, low transformation is that when we apply it only to the surface area of a developer sitting at their desk, there is a very, very low ceiling of productivity gain. This is an organizational problem. If we want organizational results, we have to think about it on an organizational level, not on a coding task level.
Fortunately, our universe is expanding right now, and that expanding is coming through the use of agents in agentic workflows. Our universe is getting bigger and so are all of the promises and all of the hype, but so is the possibility.
So, let's go back to the moon landing, right? Like the ultimate hype was that we're all going to be living on the moon by now in flying cars, like Jetsons style. Similarly here we have a little bit of crazy ideas. Gas Town, if any of you have used it, there's just so much crazy stuff to do right now. Gas Town is infinitely interesting to me. There are so many interesting things. Disclaimer, don't use Gas Town. It is unhinged. We've got OpenClaw, Moltbot, Clawdbot, whatever it's called. We've got Ralph loops. We've got all the stuff, right? There is so much experimentation and so much fun. It's just really fun to build. But me building my nail polish matching color scheme app while I'm sitting at the nail salon is not the same as a multinational bank being able to change their revenue because of AI. Those are really different things.
And I was at this retreat with Martin Fowler and Kent, who are I think back there. Hello. We'll talk about that a bit more later. We spent a lot of time trying to connect AI and the use of AI to bottom line, to profit, to P&L. And interestingly, kind of where we landed at the end was this question of like, what is the value of innovation? Was it still valuable to go to the moon even though I'm not really located on the moon right now? And I would argue that yes, it is valuable to innovate. And that can get into some murky area because this is a business, right? This isn't just society and doing things for the good of humankind. We have to do them in an economic context and that can get a little bit tricky.
So when we think about this quote, somewhere something incredible is waiting to be known, there is a sense of wonder, and AI and space are both the age of exploration and it is so exciting. But the point of going to the moon wasn't that we all need to live on the moon. In fact, the point of going to the moon and the point of exploring and doing all this crazy stuff was to improve life on Earth. It was to use the space exploration and all of this wonder to apply it to the systems-level problems that we had back on Earth. Not everyone wants to live on the moon, but we have sunglasses, we have space blankets, we have barcodes, we have quartz watches. We have so much technology and so many improvements back on Earth because of this crazy age of exploration where we all went to space. Even though we're not living on the moon, we've still used the lessons and applied it to our systems back here on Earth.
And so thinking about agentic workflows, agents expand the possibilities of what we can build, how we can build it, and who we can build it for. Not everyone goes to the moon, and it's okay not to go to the moon. Not everyone is going to be building crazy stuff with Gas Town every day in your enterprise context. And that's also okay, because the experimentation helps push the boundary of what's possible and helps us think about solving problems in new ways.
So, let's talk a little bit about how agents are being used in the industry right now. Again, this is new data that I'm sharing for the first time here. Agentic use is on the rise. There's not a lot of companies, honestly, that are so far ahead of the curve that they're already instrumenting their agentic use cases with really good telemetry. This sample is a little bit smaller. It's around 3,000 developers at six companies. Keep in mind, these companies are ahead of the curve. They're already instrumenting their agentic workflows with telemetry. We have about 80% of developers using these agentic workflows at least once a week, with over 50% using agentic workflows every single day to get their work done.
We talked about Codex I think in the previous panel. So on February 2nd, the Codex desktop app was released and since then there's been over a million downloads by now. I got this data yesterday. I'm sure it's quite different by now. There's been a 60% growth in users just in the last week. They also launched GPT-5.3-Codex last Thursday. They're processing
trillions of tokens per week. Internally at OpenAI, 95% of developers are using Codex to ship stuff. And of the developers who are using Codex versus other AI tools, the developers who use Codex are shipping about 60% more PRs per week, which is very interesting. A data point, not the only data point, but it just speaks to the very high ceiling, the high possibility, the sense of wonder that we have with building all of the stuff with cool new tools like agentic workflows.
I want to bring it back to a non-AI startup though. I want to highlight Haven Headache and Migraine Center. So, this is a company that's based here in San Francisco, actually just a few blocks away. Haven set out to answer the question, can we solve headaches with Zoom? And it turns out you can. So if you're a headache sufferer, this might be useful for you to learn about.
In healthcare, it's really, really crucial for Haven and their development team to distinguish between using agents for durable code or disposable code. One of the things that they're doing that's very cool since they are a disruptor, they are a small startup, is using agentic workflows to rapidly prototype new custom, like new patient workflows. So they're working on a patient portal building with Ralph loops, taking Linear and Figma artifacts, changing it into a PRD, you know, spitting that out in JSON and then just having Ralph loops run. What they're getting though isn't garbage disposable AI slop. What they're getting is really high quality prototypes with really excellent documentation, excellent tests, much higher quality at a way faster rate than they would have if they would have built it by hand the old-fashioned way.
The other thing that they're doing that I really admire is improving the standard of care for their patients by training a HIPAA-compliant model on hundreds of thousands of symptom logs. So Haven meets you where you're at. You get a text message, you can log your symptoms and then they can instrument your care, figure out what needs to happen from there. So they're training a HIPAA-compliant model on hundreds of thousands of these messages so that those messages can be routed to, you know, medication refill or schedule follow-up appointment, just meets you where you are. And the result of this is that they have 3x the industry average in customer satisfaction for a healthcare tool like this, but also real, real meaningful clinical outcomes. So their patients have fewer headache days per month and also the severity of their headaches is much less severe. So good job, Haven.
In the enterprise, there are lots of examples of big enterprise companies experimenting with agent workflows. So there's an enterprise manufacturing company that's using it for solely internal developer purposes. They used Copilot and Claude to build out a dev portal to accelerate developer onboarding. At Cisco, there's 18,000 engineers using Codex daily. They're using Codex for complex migrations and also code review, leading to a 50% reduction in the amount of time it takes to do code review.
There's a really cool paper as well by JPMorgan Chase's multi-agent framework for annotation, MAFA. If you Google that, you can find the source paper. It's really fascinating. What they're doing is building out like a whole business of agents. So like a true multi-agent workflow similar to Gas Town where each agent has a special job to do. What they're also doing in this model is introducing consensus among the agents. So they're taking all of these interactions and then they're annotating them. This was, you know, the intent, what was it, an FAQ, what were all these interactions. The agents are annotating them and then there's another set of agents who are responsible for reranking and calibrating and validating the output, and then of course we have to introduce consensus algorithms to the party because now we have multiple agents with maybe multiple different opinions about things. This is really fascinating and I believe consensus among agents is going to be a huge problem to solve in 2026.
I spoke about this retreat. I was lucky enough to be invited by Martin Fowler and Thoughtworks to the Future of Software Development retreat celebrating the 25th anniversary of the Agile Manifesto. Gergely joined me. A few other folks who are here also joined me. We spent a day and a half up in the mountains talking about agents. That's really all we talked about, about using agents responsibly, ethically, sustainably, how we can use them for organizations. And our conclusion, even though there was so much interesting stuff, Steve Yegge was there, we were working on Gas Town things, like there was a lot of experimentation happening, but the conclusion that we came to was that AI does not solve organizational systems problems. It only can do that when you apply AI to the system problem, which means you need to acknowledge that the system problem exists in the first place. AI is not a magic silver bullet. Even though things like Gas Town exist, even though there is so much sense of curiosity and wonder in the universe.
We kind of had a sort of off-the-cuff conversation. Kent Beck, Steve and I were just catching up outside of one of the sessions in between conversations. And here's sort of where we summarized our thoughts. Organizations are constrained by human and systems level problems. We remain skeptical of the promise of any technology to improve organizational performance without first addressing those human and systems level constraints. We remain skeptical and we also remain human, because the risk is if we don't address the systems level problems, we will just take them to space with us. We will just take them to space with us. We're not actually going to solve the human factors that are the driving force behind all of the constraints that organizations have right now. We can apply AI to those problems, but we still need to solve them. We can't just go to the moon and expect that pollution and garbage and traffic aren't going to be a problem anymore.
And so the question is not how to colonize Mars, but the question is how to get real organizational impact with agents and AI. At this retreat, we also talked a lot about common factors that we see. What do we see organizations doing? What are the common patterns that is kind of like the secret to winning? What do they have in common?
The first one is that organizations who win with AI and are winning with AI have goals and they measure their progress against those goals. Spray and pray does not work. Spray and pray, what I mean by that is just giving all of your developers licenses and hoping for the best. It does not work. I can say that very, very clearly. I have a lot of evidence that does not work. If you can point AI innovation and that experimentation to a problem, have a concrete goal and then measure if you're reaching that goal, that is what winning organizations are doing right now. Because as Spock has told us, insufficient facts always invite danger. We need to measure things. We need to have data.
And I know this is something that's really difficult for a lot of organizations right now because developer productivity and engineering excellence are also really hard problems. And this is happening all at the intersection. So I have something that can help you if that is a problem that you're facing in your organization. This is the AI Measurement Framework. This is a framework that I co-authored with Abi Noda, who's the CEO of DX. This complements our Core 4 framework, which some of you might have heard of; otherwise it's in the impact column here. What we're looking to do is track not just usage and adoption and utilization of AI, but then also translate that into real organizational impact. Is this changing your speed, your developer experience, your quality, your innovation ratio? Those are really important questions to connect the adoption to impact. Finally, we have to look at the cost. Are we getting a good deal? Maybe some of us are for now. And we need to understand, as the cost of these tools keeps going up and up, is the investment the right one?
The second thing that is helping organizations win is that developer experience matters now more than ever. Here is a piece of very unconventional advice that I'll give you: just anything that you were going to talk about with your leadership team about developer experience, just call it agent experience and you'll get money for it. It's funny but it works. It works because developer experience, feedback loops, you know, clearly defined services, great documentation, fast CI, these are all things that we have been screaming about for decades, literally, and we've been begging for pennies from our organizations to please let us invest, please let us invest in developer experience. And we've been told no over and over again. Come to find out, in fact, these are the things that make AI really successful. We need to have really solid testing and quality practices. We need to have great documentation. These are critical for agentic workflows. It is disheartening that we didn't want to spend the money when it came to human engineers, but when it comes to robot engineers, we're okay with it. But that is the world that we live in, and let's capitalize on our opportunity.
So DevEx matters more than ever. In fact, when we look at the data right now, remember we're hovering around that 4-hour mark for time savings. When we look at all of the other factors of developer experience, AI time savings is not going to make up for bad meeting culture and lots of interruptions and, you know, developers who are constantly being pulled out of their work, unplanned work, interruptions, outages, those kinds of things. AI will not make up for that. We can use AI to help solve that problem, but AI in and of itself is not going to make up for it. Then when we look kind of in the bottom half, build and test wait time, toil and dev environment, we put all that together, we realize that just the time savings from coding task speed-up isn't going to get us very far. But what will get us far is when we can take AI and point it at those problems. Can we use AI to help reduce meeting frequency? Can we use AI to improve CI wait time? Can we use AI to reduce dev environment toil? That is what winning organizations are doing right now. They are putting DevEx at the center of their universe and seeing AI as a tool to fix systems level problems.
They're doing it also on an organizational level. If you want organizational outcomes like revenue, P&L, time to market, you have to think about AI as an organizational problem, not as an individual problem that your developer needs to solve at their desk. It has to apply to workflows that span entire value streams. Back to that MIT study, when we looked at the organizational barriers to AI adoption, they weren't technical. This wasn't about the models necessarily. It wasn't even about the tools that wrap the models. It was about things like change management or lack of executive sponsorship, when you have an executive team saying go with AI, but they themselves have never cracked their laptop open and fired up Windsurf or Claude Code or Codex. Poor user experience, just very unclear expectations about AI. Those are the things that get in the way.
If this sounds familiar to you and perhaps your organization could do a better job, there's two things that I want to point you to. The first one is the DORA AI Capabilities Model. These are models that kind of communicate and help you get ready for AI. So think about this as an AI readiness model or an AI capabilities model. This has a crazy amount of data from organizations that DORA studies. They do a lot more than just the four key DORA metrics, finding correlations between practices that organizations have and good outcomes with AI. So, if you use AI and have a good, clear and communicated AI stance, you are going to do better organizationally than a company that does not have one. You can find this at dora.dev. It's the DORA AI Capabilities Model. There was just a new paper that came out last month. Nathan is here, who leads DORA over at Google Cloud. If you want to talk to him about this, he's probably the guy.
The other one is the Thoughtworks forest framework. This is similar to the AI capabilities model, kind of a different flavor on it. If you go to thoughtworks.com and look in their white papers, you can read through this. But these are both really solid, well-researched, industry-backed AI readiness models to help convince your leadership team if you need that, or just help you do an internal audit of: are we doing the right things to make ourselves ready to, you know, reap the benefits of all this experimentation?
The last thing is that organizations who are doing really well with AI right now are experimenting by solving real customer problems. Again, space exploration and going to Mars is great, but that is not sustainable for your whole entire organization to be experimenting with going to Mars. It just costs too much money. It distracts too much from the core business problem. It does not serve your customers. So, keep experimentation going. Other experimentation can be really laser-focused on real customer problems that you have. And that is how you're going to see the organizational results.
Somewhere something incredible is waiting to be known. There is so much possibility of how we can build, what we can build, who we can build it for right now with AI, and agents are just accelerating this. They are expanding our universe. We are definitely in an age of exploration. The thing I want to urge all of you to take with you into the rest of the sessions today is to find that balance between a sense of wonder and a sense of awe and aiming for Mars and aiming for your moon colony, but also understanding that we need to solve the problems here on Earth and we have to live in this reality. So please stay grounded, stay skeptical, stay human, most of all stay pragmatic. Thank you all.
Article published
