How Uber Is Leading Its Engineering Organization Through an Agentic Shift
The Pragmatic EngineerAt The Pragmatic Summit, Anshu Chada, who leads Uber's developer platform organization, and Ty Smith, a principal engineer who has been one of the leading voices on Uber's AI strategy, described how Uber moved from AI-assisted coding toward autonomous, agent-driven development. The talk had three parts. Chada explained why Uber made the push and what return it has seen. Smith walked through the internal platforms Uber built or integrated. Chada closed with the problems that remain unsolved: organizational adoption, measurement, and cost. Their position was that agents pay off most clearly when they take over engineering toil. They also acknowledged that proving business impact and controlling spend are still open problems.
From "AI-powered" to agentic, with humans still in charge
Chada began by noting that AI is not new at Uber. The fares and matching platforms have used AI methods for years, and Chada called this one of the things that sets Uber apart from competitors. What is newer is using AI across the engineering lifecycle and, more broadly, in everything Uber employees do. According to Chada, CEO Dara Khosrowshahi has named AI one of the company's six strategic shifts, moving Uber from a human or early-AI-powered company to a generative-AI-powered one. Chada joked that "agentic-powered" is now the more fashionable term but said the idea holds. Chada cited Khosrowshahi's framing that AI lets people become "superhumans" in their productivity.
Chada stressed that the goal is to augment people, not to automate them away. On the engineering side, Uber wants engineers to spend their time on creative work instead of toil. In Chada's account, handing the boring work to AI (upgrades, migrations, bug fixes) produced two results. Engineers were more satisfied, and product teams shipped features at speeds and in ways Chada said the company hadn't thought possible.
From pair programming to "peer programming"
Chada described a change in how developers work with AI. In the "good old days" of 2022–2023, GitHub Copilot offered synchronous tab completion and an IDE chat window. Uber measured roughly a 10–15% increase in overall developer velocity from this. Chada called that phenomenal, but said it did not by itself drive the change of the past year.
That change is "peer programming": developers hand off work that runs asynchronously, and the models are accurate enough that the human mostly redirects and course-corrects. Chada described the end state as developers acting as their own tech leads, directing agents built on various models and pulling them back in when direction is needed. Chada said this doesn't work for every task. It fits toil well: dead code cleanup, writing docs, library migrations. These are basic but essential for a healthy codebase, and they don't grow the business on their own. Moving them to agents, Chada argued, effectively does help grow the business.
Chada also pointed to the rapid growth in how long models can work on a task on their own, showing a logarithmic chart that went from under a second to agents running for hours. Chada said this is why "vibe coding," a joke a couple of years ago, has become a serious idea. Chada cited Ramp, which had its own talk at the summit, and Anthropic's account of shipping a product quickly with multiple agents as examples of where the industry is going.
Toil dominates agent usage, and that reinforces itself
When Uber opened its agentic workflows to developers, Chada said, about 70% of the tasks developers submitted were toil. Chada gave a reason: accuracy was much higher on these tasks than on ambiguous work. A library upgrade or migration has a clear start and end state. A new feature or an experience that needs experimentation does not. Higher accuracy led developers to submit more toil work, which Chada described as a virtuous cycle. That success is why AI-driven toil reduction became one of the developer platform organization's top AI priorities.
The platform stack: building on existing infrastructure
Smith took over to describe the technology, saying Uber sees itself as "building on the shoulders of giants." The base layer is Michelangelo, Uber's long-standing ML platform. It provides a model gateway for proxying to frontier model providers such as OpenAI and Anthropic or hosting internal models, along with the training and inference capabilities expected of an ML platform. Smith said it has leaned more into agentic use over the last couple of years.
Above that sits Uber's existing context: source code, engineering documentation, Jira tickets, Slack. Smith described these as the "organizational memory" an effective agent needs access to. The next layer is industry agent clients. Uber tries to give engineers the latest and best tools and encourages experimentation. Uber builds specialized agents on top of these clients, including the background agent platform and test generation system described later. At the top is what the industry calls engineering enablement: measurement, cost control, and education.
MCP gateway, agent builders, and AIFX CLI
Smith said Uber moved quickly on MCP when it became popular last year. A cross-company tiger team designed a strategy and built a central MCP gateway. The gateway proxies both external MCPs and internal ones exposed from Uber's service infrastructure, and it presents them to engineers consistently, handling authorization, telemetry, logging, and similar concerns. Uber also provides a registry and a sandbox where developers can try MCPs, confirm they behave as expected, and discover new ones.
Through Michelangelo, Uber also supports building agents with SDKs and no-code tools, including agent builders, visualization, telemetry, and tracing. The goal is that agents built anywhere in the company, with access to internal services and data sets, can be discovered through a registry and reused by engineers and non-engineers. Smith said these agents can be deployed consistently across Uber's remote dev environment (devpods), local laptops, the Minion background agent platform, and production.
With engineers using several agent clients (Smith named Claude Code, Codex, and Cursor), Uber saw a need to centralize how those clients are provisioned, configured, and updated. The result is AIFX CLI. It installs and discovers MCPs from the registry and configures them inside agent clients. It applies standard configuration so newcomers get effective prompts and settings from the start, and it connects to the background task infrastructure. Smith described it as the front door developers use to access the agentic infrastructure.
How the developer workflow is changing
Smith compared workflows. Traditionally, developers spent some time planning, most of it writing code, and a little reviewing, inside an edit-build-run loop in an IDE. The first agentic workflow kept the developer closely in the loop: prompting Cursor or Claude Code, approving commands, steering toward an outcome. What Smith sees emerging at Uber and elsewhere is fully autonomous background agents, several running at once. Smith described the dynamic this way: while waiting on one agent, a developer thinks about getting coffee or browsing Reddit, then decides to start another agent instead. Smith said this is where Uber and much of the industry are heading, and that it creates new problems.
Minion: Uber's in-house background agent platform
The first problem was where background agents run. Smith noted that vendors such as Cursor, Claude Code, and Codex run their background agents on the vendor's infrastructure. That may make sense long term, but Uber wanted to bootstrap on its own infrastructure so it could move fast. The result is Minion, which Smith called Uber's formal background agent platform.
Minion is built on state-of-the-art agent CLIs and SDKs and runs on Uber's CI platform. Its environments have Uber's monorepos already checked out, handle network access to the rest of the infrastructure, and connect to MCP servers through AIFX. Developers can reach it from a web interface, Slack, GitHub PRs during code review, and the CLI. APIs let other workflows and services at Uber call it. Smith highlighted its good defaults: users may not write an ideal prompt or have an ideal setup, so Uber provides per-monorepo defaults that make a task more likely to succeed than if the engineer ran it locally without that setup and context.
Smith demoed a real task. A user reported a crash on their Mac when running a command, and Smith passed the error, the platform, and the command to Minion. The interface offers prompt templates with placeholders, a choice of monorepo, running on a branch or as a follow-up to an existing PR or diff, task history, a choice of agent (Claude Code in the demo), and output as a GitHub PR. Smith explained that the output option matters because Uber is several years into a migration from Phabricator to GitHub and has to support both.
A red icon flagged the prompt as low quality and less likely to succeed. Minion includes a prompt improver that analyzes prompts and suggests changes the user can accept. Once a task starts, Minion notifies the user on Slack with tracking links. In the demo, a second Slack message arrived seven minutes later saying the task was done, with links to the PR and artifacts. The completion view shows agent logs, which can be searched if a task fails, along with options to retry or send follow-ups. Here the task succeeded. The PR was co-authored by the Minion bot and the person who started it, linked to Jira, included a test plan describing how the change was verified, and recorded which agent produced it. Smith acknowledged the fix was simple. The point was that the developer pasted in a user's problem and got a PR, avoiding the usual context switching.
Code review becomes the bottleneck: Code Inbox
Smith said developers now spend more time planning and reviewing because far more code is being generated. Review is probably not developers' favorite work, and more of it risks slowing teams down or letting bugs through. The first tool Smith described for this is Code Inbox, which targets context switching across many agent-generated PRs and agents that need attention.
Code Inbox is a single inbox for PRs a developer needs to review. It is built to cut noise, surfacing only items that are actionable and relevant to that person, when they need attention, and not while they are waiting on someone else. Smith said a lot of work went into smart assignment. It picks reviewers based on ownership and compliance, how the person has worked in the past, their time zone, and their calendar. It enforces SLOs, with visibility into how long a review has been assigned, reassignment, and automatic escalation. Slack notifications are batched and respect focus time and holidays, and they can plug into a team's existing Slack review queues.
Code Inbox also estimates the risk of each change. Smith said Uber plans to keep investing here. A small test change is very different from a change to a key service, so the tool analyzes surface area, blast radius, and service type, and shows that to reviewers so they can scrutinize more closely or bring in another person.
uReview: grading AI review comments
The second review product, uReview, helps with the review itself. Smith acknowledged external tools such as CodeRabbit and Graphite and said Uber has tried several and will keep using some. Uber's internal context, and complications like the Phabricator-to-GitHub migration, made it worthwhile to own the layer through which review comments reach developers.
uReview preprocesses code and then runs a set of plugins: general defect-finding bots, checks drawing on best practices or MCPs, and other organizational sources. An API lets external review bots plug in, so their comments appear with the rest and duplicates and noise are reduced. Output then goes through a review grader. Smith named a common problem: review bots produce many low-value comments because they would rather give the developer something to do even when it isn't needed. Uber wants only high-confidence comments that matter, not nits. The pipeline then removes duplicates across systems and categorizes the remaining comments. Smith said each stage was evaluated separately and uses whichever model performs best for that job, and that the system has been developed over most of the past year.
As it matured, Smith reported, uReview produced higher-quality comments at a higher rate. As more best practices and rules were added, the volume of comments grew while the rate at which developers addressed them stayed high. Smith treats that address rate as the key signal that comments are useful and not just annoying. Smith also showed a feedback UI built for Phabricator, a reminder that Uber has to maintain two review interfaces for now.
Autocover and the test validator
Smith said review is incomplete without verification, since more incoming code raises the risk of mistakes getting through. Uber built Autocover, a unit test generation system Smith said has been discussed publicly before. Smith acknowledged that tools like Claude Code can already generate tests. Smith's argument was that a dedicated custom agent, built on Uber's internal LangFX SDK (which is built on LangChain), produces much better tests. Smith reported about 5,000 generated tests merged per month across the company, and roughly 3x the quality of tests from a typical general-purpose agent.
Uber was worried about bad tests, such as change-detector tests, so Autocover includes a critic engine alongside generation. That critic was later split out into a standalone test validator that developers can run on any test, human-written or AI-generated. Smith said the aim is to raise test quality overall and avoid false confidence from higher coverage numbers.
Large-scale change: Auto Migrate and Shepherd
The last category Smith covered was code maintenance, where toil is heaviest. Smith noted CEOs telling the press, boards, and investors that some percentage of their code is now AI-generated, and said Uber looked into how mature companies such as Google and Meta achieve this. Smith's conclusion was that those companies had foundations Uber had not built: infrastructure for rolling out large-scale changes that AI can build on. Last year Uber started a large program, Auto Migrate, to create scalable large-scale change.
Smith described the program's four areas:
- Problem identification: judging a migration's risk and surface area, and deciding how to split the work into PRs that make sense and reduce risk.
- Code transformation: using an agent such as Claude Code, or a deterministic tool such as OpenRewrite, in which Uber has invested heavily.
- Validation: gaining confidence in an automated change without depending only on human review, using CI, unit tests, and sometimes staging or production signals.
- Campaign management: built from scratch. It covers getting perhaps 100 migration PRs to the right people, tracking them, notifying owners, and keeping them up to date.
Campaign management became the platform Shepherd. Migration authors use its web UI to track every PR in a migration. They define a migration in a YAML file, with either a prompt for an agent or a pointer to a script. Shepherd generates the PRs, refreshes them on a set schedule, notifies reviewers, routes them to the right queues, and integrates with Code Inbox.
Smith showed two examples. In the first, Shepherd used OpenRewrite to generate PRs moving Uber's Java services to Java 21. Each PR identified the correct owners, stayed within one code-owner boundary, and contained the small change needed for the upgrade. In the second, Shepherd used Minion as its agent. Tools from Uber's programming systems group had identified performance issues, and Shepherd generated many PRs and diffs to fix them. Each included a standard description of how it was tested and verified and what the reviewer needed to know to review it safely. (Smith briefly named the tool as "Dr. Fix," then corrected that it was a different tool.)
Non-technical challenge 1: a fast-moving market
Chada returned to discuss the challenges that remain unsolved. The first is business-related. Chada compared the AI market to an Olympic race whose leaders change often: which models are strongest for which tasks, and whether to build in-house or buy SaaS, have to be revisited regularly. At Uber's size, investments like Autocover or Auto Migrate commit dozens of people for months, so the company cannot change direction every quarter.
Chada described two mitigations. The first is good abstraction layers. With Minion, for example, Uber can swap the underlying model or technology if something better appears. The second is a mindset: accepting that what Uber builds will probably be replaced by something better from the industry, and not getting attached to it. Chada said the Cursor co-founder had mentioned a test coverage system that might arrive within weeks, said they were excited about it, and said it might make Autocover obsolete. That would be fine, in Chada's view, because the goal is impact for Uber.
Non-technical challenge 2: legacy systems and slow adoption
The second challenge is people. Uber has 10–15 years of infrastructure, some very sophisticated and some, in Chada's words, archaic and understood by very few people. Making those systems reachable by AI is hard. Even setting up MCP endpoints for different parts of the ecosystem has been difficult.
Adoption has been slower than Chada expected. Chada called the tooling "magic" and described a demo session in which four VPs landed code for the first time in years within 24 minutes. Many developers, though, are being asked to give up habits they are used to, such as reading code, writing it from scratch, and working in an IDE, and to take a risk on a very different way of working. Uber tried top-down mandates telling people they must adopt. Chada said this had some effect, as tracking any metric tends to push it up. What worked better was sharing wins. When engineers shared things they had tried that worked, adoption "erupted." Uber now relies on key promoters spreading techniques to their peers, because, in Chada's words, engineers trust other engineers more than directors like Chada.
Non-technical challenge 3: measuring real impact
Chada said Uber has many metrics and can say with confidence that AI is having a positive impact. Developer NPS and overall developer experience are at their highest, self-reported satisfaction and productivity have never been higher, and the volume of AI-landed code and overall velocity are strong. Chada showed a chart with an inflection point when the Minion agentic system launched and models such as Sonnet and Opus became very good. The gap between casual users and power users (those using the tools at least 20 days a week, as Chada put it) has widened sharply.
Chada acknowledged these are activity metrics, not business outcomes. Uber's CFO has asked what the impact is, and Chada said they can't just show diffs; they need to show revenue impact. Chada called this unsolved and said other companies likely face it too. This year Uber plans to instrument its feature infrastructure to measure the time from when a design is created to when an experiment launches in production, and to track whether that pipeline speeds up.
Non-technical challenge 4: cost
Chada said plainly that "the cost of AI is too damn high." Uber's AI costs have risen at least 6x since 2024. Chada said the technology is clearly valuable, but spending has gone from something Chada could fund from their own budget to something that needs CFO approval. Chada didn't blame vendors like Cursor or Anthropic, pointing instead to high GPU and memory costs.
Uber has had to be more careful with tokens and with matching models to tasks. In Minion, for example, a stronger model creates the plan and cheaper but still effective models carry it out. Chada said the goal is for the infrastructure to make these choices so developers don't have to, which reduces friction and cost together. Chada said this requires constant re-evaluation as new tools arrive. This year Uber added JetBrains AI and Warp, each with its own pricing model and its own complications in how developers use it. Chada presented balancing that spend against impact that is still hard to measure as an ongoing problem, not a solved one.
Awesome. Good afternoon, folks. Thank you so much for joining. I really appreciate it. Yes, my name is Anshu. I lead the developer platform organization at Uber. And Ty is a principal engineer. He's been one of the leading engineering voices that's led our agentic shift, but also our overall AI strategy over the past couple years.
All right, let me try to get out of the way. We have a pretty packed agenda, but I want to try to get to the end as fast as possible because I'm curious about your folks' perspective on what I'm going to talk about.
We're going to walk through what has motivated our push into agentic AI and some of the key ROI that we've manifested. And there is key ROI. I'm really excited about the impact that we realized for Uber. And then I'm going to hand off to Ty and he's going to get into the specifics, the actual technologies that we built or integrated that has resulted in that impact. And then I'm going to end the talk talking about some of the non-technical challenges that we've been dealing with. Organizational and cultural, basically people challenges, measurement, and of course cost.
Okay, so AI is not new to Uber. Our fares platform, our matching platform has been using AI methodology for years and years. It's one of the things that sets us apart from the competitors. But over the past few years, using AI as part of the engineering productivity and engineering life cycle, that's fairly new.
And it's AI's integration into not just engineering, but all aspects of what Uber employees do, that has become a bigger and bigger part of what we want to focus on. So much so that Dara has stated that AI is one of our six strategic shifts. We move from a human / early AI-powered company into a generative AI-powered company. Now, I will say that that term is really passé these days. Nowadays, it's more fashionable to be an agentic-powered company, but the concept still holds where, from the metrics, from the data that we've gathered, Dara made this quote, which is, you know, AI is enabling people to become superhumans in terms of their productivity and the impact that we can realize for our end users.
So, from that standpoint, we want to enable all tasks that people do at Uber to be supported by generative AI to augment human productivity. And that last part is really important because what we're not pushing for is AI automate all humans in the company, right? And especially on the engineering side, what we found is we want to focus on enabling our engineers to focus on creative work rather than toil.
I'm going to get into some of the metrics and the impact that we realized from that, but as we've unlocked AI, what we found is when we push some of the boring stuff to it, upgrades, migrations, bug fixes, not only does it result in much higher satisfaction for our engineers, they're able to push our product and create features for end users in ways that we didn't even think was possible and at velocities that have just been incredible. So, this has been the place that we've really been doubling down on.
Now, one of the reasons we've been able to push on this is the capabilities of the technology and the industry. And the first part of it is we're going from pair programming to peer programming. So, if you think about back in the good old days of 2022 and 2023, when GitHub Copilot first came out, it was a pretty novel way of augmenting development. You had a system where you could do synchronous tab completion and an IDE chat window that would help developers move faster. We saw it ourselves in our metrics. We saw maybe a 10 to 15% bump in overall dev velocity.
This is pretty phenomenal, but this by itself didn't push us in the direction that we've seen over the past, let's say, year, where the paradigm has shifted to peer programming, where you could hand off workloads that are running asynchronously, and the models that we use are so good, they're so accurate, that all you need to do with the AI agent is to redirect in certain ways, maybe give it some course correction.
And that's all culminated in this model where we imagine developers acting as their own tech leads. Right? Developers are directing AI agents using a variety of the different models and capabilities that are available to be able to execute asynchronously and come back for direction.
Now, this doesn't work for every single task, but again, when we think about some of the toil work that developers need to do, dead code cleanup, writing docs, library migrations, these all seem basic, these are basic operations, but they're absolutely essential for maintaining a healthy codebase. But they don't by themselves help to grow the business. So, pushing these workloads to AI agents in effect helps to grow the business.
Another thing that's really helped is the growth of the capabilities of the system. So, on the right we see a logarithmic diagram that shows how long some of the models have been able to execute over the period of time. We go from, you know, less than 1 second to agents that can operate for hours and hours.
And that's helped with this paradigm shift that's happened where, again, when Copilot was first introduced, it was still augmenting traditional development. But as the capabilities have gotten better, we've seen this concept of vibe coding, which was just a joke a couple years ago, becoming much more prominent and much more of a serious concept. And it's resulted in companies like, I have an example for Ramp. I know they have a talk today. But there are other companies that have had similar examples. In fact, even Anthropic talked about how they were able to release Cowork very quickly using a variety of agents. This example is not unique, but it is representative of where the industry has been going so far.
Okay. So, talking about toil, Ty is going to talk about the agentic system that we've deployed out. As soon as we made our agentic workflows available to developers, what we saw is 70% of the workflows that developers are pushing into the system were toil tasks. There's a couple reasons for that. One is the accuracy of these tasks was much higher compared to the more ambiguous workloads. And it makes sense, right? Like the start and end state of some of these tasks, if you think about a library upgrade or a migration, is much more straightforward versus, say, building a brand new feature or an experience out that requires an experiment.
Because the accuracy was higher, developers were more likely to push more workloads into the agentic system that were toil oriented. And it became a virtuous cycle that we saw. And based on that, based on the success, we pushed for making this one of my org's, developer platform's, top priorities in terms of AI augmentation.
Okay, I'm going to hand off to Ty now to get into specifics.
I'll start by saying that we are not building this in isolation at Uber. Uber's had a long history of building AI solutions and having a lot of engineers across the organization building infrastructure. And so we really see ourselves as building on the shoulders of giants. We have our historic Michelangelo platform, which has had some public content in the past, that provides things like a model gateway so that we can proxy and talk to the main frontier models or host internal models, traditional inference and training platforms, and all the other things you would expect in an ML platform, that within the last couple years has really started to lean more into the agentic side and the APIs that we're using to talk to OpenAI and Anthropic and those folks.
On top of that, we have a lot of the traditional infrastructure and context at Uber that we would want to take advantage of. Things like having access to our source code, our engineering documentation, Jira tickets, Slack information. These are all things that, to have an effective agent with organizational memory, it needs to start to get access to. One of the key ones that I'll dig a little more into in the later slides is our deployment of MCPs throughout Uber.
On top of that, we see a lot of industry agents. We really take a perspective of trying to enable the latest and greatest for our engineers, allowing them to experiment, allowing them to have a learning culture and use the best of class. So that means that there's a lot of clients that are coming in that folks are using, and we use a lot of those to build specialized agents. This could be our background agent platform that we're going to be talking about, our test generation platform, or many other kinds of internal ones. And then at the top of that we have, you know, the engineering enablement phrase that's been going around the industry. It's measurements and cost control and education and everything else that you would expect.
So let's dig in a little bit to how we think about MCPs. This became a very popular piece of technology in the industry last year, and we moved very quickly to make sure that this was deployed and secure for our engineers so that we could make them as productive as possible. So, we ended up putting a tiger team together from across the company. They came together and designed a strategy and built a central MCP gateway.
This allows us to proxy external and internal MCPs from our service infrastructure and expose those in a consistent way to engineers, handling things like authorization, telemetry, logging, everything else that you might expect. We also provide a registry and a sandbox so that developers can come in, they can play with these MCPs, they can make sure that it's going to do what they're expecting, and that they can discover new ones.
Continuing on with that, we also have at Uber, through our Michelangelo platform, the ability to build agents with both SDKs and with no-code solutions. Building agent builders, the ability to visualize, have telemetry, do tracing. That way, as folks around the company are building some of these solutions that have access to the internal services, the data sets, we can reuse these in other systems. This can be discoverable, and we can provide a registry that can then be found by other engineers or non-engineers alike to deploy these. And they're deployed consistently in a lot of our environments. This can be from our dev pod infrastructure, which is our remote dev environment, local laptops, through our background agent, which is called Minions, we're going to introduce that here in a minute, or deploying these in production.
So, we talked about the agents and the registry there, the MCPs and the registry there, all the different agent clients that our engineers might be using, be it Claude Code or Codex or Cursor. One thing that we recognized we needed to platformize pretty quickly was a central ability to provision and configure and update the agent clients themselves, the ability to install and discover MCPs from the registry, configure those inside of the agent clients, deploy standard configuration management so that people who are just new to the space are having more effective prompts and configurations right away, and management and connection into our background task infrastructure. So we built this tool called AIFX CLI, and it is kind of the forefront of what developers are using to access our agentic infrastructure.
So, let's take a minute before I jump into our specific product and think about the traditional developer workflow. If you looked at how people were spending time, a little bit in planning, probably a lot in code authorship historically, and then a small amount in review. And then typically they'd be in this edit, build, run loop of editing their code, building it, doing the verification, using some standard IDEs.
Now of course this has been changing significantly with the agentic world. And so if we look at what the first agent workflow looked like, it might look something like this. You have a developer who's in the middle of using Cursor or Claude Code, they're giving a prompt, it's asking for the ability to proceed and approve commands, and they're very interactive in the loop, trying to drive it to an outcome that they want.
But what we're seeing emerge now in the industry and at Uber is both background agents that are running fully autonomously, as well as the ability for multiple of these to be run at once. Right? This gets into this place where as an engineer you're giving a prompt, you're waiting for some time while it's running. You're thinking, "Oh, what am I going to do? Am I going to go have a coffee or browse Reddit? Might as well kick off another background agent." And so they get into this mode where the new flow looks like running several agents at once, right?
This sounds great. I think us and a lot of the industry are trying to push towards this. But a lot of challenges start to emerge with this different way of working. One of them for us was we wanted these background agents to be running autonomously, and looking at the external vendors that were offering, you know, tools like Cursor and Claude Code and Codex, all of them are running their background agents in other people's infrastructure. And while we can get there, while that may make sense long term, having the ability to bootstrap on our own infrastructure was really important to us and allowed us to move really quickly. And so we built a product called Minion. Minion is our formal background agent platform.
It's built on top of state-of-the-art agents, CLIs, and SDKs. This leverages all of Uber's existing infrastructure. It runs on our CI platform. This has our monorepos checked out, ready to work in quickly, handles all of the network access into the rest of the infra, and allows the connection to all of those MCP servers that we talked about earlier through AIFX.
It's integrated for the developer in a bunch of different workflows and panes of glass. There's the web interface we're looking at here, which is one of the main interaction paradigms, but it's also available through Slack, through GitHub PRs in the code review process, through the CLI that we saw earlier, and we have APIs exposed so that it can start to be connected to by other workflows and other services throughout the rest of Uber.
And one other powerful thing is this offers good defaults. So when people are coming here and kicking off these background jobs, you know, they're giving a prompt, they're expecting a PR out of this. They may not be giving the ideal prompt or have the ideal setup for it. And we can provide great defaults for each of our monorepos, make sure that this is more likely to be a successful task for the engineer authoring it than if they just did this locally and didn't have a lot of the CLAUDE.md setup or the other context that we may want to provide.
So let's walk through a demo real quick of what using Minions is like with an example. So we have this web interface, and this is one that I actually ran. We had a user report an error. They said, "Hey, this is crashing on my machine when I run this command. Here's what the error is." And I threw that into Minion. I said, "Hey, you know, we're having this issue. The user's on a Mac. Here's the error they're seeing. Here's the command that was run."
And so you can see a few things here that are cool. One, we have these existing templates that users can choose from that are well-written prompts that have placeholders they can fill in. We have the ability to choose and run in our different monorepos. I can run on a branch, or we can switch it to a follow-up task of existing PRs or diffs. We have all the task history here.
Then we have some cool things here. I can select the agent. So in this case I'm going to run it in Claude Code. Put it out as a GitHub PR. We've been in a long multi-year migration from Phabricator to GitHub, so having this dual mode is important for our internal engineers.
And one interesting thing you'll see is this red icon here. What this indicates is that this wasn't a great quality prompt and it would have less chance of success. So one tool that we built into this was a prompt improver: the ability to analyze the prompt and make suggestions that the user can accept on how to have a higher chance of success.
Now, once that kicks off, this is running. You know, background agents can take a little bit of time. We ping the users on Slack, give them links so that they can go ahead and track this. And a few minutes later, in this case it was 7 minutes later, the Slack notification pings them again. It says, "Hey, the Minion task is done. You can go look at the PR here. You can go look at the artifacts."
So, let's not go to the PR quite yet. Let's jump back into the task completion. I have a view here now where I can see what ran. I can investigate the agent logs if I need to. If this failed, I can retry or have follow-up tasks. Let's say it failed, then I can search through the logs here and start to try to understand what the agent was doing and maybe give a follow-up. But, in this case, it was successful. We got a PR out of that immediately. It was a very straightforward one.
Our Minion bot co-authors this with the person that kicked it off. Here we have a linked Jira. We have the test plan of how it verified. You can see it was authored here by which agent Minion was running, Claude. And it was a very straightforward fix that we got. So, this was a very simple workflow, but it was much easier for the developer to just dump in a prompt, "Hey, here's a problem the user is having," and get a PR of that, as opposed to all of the context switching that they would need to traditionally do.
Right now, the workflow for the developers has changed and is changing further. They're spending more and more time in planning and code review because there's so much more code being generated that they're being forced to do it. This probably isn't the favorite type of work that developers love doing, code review. And there's a lot of challenges with that. If people are doing code review and it's taking more time, they're maybe slowing down. They maybe let more bugs in because they're missing it in review because there's much more. So, let's jump into a few of the investments we made to try to improve that.
So, one of the big problems is context switching amongst all the background agents. This could be on PRs that are coming out or the agent itself needing attention. So, we built a tool called Code Inbox, which was designed to try to help with this situation. It's a unified inbox for PRs that a developer needs to review. And what's interesting about this is it's designed to try to remove noise. So, only bringing out the actionable ones that are directly relevant for a user when it needs attention, not when it's, you know, sitting there waiting for someone else.
And we put a lot of work into smart assignments with Code Inbox so that we try to find the most relevant person to review the code, both from an ownership and compliance perspective, but also the history of how that person was working, their time zone availability, their calendar availability. And we try to find the right person and assign, and then have strict SLOs that we track so they can see how long it's been assigned, help reassign, do automatic reassignment or escalation if necessary.
And then this does a smart job of the Slack notifications to devs. So, it's doing things like batching notifications so they don't see a bunch of noise, or accounting for their focus time so it's not bothering them in the middle of it, or, you know, their holiday time if they're out. It'll also handle integrating this into teams' existing processes. So, if teams have existing Slack usage with code review queues, we can plug directly into that and inject the reviews at that level as well.
Some of the other cool stuff that we built into this one was we tried to understand the risk of the change, and we're going to continue to invest in that. There's a much different risk profile to a small change in a test versus a change that's in one of our key services. And so, we try to highlight that here by analyzing the surface area, the blast, how much that's going to affect, what type of service that's hitting, and then make those estimates so that we can raise that to the developer, so they might put more scrutiny on the review or bring in another person or, you know, whatever decision they might want to make for a riskier change.
So, in the code review space, I want to move on to a second product. We talked about the notifications and bringing context awareness, but this is our product, it's called U Review, and this one is designed more at the review help itself. There's a bunch of external products right now. We've all seen them in the market, everything from CodeRabbit to Graphite. All of those are trying to solve this problem, and we've played with a bunch and we'll continue to use external ones as well, but what we found is we had a lot of internal context. We had a lot of complexity, like the migration between Phabricator and GitHub, that made it make sense for us to have a platform where we were controlling the surface area for the comments coming through.
And so how this works is we have a preprocessor for the code, and at that point we have a set of plugins that are going to run. There can be general defect bots that are analyzing it. It can be pulling from best practices or MCPs or other types of information around the organization. And we also have an API so that we can plug in external bots. So, if we were using, you know, one of those external code review tools, we can just plug it into the API here and have it surfaced with the rest of the comments that are coming in to the developer, to help minimize duplicate comments or extra noise.
That then runs through a review grader. This has been one of the common problems: we've seen a lot of low-value comments surface from those, because they'd rather give something to the developer to do even if it isn't maybe necessary. And we really only want to put the high-confidence changes that the developer really needs to focus on, not little nits. And so, this continues through the flow. It looks for duplicates from these different systems and finally categorizes these. Now, for each one of these layers we've done evaluations and have different models running based on the performance that each model has on the type of behavior.
This has been something we've been working on for most of the last year. So, we saw some growth and some progress in this system as it matured. One, we saw that we were able to get higher quality comments at a higher rate. As we invested and integrated additional best practices and other rules, we also saw the rate of the comments and the best practices increase while maintaining a high rate of comments being addressed. This is the specific piece of feedback that we're looking at to make sure that this isn't noise, that developers are actually fixing these and it's not just annoying them. And then here's a screenshot in Phabricator, not in GitHub, where, as I mentioned, we have to have the dual kind of UI because of the two systems at the moment. And so this was a custom one we built to try to have a feedback loop for the developer.
In the code review space, it wouldn't be complete if I didn't talk about the verification, the validation, CI, tests. I think that's the other big part that we are really concerned about, to make sure that those mistakes aren't slipping through code reviews as more code is coming in. So, we built a system called Auto Cover that we've talked a little bit about in the past. Actually, the author for it, I saw him around here somewhere. This was a system that we designed to generate unit tests. Now, you might say, "Well, you can just do that with Claude Code or many other products." And you can, but what we found is by really focusing on this project, building a custom agent on top of our internal LangFX SDK built on LangChain, we were able to get a much higher quality type of unit test output. And so at this point we're seeing about 5,000 tests generated and merged per month around the company from this, and almost a 3x rate of quality versus something that would be generated from your typical generic agent.
Now, as we were doing this, we were quite concerned with, you know, bad quality tests, change detector tests, things like that coming in. And so we built into this a critic engine. So it has both the generation and the critic engine. And we separated that out into an independent test validator that developers can now use independently, whether it's a human-generated test or an AI-generated test, which is great to help uplevel the test quality in general and ignore any false confidence that we might get from having the higher coverage.
So I'm going to talk about one more category before I hand it back to Anshu. We've talked about, you know, authoring code initially and the code review process, but code maintenance is a big area. And, you know, at the beginning Anshu was talking about the toil work. This is where a lot of folks kind of consider toil the heaviest.
So as we were looking at how to build this out, we looked at the space, we looked at the messages coming out from other companies, you know, where their CEOs are going up in front of the news or to their boards or their investors and saying X percent of our code is generated by AI now. And we were looking at that and saying, well, how is some of this done? And some of the companies were very mature companies, like Google or Meta. And we'd had a lot of discussions and seen that they had fundamentals that Uber hadn't invested in before, which was the ability to kind of scale out large-scale changes so that the AI can then build on top of that. And so we got together last year and we decided we needed to run a big program to create a scalable version of how we handle large-scale change. We call this Auto Migrate.
We broke this program up into kind of four key areas. We have the problem identification area, where someone would be looking at a migration or an upgrade and deciding what the risk of the migration is, or what the surface area of the changes is, or how to cut up the PRs so that they make sense and reduce risk. You have the code transformer piece, which could be an agent, you know, we could be using Claude Code or any others, but it could be something deterministic like OpenRewrite, which we've made a lot of investments with.
Then it gets into the validation phase, where we need to understand how we get confidence that that automated change is going to be successful, and that it's not just relying on human review. So, this might be CI or unit tests or sometimes even, you know, staging or production signal, but this is a key area as we think about it. And then finally, the area of campaign management was something specifically that we needed to build from scratch. You know, if you have 100 PRs that need to go out to developers for a migration, how do you get those all into the right spot? How do you track those? How do you make sure those folks are notified? How do you refresh this? And so, this became the key of the platform that we called Shepherd.
Here's introducing the experience with Shepherd. At the surface, it's a web UI where developers can go, migration authors specifically, and they can track all of the PRs that are associated with a migration. It allows them to define those simply through a YAML file, where they can either give a prompt if it's an agent or they can point it to the script that's going to be handling it. And then Shepherd is going to take care of generating those PRs, refreshing them on whatever cadence you defined, keeping those fresh for the developers, notifying the people that need to review it, getting it in the right queues, integrating with Code Inbox, the last product I showed.
So, let's walk through two quick demos of this. One, here's a PR using one of the deterministic transformers. This used OpenRewrite. We had Shepherd generate all of the PRs to move our Java services to Java 21. Here we had this PR generated that correctly found the owners, created a limited PR in just the space of the code owners that upgrades it to Java 21, and we can see that generated here in a small change needed for that upgrade.
Here's one more where it's similar, but in this case it's using the Minions platform. It's integrated with it as an agent. So, we have separate tools in our programming systems group to do analysis and find performance issues. A lot of these are generating really good data. One of these we called Dr. Fix. Actually, it's a different one, sorry. It's not about Dr. Fix. But it identified these performance issues, and it was able to generate a lot of PRs, or diffs in this case, to account for those, run those through Shepherd, and have a standard thing that accounts for how it was tested, how it was verified, and what the developer needs to know to review it safely.
And with that, we've walked through kind of the major deep dive. I want to hand it back to Anshu to talk about the last couple of topics.
All right. So Ty talked through a lot of the engineering investments that we made to make, you know, this agentic shift. I'm going to talk through a few non-technical challenges that we're still dealing with.
First up is on the people side and the business side. So, on the business side, I have a diagram here that's very topical since the Olympics are on right now. The leaders when it comes to AI tech are changing pretty frequently. You know, the models that are the most powerful for certain tasks, whether you should build something in-house versus use a SaaS provider, these decisions need to be revisited on a pretty regular basis. Unfortunately, in a large organization like ours, some of the investments that Ty talked about, whether it's Auto Cover or Auto Migrate, these are not trivial decisions to make. We need to commit dozens of people on projects that might be running for months. So, we can't just change our mind after a quarter.
There's two things that we've done to mitigate this. One is seemingly pretty basic: making sure that we have the right abstraction layers in place. Ty showed the Minions infrastructure. Under the covers, if we need to swap out the model or we need to swap out the technology that we're using, we're now able to do so. If a better technology comes around that can solve some of the underlying pieces more effectively, we can do so. But the second part is just having this belief that the tech we're building will likely be replaced with something better in the industry. And so it's really important for us to not be married to the tech that we're building, and to be okay if something comes along. Like, the co-founder of Cursor talked about the test coverage system that might be coming in a couple weeks. I'm really excited about that. It might make our Auto Cover infrastructure obsolete, and that's okay, because at the end of the day we need to deliver impact for Uber.
The second part is another people problem. So, a lot of the challenges that Ty alluded to deal with, you know, historic infrastructure that's been built out over the last 10 to 15 years at Uber. We have some really sophisticated code that we built out, and then we have some really, I would say, archaic code that, you know, very few people know about. Getting that technology integrated into places where AI can reach it is challenging. Even just getting MCP endpoints set up to reach different parts of our ecosystem has been a challenge.
Similarly, the tech that Ty talked about, like I've seen in action, it's magic. I ran a
demo session with some of my VPs, and in 24 minutes I had four VPs land code for the first time in years. It was a pretty amazing experience. They were pretty satisfied by it, too.
But our adoption for this technology has been relatively slow. It's been slower than I've expected. And part of it is because we're trying to have developers do something that they're so not used to. They're used to looking at code and generating from scratch, operating in their IDE, and we're telling them to take a risk by operating in a very different way.
In both of these cases, you know, we've tried different tactics to get around this people issue. We've tried a top-down approach, you know, directives from leaders to say, "You must do X, Y, and Z. You must adopt." It's had some impact, as you folks know. You track a metric, it's going to go up, or it's going to improve.
The more successful technique that we've applied is actually just sharing wins. So, as we share examples between different engineers, cool things that they've tried that have resulted in wins, adoption of that technology has erupted. So, that's been the tactic that we're pushing on now: key promoters pushing techniques to their peers, because those promoters are typically engineers, and engineers trust other engineers as opposed to directors like me.
Okay. I'm going to touch on measurements now. So, we have tons and tons and tons of metrics. I can say with confidence that objectively AI is having positive impact. Our net promoter score, our overall developer experience at Uber, has never been higher. The self-reported net satisfaction developers have and their productivity has never been higher.
The amount of code that we're landing through AI is amazing. The overall engineering velocity is fantastic. And you can see the graph over here. We see the inflection point where we introduced the Minions agentic system, along with when the models became really, really good, like Sonnet and Opus being introduced. The delta between developers that are using it very casually versus the power users that are using it at least 20 days a week has only exploded. The deviation has only gone up.
So I'm really pleased about this. Now, the issue is that these are activity metrics, right? These are not necessarily business outcomes. And when we start talking about the costs of this technology, you know, our CFO has asked me what is the impact of this? Right? You know, I can't point him to diffs. I need to show him what's the impact on revenue.
I'm sure you folks are dealing with the same problem. This is not necessarily a solved problem for us. One of the tactics that we're taking this year is to instrument our overall feature infrastructure so that we can time from when, you know, a design is first created to when an experiment is launched in production. And then seeing how we're able to speed that pipeline up.
And then speaking of costs, the cost of AI is too damn high. You know, since 2024, our costs have gone up at least 6x. Now, we'll say that this technology is amazing. There's, again, no question that it's had positive impact. But it's gone from something that I can self-fund using my own budget to something that I need to ask permission for from, you know, the CFO.
What that's necessitated, especially where we went from a model where, you know, it's not necessarily Cursor's or Anthropic's fault that it's going up, the GPU costs are high and memory costs are really high. So, we've had to be more responsible about how we use tokens, how we think about what's the right model for the job, and then helping developers select those models.
So again, going back to the example with Minions, we helped developers think about the right model to form the plan for the project, and then lower cost but still pretty effective models to do the execution. We don't necessarily want developers to think about it, but we want to be able to have the infrastructure decide for them so that we reduce the friction for them, but then we also optimize our costs.
But this is something that we continuously have to keep on evaluating and adjusting, especially as new technologies are introduced. So, like this year we introduced JetBrains AI and Warp, which we hadn't introduced in the past, and they all have their own costing model and all have their own complexities with regards to how developers are using them.
Article published
