How Uber Is Leading Its Engineering Organization Through an Agentic Shift

Open on YouTube ↗
Overview

At The Pragmatic Summit, Anshu Chada, who leads Uber's developer platform organization, and Ty Smith, a principal engineer who has been one of the leading voices on Uber's AI strategy, described how Uber moved from AI-assisted coding toward autonomous, agent-driven development. The talk had three parts. Chada explained why Uber made the push and what return it has seen. Smith walked through the internal platforms Uber built or integrated. Chada closed with the problems that remain unsolved: organizational adoption, measurement, and cost. Their position was that agents pay off most clearly when they take over engineering toil. They also acknowledged that proving business impact and controlling spend are still open problems.

21 min read

From "AI-powered" to agentic, with humans still in charge

Chada began by noting that AI is not new at Uber. The fares and matching platforms have used AI methods for years, and Chada called this one of the things that sets Uber apart from competitors. What is newer is using AI across the engineering lifecycle and, more broadly, in everything Uber employees do. According to Chada, CEO Dara Khosrowshahi has named AI one of the company's six strategic shifts, moving Uber from a human or early-AI-powered company to a generative-AI-powered one. Chada joked that "agentic-powered" is now the more fashionable term but said the idea holds. Chada cited Khosrowshahi's framing that AI lets people become "superhumans" in their productivity.

Chada stressed that the goal is to augment people, not to automate them away. On the engineering side, Uber wants engineers to spend their time on creative work instead of toil. In Chada's account, handing the boring work to AI (upgrades, migrations, bug fixes) produced two results. Engineers were more satisfied, and product teams shipped features at speeds and in ways Chada said the company hadn't thought possible.

From pair programming to "peer programming"

Chada described a change in how developers work with AI. In the "good old days" of 2022–2023, GitHub Copilot offered synchronous tab completion and an IDE chat window. Uber measured roughly a 10–15% increase in overall developer velocity from this. Chada called that phenomenal, but said it did not by itself drive the change of the past year.

That change is "peer programming": developers hand off work that runs asynchronously, and the models are accurate enough that the human mostly redirects and course-corrects. Chada described the end state as developers acting as their own tech leads, directing agents built on various models and pulling them back in when direction is needed. Chada said this doesn't work for every task. It fits toil well: dead code cleanup, writing docs, library migrations. These are basic but essential for a healthy codebase, and they don't grow the business on their own. Moving them to agents, Chada argued, effectively does help grow the business.

Chada also pointed to the rapid growth in how long models can work on a task on their own, showing a logarithmic chart that went from under a second to agents running for hours. Chada said this is why "vibe coding," a joke a couple of years ago, has become a serious idea. Chada cited Ramp, which had its own talk at the summit, and Anthropic's account of shipping a product quickly with multiple agents as examples of where the industry is going.

Toil dominates agent usage, and that reinforces itself

When Uber opened its agentic workflows to developers, Chada said, about 70% of the tasks developers submitted were toil. Chada gave a reason: accuracy was much higher on these tasks than on ambiguous work. A library upgrade or migration has a clear start and end state. A new feature or an experience that needs experimentation does not. Higher accuracy led developers to submit more toil work, which Chada described as a virtuous cycle. That success is why AI-driven toil reduction became one of the developer platform organization's top AI priorities.

The platform stack: building on existing infrastructure

Smith took over to describe the technology, saying Uber sees itself as "building on the shoulders of giants." The base layer is Michelangelo, Uber's long-standing ML platform. It provides a model gateway for proxying to frontier model providers such as OpenAI and Anthropic or hosting internal models, along with the training and inference capabilities expected of an ML platform. Smith said it has leaned more into agentic use over the last couple of years.

Above that sits Uber's existing context: source code, engineering documentation, Jira tickets, Slack. Smith described these as the "organizational memory" an effective agent needs access to. The next layer is industry agent clients. Uber tries to give engineers the latest and best tools and encourages experimentation. Uber builds specialized agents on top of these clients, including the background agent platform and test generation system described later. At the top is what the industry calls engineering enablement: measurement, cost control, and education.

MCP gateway, agent builders, and AIFX CLI

Smith said Uber moved quickly on MCP when it became popular last year. A cross-company tiger team designed a strategy and built a central MCP gateway. The gateway proxies both external MCPs and internal ones exposed from Uber's service infrastructure, and it presents them to engineers consistently, handling authorization, telemetry, logging, and similar concerns. Uber also provides a registry and a sandbox where developers can try MCPs, confirm they behave as expected, and discover new ones.

Through Michelangelo, Uber also supports building agents with SDKs and no-code tools, including agent builders, visualization, telemetry, and tracing. The goal is that agents built anywhere in the company, with access to internal services and data sets, can be discovered through a registry and reused by engineers and non-engineers. Smith said these agents can be deployed consistently across Uber's remote dev environment (devpods), local laptops, the Minion background agent platform, and production.

With engineers using several agent clients (Smith named Claude Code, Codex, and Cursor), Uber saw a need to centralize how those clients are provisioned, configured, and updated. The result is AIFX CLI. It installs and discovers MCPs from the registry and configures them inside agent clients. It applies standard configuration so newcomers get effective prompts and settings from the start, and it connects to the background task infrastructure. Smith described it as the front door developers use to access the agentic infrastructure.

How the developer workflow is changing

Smith compared workflows. Traditionally, developers spent some time planning, most of it writing code, and a little reviewing, inside an edit-build-run loop in an IDE. The first agentic workflow kept the developer closely in the loop: prompting Cursor or Claude Code, approving commands, steering toward an outcome. What Smith sees emerging at Uber and elsewhere is fully autonomous background agents, several running at once. Smith described the dynamic this way: while waiting on one agent, a developer thinks about getting coffee or browsing Reddit, then decides to start another agent instead. Smith said this is where Uber and much of the industry are heading, and that it creates new problems.

Minion: Uber's in-house background agent platform

The first problem was where background agents run. Smith noted that vendors such as Cursor, Claude Code, and Codex run their background agents on the vendor's infrastructure. That may make sense long term, but Uber wanted to bootstrap on its own infrastructure so it could move fast. The result is Minion, which Smith called Uber's formal background agent platform.

Minion is built on state-of-the-art agent CLIs and SDKs and runs on Uber's CI platform. Its environments have Uber's monorepos already checked out, handle network access to the rest of the infrastructure, and connect to MCP servers through AIFX. Developers can reach it from a web interface, Slack, GitHub PRs during code review, and the CLI. APIs let other workflows and services at Uber call it. Smith highlighted its good defaults: users may not write an ideal prompt or have an ideal setup, so Uber provides per-monorepo defaults that make a task more likely to succeed than if the engineer ran it locally without that setup and context.

Smith demoed a real task. A user reported a crash on their Mac when running a command, and Smith passed the error, the platform, and the command to Minion. The interface offers prompt templates with placeholders, a choice of monorepo, running on a branch or as a follow-up to an existing PR or diff, task history, a choice of agent (Claude Code in the demo), and output as a GitHub PR. Smith explained that the output option matters because Uber is several years into a migration from Phabricator to GitHub and has to support both.

A red icon flagged the prompt as low quality and less likely to succeed. Minion includes a prompt improver that analyzes prompts and suggests changes the user can accept. Once a task starts, Minion notifies the user on Slack with tracking links. In the demo, a second Slack message arrived seven minutes later saying the task was done, with links to the PR and artifacts. The completion view shows agent logs, which can be searched if a task fails, along with options to retry or send follow-ups. Here the task succeeded. The PR was co-authored by the Minion bot and the person who started it, linked to Jira, included a test plan describing how the change was verified, and recorded which agent produced it. Smith acknowledged the fix was simple. The point was that the developer pasted in a user's problem and got a PR, avoiding the usual context switching.

Code review becomes the bottleneck: Code Inbox

Smith said developers now spend more time planning and reviewing because far more code is being generated. Review is probably not developers' favorite work, and more of it risks slowing teams down or letting bugs through. The first tool Smith described for this is Code Inbox, which targets context switching across many agent-generated PRs and agents that need attention.

Code Inbox is a single inbox for PRs a developer needs to review. It is built to cut noise, surfacing only items that are actionable and relevant to that person, when they need attention, and not while they are waiting on someone else. Smith said a lot of work went into smart assignment. It picks reviewers based on ownership and compliance, how the person has worked in the past, their time zone, and their calendar. It enforces SLOs, with visibility into how long a review has been assigned, reassignment, and automatic escalation. Slack notifications are batched and respect focus time and holidays, and they can plug into a team's existing Slack review queues.

Code Inbox also estimates the risk of each change. Smith said Uber plans to keep investing here. A small test change is very different from a change to a key service, so the tool analyzes surface area, blast radius, and service type, and shows that to reviewers so they can scrutinize more closely or bring in another person.

uReview: grading AI review comments

The second review product, uReview, helps with the review itself. Smith acknowledged external tools such as CodeRabbit and Graphite and said Uber has tried several and will keep using some. Uber's internal context, and complications like the Phabricator-to-GitHub migration, made it worthwhile to own the layer through which review comments reach developers.

uReview preprocesses code and then runs a set of plugins: general defect-finding bots, checks drawing on best practices or MCPs, and other organizational sources. An API lets external review bots plug in, so their comments appear with the rest and duplicates and noise are reduced. Output then goes through a review grader. Smith named a common problem: review bots produce many low-value comments because they would rather give the developer something to do even when it isn't needed. Uber wants only high-confidence comments that matter, not nits. The pipeline then removes duplicates across systems and categorizes the remaining comments. Smith said each stage was evaluated separately and uses whichever model performs best for that job, and that the system has been developed over most of the past year.

As it matured, Smith reported, uReview produced higher-quality comments at a higher rate. As more best practices and rules were added, the volume of comments grew while the rate at which developers addressed them stayed high. Smith treats that address rate as the key signal that comments are useful and not just annoying. Smith also showed a feedback UI built for Phabricator, a reminder that Uber has to maintain two review interfaces for now.

Autocover and the test validator

Smith said review is incomplete without verification, since more incoming code raises the risk of mistakes getting through. Uber built Autocover, a unit test generation system Smith said has been discussed publicly before. Smith acknowledged that tools like Claude Code can already generate tests. Smith's argument was that a dedicated custom agent, built on Uber's internal LangFX SDK (which is built on LangChain), produces much better tests. Smith reported about 5,000 generated tests merged per month across the company, and roughly 3x the quality of tests from a typical general-purpose agent.

Uber was worried about bad tests, such as change-detector tests, so Autocover includes a critic engine alongside generation. That critic was later split out into a standalone test validator that developers can run on any test, human-written or AI-generated. Smith said the aim is to raise test quality overall and avoid false confidence from higher coverage numbers.

Large-scale change: Auto Migrate and Shepherd

The last category Smith covered was code maintenance, where toil is heaviest. Smith noted CEOs telling the press, boards, and investors that some percentage of their code is now AI-generated, and said Uber looked into how mature companies such as Google and Meta achieve this. Smith's conclusion was that those companies had foundations Uber had not built: infrastructure for rolling out large-scale changes that AI can build on. Last year Uber started a large program, Auto Migrate, to create scalable large-scale change.

Smith described the program's four areas:

  • Problem identification: judging a migration's risk and surface area, and deciding how to split the work into PRs that make sense and reduce risk.
  • Code transformation: using an agent such as Claude Code, or a deterministic tool such as OpenRewrite, in which Uber has invested heavily.
  • Validation: gaining confidence in an automated change without depending only on human review, using CI, unit tests, and sometimes staging or production signals.
  • Campaign management: built from scratch. It covers getting perhaps 100 migration PRs to the right people, tracking them, notifying owners, and keeping them up to date.

Campaign management became the platform Shepherd. Migration authors use its web UI to track every PR in a migration. They define a migration in a YAML file, with either a prompt for an agent or a pointer to a script. Shepherd generates the PRs, refreshes them on a set schedule, notifies reviewers, routes them to the right queues, and integrates with Code Inbox.

Smith showed two examples. In the first, Shepherd used OpenRewrite to generate PRs moving Uber's Java services to Java 21. Each PR identified the correct owners, stayed within one code-owner boundary, and contained the small change needed for the upgrade. In the second, Shepherd used Minion as its agent. Tools from Uber's programming systems group had identified performance issues, and Shepherd generated many PRs and diffs to fix them. Each included a standard description of how it was tested and verified and what the reviewer needed to know to review it safely. (Smith briefly named the tool as "Dr. Fix," then corrected that it was a different tool.)

Non-technical challenge 1: a fast-moving market

Chada returned to discuss the challenges that remain unsolved. The first is business-related. Chada compared the AI market to an Olympic race whose leaders change often: which models are strongest for which tasks, and whether to build in-house or buy SaaS, have to be revisited regularly. At Uber's size, investments like Autocover or Auto Migrate commit dozens of people for months, so the company cannot change direction every quarter.

Chada described two mitigations. The first is good abstraction layers. With Minion, for example, Uber can swap the underlying model or technology if something better appears. The second is a mindset: accepting that what Uber builds will probably be replaced by something better from the industry, and not getting attached to it. Chada said the Cursor co-founder had mentioned a test coverage system that might arrive within weeks, said they were excited about it, and said it might make Autocover obsolete. That would be fine, in Chada's view, because the goal is impact for Uber.

Non-technical challenge 2: legacy systems and slow adoption

The second challenge is people. Uber has 10–15 years of infrastructure, some very sophisticated and some, in Chada's words, archaic and understood by very few people. Making those systems reachable by AI is hard. Even setting up MCP endpoints for different parts of the ecosystem has been difficult.

Adoption has been slower than Chada expected. Chada called the tooling "magic" and described a demo session in which four VPs landed code for the first time in years within 24 minutes. Many developers, though, are being asked to give up habits they are used to, such as reading code, writing it from scratch, and working in an IDE, and to take a risk on a very different way of working. Uber tried top-down mandates telling people they must adopt. Chada said this had some effect, as tracking any metric tends to push it up. What worked better was sharing wins. When engineers shared things they had tried that worked, adoption "erupted." Uber now relies on key promoters spreading techniques to their peers, because, in Chada's words, engineers trust other engineers more than directors like Chada.

Non-technical challenge 3: measuring real impact

Chada said Uber has many metrics and can say with confidence that AI is having a positive impact. Developer NPS and overall developer experience are at their highest, self-reported satisfaction and productivity have never been higher, and the volume of AI-landed code and overall velocity are strong. Chada showed a chart with an inflection point when the Minion agentic system launched and models such as Sonnet and Opus became very good. The gap between casual users and power users (those using the tools at least 20 days a week, as Chada put it) has widened sharply.

Chada acknowledged these are activity metrics, not business outcomes. Uber's CFO has asked what the impact is, and Chada said they can't just show diffs; they need to show revenue impact. Chada called this unsolved and said other companies likely face it too. This year Uber plans to instrument its feature infrastructure to measure the time from when a design is created to when an experiment launches in production, and to track whether that pipeline speeds up.

Non-technical challenge 4: cost

Chada said plainly that "the cost of AI is too damn high." Uber's AI costs have risen at least 6x since 2024. Chada said the technology is clearly valuable, but spending has gone from something Chada could fund from their own budget to something that needs CFO approval. Chada didn't blame vendors like Cursor or Anthropic, pointing instead to high GPU and memory costs.

Uber has had to be more careful with tokens and with matching models to tasks. In Minion, for example, a stronger model creates the plan and cheaper but still effective models carry it out. Chada said the goal is for the infrastructure to make these choices so developers don't have to, which reduces friction and cost together. Chada said this requires constant re-evaluation as new tools arrive. This year Uber added JetBrains AI and Warp, each with its own pricing model and its own complications in how developers use it. Chada presented balancing that spend against impact that is still hard to measure as an ongoing problem, not a solved one.