Thuan Pham on Scaling Uber: Surviving Hypergrowth, Rewriting Under Pressure, and What Still Makes Engineers Great

Open on YouTube ↗
Overview

Thuan Pham joined Uber as its first CTO in April 2013. The company then had about 40 engineers and roughly 30,000 rides a day, and its systems crashed several times a week. Pham stayed seven years. In this conversation with the host of The Pragmatic Engineer, who worked at Uber for almost four of those years, Pham explains how the engineering organization got through that period: rewrites made against hard deadlines, an org structure invented because the old one had stalled, a China launch that outside peers called impossible on that timeline, and a microservices explosion nobody planned. The same view runs through all of it. Under violent growth, the job is to buy enough runway to survive, not to build the perfect system. Pham also talks about his current role as CTO of Faire, how the company uses AI, and why he thinks the traits that make a great engineer have not changed.

30 min read

From refugee boat to MIT

Pham was born in southern Vietnam. His father was tied to the southern military, and after the country was unified in 1975, families connected to the old regime had few opportunities, including in education. Pham notes that this is not necessarily true today. His mother decided her two sons would not grow up that way, and the family joined the "boat people" exodus. By Pham's account, about two million people left and about a million survived the crossing, because the boats were not seaworthy. People did not dwell on the odds, he says. If they had, they probably wouldn't have gone.

Getting out took four attempts and drained the family's savings, because several of the earlier arrangements were scams: pay half now, and the boat never comes. On the fourth try they had a good captain who got them through storms and past Thai pirates. Pham was about 11 or 12. After three days and four nights crossing the South China Sea they reached Malaysia. A week later they were towed back out to sea and left near Indonesia, which took them in and placed them on a deserted island that became a refugee camp. The United States accepted them because of the family's ties to the US-backed regime. They arrived with no English and no money, sponsored by a church, wearing clothes from its donation closet.

Pham found computers through a high school friend whose father gave him an IBM PC with two floppy drives in 1982. They wrote BASIC programs and learned Lotus and WordStar, and Pham found that thinking algorithmically came naturally. He calls himself a procrastinator who hates doing the same thing twice, and says programming suited him: you solve the problem once creatively, and then the machine repeats it faster than any person could. He volunteered at a government agency and tied Lotus, dBASE III and scripting languages together to automate its financial accounting. Two accountants had spent about three weeks every quarter reconciling the books. Pham turned that into a batch job started with one button that ran in about three hours. The agency's recommendation letter helped him get into MIT, where, he says, he learned the actual fundamentals of computer science.

Early lessons: research that never shipped, and technology ahead of its time

An MIT co-op program matched Pham with Hewlett-Packard Laboratories. He did joint bachelor's and master's thesis work there and was hired into the lab afterward without a PhD. In the mid-to-late 1980s the lab was working on medical informatics: a networked, distributed architecture where a patient's records and X-rays followed them to any physician workstation, plus a knowledge base that checked for drug interactions. Pham was frustrated that the work ended in published papers. Product divisions picked research up only at an annual tech fair, and he wanted to write code people would actually use.

He moved to Silicon Graphics, which was building interactive TV: video on demand, online shopping and online games delivered to about 4,000 cable-connected homes in a trial. They invented network protocols along the way and put an SGI box on top of a tube TV. Celebrities, including Michael Jackson and Steven Spielberg, came to see the demos. Pham says the team really believed this was the future, and it was, but "way way ahead of the time." About a year in, the economics were clearly unworkable. Provisioning the head-end cost around $100 million in 1994, and the set-top box was a $45,000 workstation. A similar trial with NTT in Japan went well and then fizzled for the same reason. His lesson was that technology alone isn't enough. You need the right place, the right time and the right price point.

NetGravity: when a competitor chose growth over profit

Pham next joined a startup founded by a former SGI officemate. It was first called Netvertiser and quickly renamed NetGravity, and it built enterprise software for serving dynamic banner ads on sites like CNN and Netscape. The premise was that advertising had funded TV, so it would fund the free internet. Pham believes he was the fourth engineer. He says he and a colleague put the first dynamically targeted ad on the Yahoo page, moving from a script that rotated a static banner every hour to targeting based on page content, cookies and ad sequences.

The company went public, but Pham draws a second lesson from it. A later competitor built an ad service bureau: publishers pasted a tag into their HTML and revenue came in, with far less investment required on their side. NetGravity had wanted to do something similar, but a board member pushed it to reach profitability before expanding. NetGravity became a larger, more robust enterprise product while the other company spread across the internet and was eventually bought by Google. When the host sums it up as a growth-focused, unprofitable player being able to swallow a profitable one, Pham agrees and points to Uber as a company that made the other choice.

Over seven or eight years Pham went from individual contributor to VP. He moved into management because his SGI experience taught him that big things require leveraging other people.

The dot-com bust and investing in yourself during "peacetime"

Asked what the bust felt like, Pham describes a period of exuberance when everything was a ".com", including Pets.com and Webvan (he still has a Webvan bin in his garage), followed by a shakeout. Growth alone eventually burns through money, he argues. Durable companies need a value proposition customers will pay for. He expects a similar sorting among today's AI companies, with the market deciding in the end.

The downturn lasted a couple of years, and hiring was hard, especially for new graduates, who he says are always hit first because companies retrench toward experienced people. His advice is to invest in your skills in peacetime and never become complacent. Strong, hungry people who punch above their weight stay marketable even in a downturn, and people who let their skills atrophy find hard times very difficult to recover from.

Going small again, then VMware

Pham tested himself on purpose. After reaching VP at a company of hundreds, he wondered whether his success was "a fluke." He joined a four-person startup, a classic leaky-roof operation, which grew to 40–50 people in about three years before running out of money. It was acquired by a company building a security appliance for intermediating web services traffic, a niche Pham says was hard to break out of. The venture didn't succeed commercially, but he says his skills kept improving, and that you have to trust that working hard makes you better whether or not the vehicle wins.

He then joined VMware while it was still fairly small, in a 40-person division building enterprise software to tie hypervisors into a management and cloud platform. That team built VMware's first product suite, VirtualCenter, which worked with ESX. Pham considers vMotion the key feature: live migration of a virtual machine between hardware with no perceptible downtime. In his view it made a thousand machines look like one and turned VMware into something like a cloud operating system.

He eventually ran an 800-person engineering team. After eight years the product was in its third generation and mostly getting new features, the original founders had left, and the leadership had changed. Pham says he gets nervous when he feels too comfortable, and he decided it was time to go.

How reputation led to Uber, and a 30-hour interview

Pham says he never ran a job search in the usual sense. Doing good work and treating colleagues well slowly builds a reputation, and when you become available, people come to you with options. Uber came through Bill Gurley of Benchmark, who knew Pham from NetGravity a decade earlier. Gurley showed him Uber's board deck, and Pham understood the business model. Later, when he tried to recruit people, the constant question was why he was joining a taxi company.

According to the host, Uber had just raised a $30 million Series B at a $300 million valuation. The interview process with Travis Kalanick is the notable part. Kalanick spent more than 30 hours interviewing Pham one-on-one over about two weeks, at least a couple of hours a day. The first meeting, planned for an hour, ran two. The two went straight to a whiteboard, and Kalanick wrote out a long list of general topics (hiring, firing, communication, design), a shorter list of engineering-specific ones (code quality, QA, design) and a list of five things he wanted in an engineering team and its culture. Pham had barely left the building when the recruiter called to set up more time. From then on, they held a two-hour Skype session each day while Kalanick traveled between regional offices, taking one topic at a time. Pham still has a photo of the whiteboard on his phone.

Pham says that after a while he forgot he was being interviewed. It felt like exchanging ideas, disagreeing and working things out. He later came to see it as a simulation of what working together would be like. Kalanick, he says, viewed Uber as having two engines, physical operations and technology, with neither more important than the other. Pham was struck by the commitment: in one session that ran long, Kalanick called his assistant mid-conversation to move a flight so they could keep going.

Dispatch: five months until the brick wall

When Pham started, Uber ran only black cars in about 20–30 cities. The roughly 40 engineers were young, scrappy and talented, and the product experience was excellent, which drove word of mouth. But the system had been built for functionality, not scale, and it went down multiple times a week. Kalanick had told Pham in the interviews to "see around corners," so after his first weeks of building relationships and trust, Pham started asking what would break first. The answer was dispatch, the service that matches riders and drivers, without which there is no business.

Dispatch was a single-threaded Node.js process. As a city grew, the three- or four-person team moved its process to a faster machine. Pham says part of his role was to teach, so instead of declaring the design broken he asked leading questions. What happens as the city grows? Move to a faster processor. What happens when you reach the fastest one? Use multiprocessor boxes and run several processes. Do those processes share state? Not really. The engineers worked out for themselves that the approach was hitting a wall.

He then asked for the biggest city by ride volume, which was New York, and when it would exceed even the largest available machine. It was about May, and the answer was October. Dispatch had to be rewritten, and Pham gave only two requirements: a city must be served by multiple boxes, and a box must be able to serve multiple cities. No new features. With that N-by-M property the company could add hardware and, in principle, scale indefinitely. The simple requirements let the team move fast, and the new system was deployed around August or September, shortly before the limit. The database and the API monolith were next in line, and the pattern repeated: estimate how much runway remains before a threat becomes fatal, then get ahead of it.

That is why Uber rewrote systems so often, Pham explains. The faster you grow, the less runway any given architecture gives you. Speculating about eventual size wasn't useful, he says. What mattered was how long the company had before hitting a wall. If he had told engineers to build a system that would scale forever, it might have taken a year, and "we'll die before then." A quick rewrite bought perhaps 12 months, during which the team could plan the next phase. His phrase for this is buying enough runway to "live to fight another day."

China: from "two months" to five

Around Christmas 2014, at Uber's office at 1455 Market Street in San Francisco, Kalanick announced that Uber would launch in China in the new year. A requirement had emerged that services run on physical servers in China. Until then Uber had tested the waters by serving China from the US. Kalanick asked for two months and asked why it couldn't simply be done by racking machines and copying the software.

Pham explained that a copied system works on day one and then diverges, and Uber didn't have twice the engineers to maintain two drifting systems. The right approach was to re-architect into one system that could be partitioned. There were also serious security concerns. Nothing running in China could be assumed to have any privacy, while everything elsewhere had to be protected, so data and controls had to be fully partitioned while every deployment still went everywhere. The technical program managers scoped it at about six months, which was the fastest they could imagine. Industry friends Pham benchmarked with laughed and said 18 months at minimum. Kalanick didn't like six months either, so they "split the difference" at four and got to work.

At four months they were about a month short, and Kalanick was unhappy. At five months they were close but about to slip again. Pham negotiated: the team was confident it could launch within a month if allowed to roll out incrementally, a batch of cities per week over a few weeks. Kalanick agreed on the condition that the biggest city, which Pham names as Chengdu, go first.

Pham has come to think this was the brilliant part. Starting with the hardest city meant everything afterward was downhill, and the team went into each new batch with confidence. Starting with the smallest city looks like good risk control, he says, but it would have meant holding their breath every week. After launch, some burnt-out engineers took a month off and "stared at the water." Afterward, Pham says, the team wasn't afraid of anything. The host adds that a friend on the IT team had about two weeks to physically set up the servers, and that every engineer he later spoke to described their own piece as impossible, yet it came together. Uber then competed head-on with Didi. Pham connects the episode to Kalanick's idea that you sometimes have to be willing to "redline" yourself to find out you can do more than you thought.

The program/platform split

The org structure came before the microservices, and Pham says it came out of necessity. Between April and the end of 2013, engineering grew from 40 engineers and three product managers to about 100 engineers and a dozen PMs. Even at that size the functional structure ground work to a halt. Every feature had to be queued on the capacity of the mobile team (eight to ten developers), dispatch (five to eight people), backend and infrastructure. Trade-offs couldn't be managed because each feature needed negotiations with many teams, and engineers complained right away, which Pham considers healthy.

Kalanick, Jeff Holden and Pham spent a couple of days working it out with sticky notes, one color per function and one name per note. Kalanick listed what he saw as the business's most important areas, which came to 17 at the time (there turned out to be far more). They had enough people to fund seven, plus part of the next four, and the rest stayed empty until hiring caught up. The principle was that each team had to be cross-functional, with every skill needed to get its work done on its own. Teams building what end users touch became "programs." Teams building tools and layers that programs use became "platforms," a vertical-versus-horizontal split. Then they placed the sticky notes into the boxes.

Why Uber ended up with thousands of microservices

"None of us wanted to go through that extreme," Pham says. Under constant pressure, the priority was speed, and the backend API monolith was clearly what would slow things down. The rule became that anything new had to be built outside it as a microservice, while a dedicated team broke the monolith apart in a project called Darwin.

Pham estimates that with time frozen, the decomposition would have taken three to six months. It took two years because the business kept moving: new cities, new products such as UberX, and overlapping hockey-stick curves. The operating philosophy was that nobody could block anybody else, so teams that needed a feature in code not yet extracted added it to the monolith. The monolith grew even as pieces were pulled out, because the remainder grew faster than the extraction, until it eventually peaked and began to shrink. Meanwhile everything new fanned out into services. Once growth became less violent, a project called ARC cleaned things up by putting domain interfaces over groups of related services. The host notes that a 2016 Uber blog post cited about 5,000 microservices and a recent one about 4,500, so the count has fallen slowly while the business has become more complex. That complexity also required new tooling, such as the tracing tool Jaeger, which Uber open-sourced.

Internal tools, born from breaking open source

The host lists some of Uber's internal and open-sourced infrastructure: Jaeger, the Schemaless trip datastore, the TChannel RPC protocol, Ringpop, observability tooling, and hundreds more. Were they all needed? Pham says he can't claim every one was, "but all the important ones were absolutely necessary." Early Uber used off-the-shelf open source such as Redis. In 2013–2016, he says, open source was less mature, and big companies like Google and Facebook kept their infrastructure internal.

His most painful example is PostgreSQL. At a certain scale it began failing randomly, taking services down, with the problem deep in the kernel. Pham recalls asking people on LinkedIn for anyone with Postgres expertise to consult, and spending several weeks on it. What frightened him was depending on something nobody owned. He would have paid anything for an answer, and there was no company to pay. That experience helped motivate Uber to build its own data layer, using MySQL only as a table store with Uber-built logic on top, so it controlled its own state and built only the features it needed.

Other limits followed. Around the 2015 holidays Pham took an Uber to the airport and got the receipt two days later, because data processing was at capacity and work was queuing up. Delayed receipts weren't a dealbreaker, since the ride itself had happened, but it meant more rewrites. Monitoring built on open-source tools was also near its breaking point, which led to M3. Uber, he says, had reached a scale that broke the open-source tools it used.

Helix, and the "not a Mickey Mouse shop" email

Helix was the full rewrite of the rider app, where the host first met Pham. The host describes a codebase of one to two million lines and two to three hundred mobile engineers. Pham corrects the idea that it was mainly a design cleanup. The vision came from Kalanick and lead designer Yuki, who storyboarded it together. The old app did "push a button, get a ride" well but was too limiting to host more services, such as messaging during a ride, and the new architecture was far more open. The backend changed too, including the move from polling every five seconds to a push-based real-time system. Pham says it took perhaps 600–700 engineers across seven to eight months. He notes the app still runs on that architecture today, calls the design almost future-proof, and gives the credit to Kalanick and Yuki.

Pham also recalls an all-engineering email about naming. The trigger was a service named "Mustafa," whose purpose he couldn't tell. As the company kept onboarding engineers and tooling for mapping names to meaning was still weak, whimsical names with no context slowed people down. He wrote that Uber is "not a Mickey Mouse shop." He admits mass emails can have side effects without solving the problem, but says the point was that at that scale the company had to take itself seriously.

Engineering levels and making internal transfers easy

Pham stands by splitting the L5 senior level into L5A and L5B: "I'm not apologizing for it." Uber benchmarked its staff-engineer bar against companies like Google and Facebook, and getting from engineer II through senior to staff could take five years. Splitting the level gave people a sense of progress and a meaningful stopping point for those who might never reach staff. It worked for a while, then people adjusted and pushed to reach staff faster. Pham held the line while he was there. After he left, he says, the levels shifted down so that L5B became staff, which he calls inflation he didn't want.

In 2016 he announced an easy internal transfer process after hearing engineers felt unsupported by their managers. His reasoning was that people who resign to join another company don't ask their manager's permission to interview, so requiring permission to move internally only made leaving easier than staying. He also expected it to push managers to develop their people so they would want to stay. It met considerable pushback, but they did it anyway, along with an internal job board mirroring external postings. His line: "It's not a jail."

Relationships as the real network

The host asks about a talk former colleagues remember, about seeing work from the perspective of death. Pham doesn't recall the exact speech but says the idea is always with him. The most accomplished people don't take themselves too seriously. Deference comes from the position, not the person, and the world forgets you once you leave the role. So he measures himself by how many people remember him as good or helpful to them. It shouldn't be a networking tactic, he stresses, because doing it for that goal is artificial. Just be genuine.

He adds that the relationships also let him deliver. Uber's engineers were young and hadn't built reliable systems at scale, while Pham's VMware network could. His first pull was an engineer named George for dispatch. More came for payments, and for Schemaless he recruited the top four engineers from his VMware team in Denmark, which is how Uber's Denmark infrastructure office started. Everyone he called for help came, after first asking "why a taxi company?" Some people have worked with him at five companies over 28 years. The same logic explains Uber's nine engineering offices: great talent doesn't necessarily move to San Francisco, so Uber brought work to it. That included a small but excellent infrastructure team in Denmark and a strong DevOps team in Lithuania, each given first-class ownership. Pham says cost savings weren't the reason.

Three tours of duty, and leaving

Pham describes his Uber years as three tours, each with a purpose, re-evaluated at the end. The first 18–24 months were about fixing what was broken and making things reliable. The second was worldwide scale, including China. Around 2017, with things stabilized, he was ready to leave that summer. A new senior technical leader had been hired, someone Pham rated highly who had done bigger things at Google, and he felt at peace handing over. That didn't work out, Uber had a very rough year, and he signed up for a third tour: getting the company through the turbulence. He knew only the condition for it to end, the arrival of a new CEO. He says he felt he owed it to the tens of thousands of people, past and present, who had built Uber. After the new CEO came, he stayed until 2020.

Pham says his departure wasn't about COVID. With money no longer a factor, he asks three questions: does he love the mission, is he making a big impact, and does he enjoy the people? When several of those are lacking, the whole stops being enjoyable. He was running a big job rather than building, and felt it was better for someone else to take it on.

Coupang, Nubank and Faire

The retirement didn't last, and Pham blames COVID. A summer of travel, including an African safari with his daughter before high school, was cancelled. Bored at home, he took many calls, including one with Coupang's founder, and joined to help. He learned about Amazon-style logistics with deliveries where an order placed before midnight arrives by 5 a.m., and he rode along on delivery trucks at 2 or 3 in the morning. He also joined Nubank's board and for a while mentored a couple of its CTOs. He compares its energy to early Uber and credits its success to solving the right problem at the right time for a large unbanked population, a well-liked product with very high NPS, and a strong culture. He attends an annual all-hands in Brazil and sometimes holds Uber-style AMAs with engineering.

He then took a couple of years off while his daughter finished high school, driving her to school, cooking and helping with college applications. He says he would not take that time back. As she left for college he nearly joined another board, but a Sequoia partner introduced him to Faire's CEO, Max. Faire is a B2B wholesale marketplace between brands and retailers, and Pham sees its mission of helping local businesses flourish as similar to Uber's. He was drawn by how fast the company moved (the process, including a homework presentation, finished within a week) and by a kind, low-politics culture. Faire has about 1,000 people, about 300 in engineering and data science, works in the office three days a week, and has engineers in San Francisco and large offices in Canada, including Toronto, which Pham visits every five or six weeks.

How Faire uses AI, and what separates great engineers now

Pham calls AI the most exciting challenge. Faire uses it to raise productivity across the company, to improve search and recommendations (imagining AI as a shopping consultant), and for coding. The team uses "swarm coding," many agents working in parallel, and is building an orchestrator for them. Early adopters showed a dramatic lift, and after more robust tooling the bulk of engineers followed. Pham compares the shift to going from single-threaded to multi-threaded programming. You prompt many actions, review and stitch together what comes back, and carry a higher cognitive load. He says their best engineers have doubled their output, meaning impact, not lines of code. For now, he says, AI makes large-scale changes and cleanups easy. The frontier they are still working on is getting similar gains when building new features on top of millions of lines of older, entangled code.

He calls this change faster than anything he has seen, including the internet. Programming once required knowing machine architecture and virtual memory. Now people who can't program can produce decent-looking apps. Still, he reports that great engineers differ from average ones by about 2–3x, because they are more inquisitive, stay at the bleeding edge and keep pushing boundaries, while others settle for the productivity boost the tool gives them. The traits are the ones that always marked standout engineers: fearlessness and willingness to stretch and try new things. "Complacency is death," he says. AI is powerful, but it is still a tool you can use well or in a mundane way.

The CTO's job, and advice by career stage

Pham sees two sides to the CTO role. The first is building a high-performing team: structure, talent development and pruning, until talent density sustains itself because A players want to hire A players and won't tolerate underperformance. Talent alone isn't enough without trust and cultural alignment, and if the organizational side is right, he believes, good outcomes follow. The second is seeing around corners. Pham looks 18–24 months out while his teams handle the six-month horizon, which he considers essentially locked in and a matter of execution. At Faire that means cleaning up an old data ecosystem where upstream changes break things downstream (a problem he also saw at Uber), defining the next generation of AI-driven search and discovery, and working out how AI can double output on feature development. Then he asks whether the team has the expertise and leadership to get there, and whom to recruit if not.

For new graduates, Pham acknowledges it is a scary, bumpy time. Faire no longer hires large numbers of new grads directly. It hires them through a co-op program that brings in a cohort every four months, and the best receive offers, because without them there will be no senior engineers years from now. His advice is staged. As a student, volunteer and solve hard problems early. In the first five to ten years, seek the places where you learn most and are pushed hardest. At senior or staff level, go where you can make the biggest impact, possibly a smaller company with a bigger stage. At principal, senior director or VP level, shift toward coaching and bringing others along.