Bryan Cantrill on Thirty Years of Servers and the Cloud, and Building Oxide From a Clean Sheet of Paper

Open on YouTube ↗
Overview

Bryan Cantrill joined Sun Microsystems in 1996, stayed through the dot-com boom and bust, co-founded the cloud company Joyent, and is now co-founder and CTO of Oxide Computer. In this conversation on The Pragmatic Engineer, the host traces the history of server infrastructure with Cantrill: why the web ran on Sun and Cisco, how Linux, x86, and AWS changed things, why the hyperscalers stopped buying commodity servers, and why Oxide decided to design its own rack-scale computer, including its own switch, from first principles. The second half turns to how Oxide builds software, how it uses AI tools and where they fail, and how an 85-person company with uniform pay and remote work plans to keep its culture as it grows. Cantrill's recurring theme is that constraint and desperation produce better engineering than abundance does.

34 min read

The 1990s: why building a website meant buying Sun

Cantrill interviewed at Sun in 1995 and started in 1996. HTTP was only a few years old, early browsers existed, and Java had just come out and taken off right away. The energy in Silicon Valley was "extraordinary," but, as he recalls, nowhere near how frothy things became a couple of years later. Sun was in the right place with the right technology at the right time. If you wanted to build a website during the boom, you bought Sun servers and Cisco switches.

The host asked why you couldn't just run a server on a PC. Cantrill's answer was that PCs lacked the system software. Linux was then comparable to Haiku today: a hobbyist operating system most people hadn't heard of. The BSDs existed but were still under the shadow of the AT&T lawsuit. GNU's Hurd was "the Duke Nukem Forever of its time," a microkernel-based OS that was always arriving next year. So there were no serious open-source Unix options, and the server market belonged to the systems vendors.

Sun was a systems company. It made SPARC-based servers, desktop workstations, and some "ill-advised laptops." What exploded in the 1990s was servers, from workgroup machines up to very large systems physically about the size of what Oxide builds today. Cantrill remembers Sun's CTO at the time, Greg Papadopoulos, telling the company around 1997–98 that Sun's top three applications were "databases, databases and databases." A serious web presence then meant Java on Solaris on Sun hardware.

The boom: frenzy, a 1952 Sauternes, and a sudden collapse

The boom was frenetic, Cantrill says, and not always in a good way. He claims his group did "much more technically interesting work in the bust than we did in the boom." His explanation is that in boom times everyone privately believes the success is because of their own piece of the stack. One of the early engineers behind Java once told him, with a straight face, that every server Sun sold was sold because of Java. That was obviously false given "databases, databases, databases," but it captured the mood. Cantrill argues this attitude doesn't lead to real innovation. Innovation, he suggests, needs some desperation, and desperation is hard to summon when the economy is good.

He also learned that booms last longer than seem possible. For a while he believed the gloomy Economist covers predicting an apocalypse, then stopped believing them. The Economist turned out to be right; the boom just ran longer. And when it turned, it collapsed "faster than you can fathom."

Day to day, the boom meant terrible traffic, scarce housing, and customers buying at enormous scale. One customer planned to buy 19,000 one-rack-unit servers for a broadband initiative. The customer was Enron. Cantrill's most vivid memory is a dinner in September 2000 at Aqua, an expensive San Francisco restaurant that he doesn't think survived the bust, hosted with a bank that spent a "galactic amount of money" with Sun. The meal had many courses and ended with a 1952 Château d'Yquem, a Sauternes that wine lovers spend their lives hoping to taste. Cantrill, not much of a drinker, was too drunk by then to appreciate it. Back in his apartment he thought: this can't last. He says it felt as though the boom turned to bust that very night.

The signs had come earlier. Pets.com and the NASDAQ had crashed in early 2000, and traffic went from gridlock to "COVID-like" within about a month, with no pandemic, only a market collapse. What really stopped was the telecom buildout. Companies like JDS Uniphase, Global Crossing, and MCI WorldCom had been building infrastructure on the belief that the internet was the future. Cantrill stresses that they were right about that. Webvan delivered groceries, which many people do today through services like Instacart. Their timing was wrong and they lost sight of the underlying economics. In November 2000, Sun received zero orders from telecoms. After that came layoff after layoff, as companies built for endless good times now faced pessimism as extreme as the earlier optimism.

The bust: less money, more focus, and ZFS and DTrace

As a software engineer, Cantrill watched many people leave. He cites the statistic that U-Hauls leaving the Bay Area outnumbered those arriving 10 to 1. The people who had come because they cared about the technology stayed, and he says they were honestly not badly hurt. Everyone's equity evaporated; Sun lost 98% of its value. His coping strategy was to tell himself he never really had that money. The bust also reminded him that a boom can make you care about things you don't actually care about, because everyone around you is financially driven.

Fewer resources, he says, forced more creativity. At Sun in roughly 2001–2005, the systems software group produced ZFS, DTrace, and the Service Management Facility. Cantrill had joined Sun to work with Jeff Bonwick, who had wanted to rethink file systems since the mid-1990s. In the early 2000s, Bonwick and Matt Ahrens finally started from a clean sheet and built ZFS. Cantrill had "a chip on my shoulder" about how systems are observed and debugged, and with two colleagues built DTrace for dynamically instrumenting running systems. He doesn't claim all of this was caused by the bust, only that the timing lined up. He does think the bust made Sun open to new approaches, including open-sourcing the operating system in 2005, which he says gave these technologies "eternal life." Ideally, he adds, the industry would just be economically normal, but high tech seems to be either on or off.

Linux, x86, and the cloud displace the integrated server

By the time the host started in software in the late 2000s, Solaris was no longer the default. Cantrill names three shifts.

The first was open source. Linux grew up because companies like IBM, SGI, and Data General "backed up the truck" and contributed technology. XFS, still widely used on Linux, came from SGI's IRIX. Google was built on Linux from the start, and the companies of the next boom depended economically on open source. Linux, the BSDs, and eventually open-sourced Solaris gave builders many options.

The second was that SPARC "bluntly lost to x86." In the 1990s the fastest microprocessors were RISC chips: SPARC, MIPS, Alpha. Because Sun ran Solaris on both SPARC and x86, Cantrill could see how fast x86 machines were becoming, while Sun's microelectronics people dismissed Intel. In his telling, Intel focused on the "memory wall" and used speculative execution to pass the RISC chips. By 2004–2005, a leading-edge microprocessor meant x86, and the normal path became a Dell or Supermicro box running Linux or FreeBSD.

The third was the cloud, starting in 2006 with S3 and especially EC2 over the next few years. The host remembered a small mid-2000s company with its own hot server room and admins whom developers had to befriend to get anything deployed. Cantrill agrees that elastic, API-driven infrastructure mattered enormously, which he calls "not a deep thought."

Oracle, the lawn mower, and a nervous conference

In 2006 Cantrill started a storage group inside Sun. It succeeded well enough to attract Oracle as a customer for the first time in a long while. He jokes about lingering "shame" at possibly having drawn the "marine apex predator" that ate the company. Oracle's acquisition of Sun closed in early 2010, and he left soon after.

In a 2011 talk he warned people not to anthropomorphize Larry Ellison: treat Ellison like a lawn mower, which will cut off your hand if you put it in, without any anger. When audience members asked whether he feared retribution, he said they misunderstood; the lawn mower isn't angry at you. Then every video from the conference was posted except his. Colleagues suspected an Oracle conspiracy. Cantrill didn't, and says it wasn't one. What he had underestimated was the organizers' own fear of offending Oracle. When the video finally appeared, it opened with a disclaimer that the talk didn't represent the association's views, and the disclaimer was also shown above his head for the whole talk. He notes that people on Hacker News still cite "minute 33" whenever Oracle or Ellison comes up. He doesn't consider it a rant, just a description of what everyone knows.

2010–2014: AWS's relentless execution, and Kubernetes as the escape hatch

Cantrill describes 2010–2014 as a period of "relentless execution" by AWS with little real competition. Azure was "drifting out there," and GCP, though technically around since about 2009, was in his words "a joke." Every re:Invent brought a price cut that competitors dreaded and a new service that partners dreaded, because it competed with what they sold.

He calls Jeff Bezos "the apex predator of capitalism." The masterstroke, in his view, was persuading everyone that cloud was a terrible business. AWS didn't break out its financials, and annual price cuts made the market look like a bloody red ocean. Joyent was competing head-to-head, running a public cloud and also selling its cloud software to customers who ran it on Dell, HP, or Supermicro hardware. So Joyent knew the margins were actually good. Cantrill's reading was that S3 was funding Amazon's war on big-box retail, paying for Prime shipping. Several of Joyent's biggest customers were retailers who didn't want their spending funding a competitor.

For a while it seemed that competing in cloud required EC2 API compatibility. Eucalyptus tried and, in Cantrill's words, "it was just a disaster," and people assumed GCP and Azure could never catch up for that reason. What changed, he believes, was Kubernetes in 2015. Many customers used only basic infrastructure, not Elastic Beanstalk, Greengrass, or Redshift, and Kubernetes gave them a layer on which they could deploy anywhere. He argues that multicloud didn't really exist before Kubernetes and that much of its early momentum came from customers who felt locked into AWS.

The host mentioned that Kat Cosgrove, a Kubernetes release lead, had speculated on an earlier episode that Google open-sourced Kubernetes to help Google Cloud by making workloads portable. Cantrill agrees that this was likely the argument Kubernetes advocates made inside Google, which he describes as bottom-up at the time; "nobody prevented it." He recalls Craig McLuckie, who pushed to create the CNCF, saying a foundation would help Kubernetes get marketing dollars. Cantrill found that funny given how much cash Google had. Calling Kubernetes a master stroke gives Google "slightly too much credit, but only slightly." GCP is now a large, important business, and while Kubernetes isn't the only reason, he thinks it played a real role.

Why the hyperscalers build their own machines

The hyperscalers "never were" meaningful buyers of Dell or HP servers, Cantrill says. Early Google assembled machines from parts bought at Fry's and famously held them together with Velcro, reasoning that distributed software made hardware quality irrelevant, even skipping ECC memory. Cantrill's point is that DIMMs failing outright is survivable, but memory silently returning wrong data is not, because corrupted values end up in a database. Google overcorrected, and once the business was established, built machines far better engineered than commodity servers, including DC bus bars and power planning across the whole data center, as described in its book on the warehouse-scale computer. Facebook/Meta, Microsoft, and Amazon each independently reached the same conclusion.

The reason was scale. Dell, HP, and Supermicro designed for a server room with a rack of six, then twelve, then maybe 24 machines. For someone buying thousands of servers for a public cloud, those vendors had no product. At every level their machines were personal computers stacked together, not infrastructure designed for scale.

Joyent was acquired by Samsung in 2016, because Samsung's cloud bill was enormous and there was no product to buy that would let them bring it in-house, so they bought a company. That also meant one fewer company for the next Samsung to buy. When Cantrill and his co-founders considered what to do next in 2019, their thesis had two parts. First, cloud computing, meaning elastic, API-driven infrastructure, is the future of all computing. Second, you shouldn't only be able to rent it. You should be able to buy it and run it in your own data center for risk management, security, or economics.

The host noted that mid-sized companies like Basecamp did this with off-the-shelf servers in colocation facilities. Cantrill agrees Basecamp became a poster child for the economic advantage, and credits DHH's outspokenness. Basecamp is smaller than Oxide's target, though, and he says the economics are even more compelling at larger scale. He enjoys it when VCs who passed on Oxide for lack of a market send him DHH's blog posts. Oxide couldn't predict exact trends, but believed companies born on the public cloud would outgrow its economics and want to go on-prem.

Designing a rack from first principles: power, blind-mated networking, and a custom switch

The Oxide rack holds 32 compute sleds. From the start, Cantrill says, the team refused to build it from Dell, HPE, or Supermicro parts and instead began from the problem itself, and found a great deal of technical debt in the PC ecosystem.

Power is one example. A conventional rack has AC power into every server, with two power supplies per 1U or 2U chassis. Each supply has its own fans, which are densely packed, fight high static pressure, and are often what wears out first. At scale, the practice is an AC bus bar feeding an efficient power shelf that rectifies AC to DC and distributes DC along the rack, with sleds blind-mating into power at the back. Oxide planned this from the start.

Starting from a clean sheet also produced opportunities Oxide hadn't anticipated. The team assumed it would copy the hyperscalers and put networking cables at the front, in the cold aisle. Connectivity vendors asked why, if Oxide was starting over, it didn't blind-mate the networking too. According to the vendors, the hyperscalers would all do that if they started fresh but were now too afraid to change. That was "catnip" to Oxide, and blind-mated networking became an early bet-the-company decision: if it failed, they had nothing.

In the Oxide rack, sleds slide into a cabled backplane wired at the factory. There are no cables for the operator to handle and no miscabling. Cantrill notes that each computer actually sits on three networks: a power/presence-detect network, a service processor network, and the high-speed data network. Cabling all of that by hand in a facility is very error-prone.

Blind-mated networking depended on an earlier bet: building Oxide's own switch. No investor on Sand Hill Road asked about it, even though the team worried about it most. Without its own switch, Cantrill says, Oxide would face a third-party integration nightmare and couldn't meet its goal: a rack that comes out of the crate, gets wheeled into place, gets power and networking, and runs with minimal operator involvement. When the host suggested a switch sounds simple, Cantrill said to keep that attitude as long as possible if you want to build one, because otherwise you never will.

Switching silicon comes from "one and a half providers," essentially Broadcom, which he calls very proprietary. Oxide chose Intel's Tofino, from Intel's acquisition of Barefoot, because it offered truly programmable networking. Intel later killed Tofino, so the relationship is "complicated," but Oxide bought enough parts to buy itself time to design its next-generation switch. Cantrill says the custom switch paid off in many unexpected ways, and owning both sides of the connection is what made blind-mated networking possible. His broader lesson is that a big risk forces real deliberation, and once you commit, you often find dividends you didn't expect.

What designing a computer involves, and finding fearless electrical engineers

Asked what designing a computer actually involves, Cantrill says it's "very involved," mainly because everything is so fast. Signal integrity for DDR5 memory and PCIe is extremely complicated. "Digital is like a lie" that electrical engineers let the rest of us believe; these are analog signals racing through a substrate. He compares the CPU to a large airliner that needs an airport and runway: it needs a surrounding system to handle power sequencing, the power distribution network, environmentals, and its connections to memory and I/O. It is "fractally complicated," so most designers take vendors' reference designs and iterate on them instead of innovating.

Oxide needed electrical engineers willing to work from first principles, and they were hard to find. Engineers from traditional server makers created friction. Cantrill describes people used to asking a vendor's field applications engineer for the voltage regulator design, with no way to judge whether it was right. His reaction: then hire that person instead. The EE team Oxide eventually built, which he calls extraordinary and "absolutely fearless," came from outside the server industry, including people who had worked on CT systems at GE Medical.

Uniform pay as a recruiting signal

What changed Oxide's hiring was compensation. The team was brainstorming how to reach people outside Cantrill's personal network. An engineer pointed out that when explaining Oxide's values to outsiders, people shrugged until they heard the pay was transparent and uniform. Cantrill had assumed pay simply wasn't talked about publicly. The March 2021 blog post on the policy, he says, made hiring "nonlinear."

At the time the salary was $207,000. It has since risen, and Cantrill has lost track of the exact number; one employee thought an unexpectedly larger paycheck was an error after missing the end of an all-hands where the raise was announced. Cantrill says the attraction wasn't the equal pay itself. It was that a company would be "so nuts" as to do it, which convinced people Oxide took its values seriously. Everyone, including electrical, software, and support engineers, earns the same base salary as Cantrill.

To the common question of whether support engineers get the same pay, the answer is yes, and he says the result is what he believes are the best support engineers in the business. The host compared this to Gumroad paying support staff like software engineers and getting people who could fix code. Cantrill describes support as a technical challenge with immediate gratitude, and says several Oxide support engineers told him their hearts were always in support but career paths had pushed them elsewhere. He makes the same argument for QA. At companies where QA sits at a lower pay grade, as the host recalled from Microsoft about 15 years ago, the message is that QA matters less. Paying the same attracts "the best of the best."

The software stack: Hubris, the control plane, and the hardest problem, updates

Oxide's entire stack is open source. Cantrill jokes that Oxide has "God's own revenue model": anyone can run the software on other hardware, but the best place to run it is Oxide's machines, which aren't free. For the service processor, Oxide wrote its own operating system from scratch in Rust, called Hubris "because we had the hubris to do it," with a debugger called Humility. On the host CPUs (then AMD Milan, now AMD Turin), Oxide built its own hypervisor and its own control plane, named Omicron before the COVID variant made the name briefly awkward.

When customers power on a rack, they get a console that Cantrill says looks like AWS "if AWS looked better," plus an API and CLI. The control plane decides where instances are placed and where virtual storage lives. Customers can use Terraform and run Kubernetes on top without needing to know those details.

Among several "thorny" problems, Cantrill singles out updates. Public clouds rely on automation but also on humans with runbooks who step in when an update goes wrong. Oxide ships a distributed system that may run behind an air gap in a secure facility, often for customers who bought it specifically to run it themselves. If something goes wrong, Oxide can't be there.

To show why this is harder than updating a phone app, Cantrill lists what must be updated: the service processor, the root of trust, drive firmware, the host OS, and every component of the distributed system that talks to the others. During an update the system is partly on the old version and partly on the new. How should it behave in that hybrid state? What happens when the database schema changes between versions, which it has? How is each component updated, and how does the process stay robust?

Oxide's approach was to ship first a "minimum viable update," called MUPdate, which required parking the control plane: take the rack offline, update it, bring it back. It was robust and it worked, but it isn't what cloud users want, since their instances need to stay up. It did provide the foundation for building live update step by step, lighting up different parts of the system and automating more over time. The first fully automatic update ran on Oxide's internal "dogfood" rack. Cantrill credits Dave Pacheco and team for delivering in about the time they expected, which he calls rare for software, by carefully trading scope against schedule while treating quality as the fixed constraint. He calls Pacheco's internal talk looking back on two years of update work one of the best single talks on software you'll ever see.

How Oxide uses AI, and why "intelligence is not enough"

Cantrill says Oxide adopted AI tools early, but "no part of the Oxide stack is vibe coded." People use them in different ways: for tedious tasks, for generating test cases, and above all, in his view, for document comprehension. Oxide has a writing-heavy culture built around RFDs (Requests for Discussion), which he says makes a company "LLM ready," not for producing documents but for consuming them. In 2020 he spent about three hours trying to build an RFD glossary and gave up because the job "spreads to the horizon." That's now something an LLM could produce, though he hasn't found time to do it.

He is "definitely not a doomer," but thinks many people are being reductive. He gave a talk called "Intelligence Is Not Enough" about problems in building the Oxide rack that an LLM could never have solved. A prominent AI doomer made a reaction video, which his 11-year-old daughter found hilarious. What frustrated him was that the video fast-forwarded past the concrete technical examples, which were the substance of the talk.

The example he gives here is from Oxide's first board bring-up, the first time a new board is powered on. The CPU wouldn't come out of reset; after 1.25 seconds it reset itself. The team suspected a marginal power network, but AMD looked at the numbers and said the margins were very good. They eliminated one hypothesis after another for weeks, and Cantrill says they felt the company was "dead." What would you even tell an LLM? That it's not working? A desperate engineer finally examined the protocol between the CPU and the voltage regulator and noticed that when the CPU requested a voltage, say 0.9 volts, the regulator set it but never sent the acknowledgement packet. AMD's test tool, the SDLE, didn't care about the acknowledgement; the CPU did, so it kept resetting and retrying. The cause was a firmware bug in the Renesas controller, fixed with a firmware update. The Renesas FAE, whom Cantrill praises, said Oxide should have reached out much sooner. "Building a board is not an IQ test," Cantrill says. Intelligence is necessary but not sufficient.

He also stresses teams. Sometimes someone joins a Google Meet just to follow along, asks what they call a dumb question, such as whether some addresses look like similar virtual addresses, and that opens a new line of investigation. The host mentioned Armin Ronacher, creator of Flask, who is building a startup with a co-founder and "an army of AI interns" but wants to hire people because they bring energy. Cantrill agrees, citing Richard Sutton, the reinforcement learning pioneer, on the distinction between LLMs and artificial intelligence: LLMs don't have goals. "A prompt is not a goal and guessing the next word is not a goal." A startup team trying together not to die is a goal, and LLMs can be tools in service of it.

Where LLMs help: small, not large, and almost never in hardware

Cantrill uses LLMs heavily as an editor. When a blog post of his reached Hacker News and someone called it LLM-written, his response was that it was LLM-edited: the only change he made on an LLM's advice was deleting a paragraph it said wasn't working. For Rust, especially for newcomers, he finds it valuable to ask whether a short snippet is idiomatic. His summary is that LLMs are "more valuable in the small than in the large." He tips his hat to people who want to spend their lives as "middle management for robots," but that's not for him. At Oxide people own their work: "the LLM broke my code" is not an acceptable excuse, because LLMs have no accountability.

Across the team, Claude Code is used "a bunch," and Cantrill encourages experimentation. For much of Oxide's work, though, including C code in the OS kernel, AI is helpful "as maybe a polishing tool but less as at the epicenter of its creation," with some exceptions. When the host suggested that little has changed despite executives' productivity claims, Cantrill called it a powerful tool, not the only one. He pushes people who refuse to use LLMs to try them, and cites Simon Willison's advice to run models locally, where they're slow and poor, to understand their limits.

For hardware, his first answer to whether AI helps was "No. Zero," which he then softened: an LLM can help interpret an I²C transaction waveform and spot non-compliant behavior, but only at the edges. The host called this 0.01. Cantrill adds that hardware already relies on software: EDA tools with automatic signal-integrity rule checks, plus extensive simulation. He finds it frustrating that because programming is such a good fit for LLMs, some programmers conclude that AI will replace every job. "Not even close," he says; they need to "get outside a little bit more."

An 85-person company of many disciplines, working mostly remotely

Oxide will soon be around 85 people. Applicants write detailed materials about their past work, what matters to them, and why they want to join. Cantrill says much of his own LLM use is reviewing these materials, and he now sees heavily LLM-written applications. He asks applicants to stop. The worst case is someone who writes everything themselves and then has an LLM answer "Why do you want to work at Oxide?" His conclusion is that such a person doesn't really want to work there. The process, he says, attracts people drawn to the culture, the problem, and the team.

He believes every company has something to teach, even when that means "scraping the bottom of the barrel." Challenged to find something positive about Oracle, he points to Larry Ellison approving every hire. He disagrees with how Ellison does it but thinks a CEO bears responsibility for every hire and should look at each one, and connects this to Paul Graham's "founder mode" essay about founders losing track of hiring. He also welcomes the question of what he didn't want to copy from Sun. Oxide is often seen as Sun's second coming, but there was much about Sun he didn't love.

The host was surprised that most of Oxide works remotely despite building hardware. Cantrill explains that a server, unlike a tractor or wind turbine, can be modeled in a basement, and much hardware work, including layout in tools like SolidWorks and Altium, is software that can be done anywhere. Bring-up happens at the manufacturer anyway, so it wouldn't be in an office either. Engineers from the electronics industry often ask whether they'll have to spend weeks in windowless offices in Taipei, Beijing, or Shenzhen. Oxide's assembly is done in the US, at Benchmark Electronics in Rochester, Minnesota, where a group of employees was working the week of the recording.

Scaling without losing the culture

With a large Series B and, which Cantrill considers more important, strong customer traction, including customers who bought one rack and now want many more, Oxide is growing quickly. His main concern is that companies often "take their eye off the ball" on hiring. Oxide intends to keep absolute discipline, and he says it has an advantage: every employee went through the same values-based process, so no one needs convincing of its importance. The company will get bigger, but "the bones aren't changing."

Only Cantrill and Steve now know everyone at the company. Cantrill compares all-hands gatherings to the parties he threw in college, which were great not because of him but because his roommates came from unrelated groups: a computer science student who played ultimate, an engineer on the water polo team, a history student in the chorus, plus the women's swim team, always invited. People loved meeting others they'd never have met otherwise, and Oxide gatherings produce the same delight when colleagues discover each other. He tells the team they have "lightning in a bottle," must not take it for granted, and must meet customer needs while protecting what got them here.

AI, anxiety, and advice for junior engineers

Looking ahead, Cantrill calls AI a revolution that will let everyone do more, but finds talk of AGI and replacing all jobs "distracting kind of nonsense." The focus should be on putting more powerful tools in human builders' hands. He adds that humanity is running many AI experiments that may not make economic sense, and that will have to be worked out.

The host raised the anxiety among junior and even experienced engineers, fed by stories of layoffs attributed to AI that often turn out to have other causes. Cantrill notes that the dot-com bust also eliminated many jobs. He acknowledges the current disruption feels broader and more permanent, and argues society should encourage new company formation, since small teams like Ronacher's can now do much more. Engineers find livelihood and meaning by building useful things, and the question to ask is: if you could build anything, what would you build? That's scarier than the old path of school, the right concentration, and "mama Google" hiring you and feeding you breakfast. There's less job security, he concedes, but more opportunity.

For a student aiming to work somewhere with a high bar like Oxide within five years, Cantrill's advice is a mindset shift from "how do I create as much as possible" to "how do I get better every day." He compares it to a strong high school player aiming for Major League Baseball: it's hard, uncertain, and requires realistic, daily focus on improvement. People should admit how much they don't understand. He jokes that he keeps waiting for the day he'll know how computers work, and says he learns new things daily, not only about computers but about delivering them to customers. LLMs should be seen not as taking your job but as a private tutor you can ask anything, while fact-checking its answers. It's easier than ever to get into a new domain, which is powerful and also scary.

Three books

Cantrill closes with three books he recommended to his son for a high school assignment. The Soul of a New Machine by Tracy Kidder, a Pulitzer winner about building a computer at Data General, is one he thinks every engineer should read and will see themselves in. Skunk Works by Ben Rich, about Lockheed's Skunk Works founded by Clarence "Kelly" Johnson, shows what engineers can do when they take on the impossible. Steve Jobs and the NeXT Big Thing by Randall Stross was written before Apple bought NeXT, at Jobs's lowest point; it is "not here to praise him" but "to bury him." Cantrill notes that NeXT gets only a few pages in the Isaacson biography, and believes Jobs's failures across the 13 years at NeXT were essential to Apple's resurrection, because the Jobs who returned behaved very differently from the one who was fired. He finds Jobs enigmatic, admires some of what Jobs did, strongly disagrees with other parts, and thinks people should examine Jobs critically instead of simply lionizing him.