Bryan Cantrill on Thirty Years of Servers and the Cloud, and Building Oxide From a Clean Sheet of Paper
The Pragmatic EngineerBryan Cantrill joined Sun Microsystems in 1996, stayed through the dot-com boom and bust, co-founded the cloud company Joyent, and is now co-founder and CTO of Oxide Computer. In this conversation on The Pragmatic Engineer, the host traces the history of server infrastructure with Cantrill: why the web ran on Sun and Cisco, how Linux, x86, and AWS changed things, why the hyperscalers stopped buying commodity servers, and why Oxide decided to design its own rack-scale computer, including its own switch, from first principles. The second half turns to how Oxide builds software, how it uses AI tools and where they fail, and how an 85-person company with uniform pay and remote work plans to keep its culture as it grows. Cantrill's recurring theme is that constraint and desperation produce better engineering than abundance does.
The 1990s: why building a website meant buying Sun
Cantrill interviewed at Sun in 1995 and started in 1996. HTTP was only a few years old, early browsers existed, and Java had just come out and taken off right away. The energy in Silicon Valley was "extraordinary," but, as he recalls, nowhere near how frothy things became a couple of years later. Sun was in the right place with the right technology at the right time. If you wanted to build a website during the boom, you bought Sun servers and Cisco switches.
The host asked why you couldn't just run a server on a PC. Cantrill's answer was that PCs lacked the system software. Linux was then comparable to Haiku today: a hobbyist operating system most people hadn't heard of. The BSDs existed but were still under the shadow of the AT&T lawsuit. GNU's Hurd was "the Duke Nukem Forever of its time," a microkernel-based OS that was always arriving next year. So there were no serious open-source Unix options, and the server market belonged to the systems vendors.
Sun was a systems company. It made SPARC-based servers, desktop workstations, and some "ill-advised laptops." What exploded in the 1990s was servers, from workgroup machines up to very large systems physically about the size of what Oxide builds today. Cantrill remembers Sun's CTO at the time, Greg Papadopoulos, telling the company around 1997–98 that Sun's top three applications were "databases, databases and databases." A serious web presence then meant Java on Solaris on Sun hardware.
The boom: frenzy, a 1952 Sauternes, and a sudden collapse
The boom was frenetic, Cantrill says, and not always in a good way. He claims his group did "much more technically interesting work in the bust than we did in the boom." His explanation is that in boom times everyone privately believes the success is because of their own piece of the stack. One of the early engineers behind Java once told him, with a straight face, that every server Sun sold was sold because of Java. That was obviously false given "databases, databases, databases," but it captured the mood. Cantrill argues this attitude doesn't lead to real innovation. Innovation, he suggests, needs some desperation, and desperation is hard to summon when the economy is good.
He also learned that booms last longer than seem possible. For a while he believed the gloomy Economist covers predicting an apocalypse, then stopped believing them. The Economist turned out to be right; the boom just ran longer. And when it turned, it collapsed "faster than you can fathom."
Day to day, the boom meant terrible traffic, scarce housing, and customers buying at enormous scale. One customer planned to buy 19,000 one-rack-unit servers for a broadband initiative. The customer was Enron. Cantrill's most vivid memory is a dinner in September 2000 at Aqua, an expensive San Francisco restaurant that he doesn't think survived the bust, hosted with a bank that spent a "galactic amount of money" with Sun. The meal had many courses and ended with a 1952 Château d'Yquem, a Sauternes that wine lovers spend their lives hoping to taste. Cantrill, not much of a drinker, was too drunk by then to appreciate it. Back in his apartment he thought: this can't last. He says it felt as though the boom turned to bust that very night.
The signs had come earlier. Pets.com and the NASDAQ had crashed in early 2000, and traffic went from gridlock to "COVID-like" within about a month, with no pandemic, only a market collapse. What really stopped was the telecom buildout. Companies like JDS Uniphase, Global Crossing, and MCI WorldCom had been building infrastructure on the belief that the internet was the future. Cantrill stresses that they were right about that. Webvan delivered groceries, which many people do today through services like Instacart. Their timing was wrong and they lost sight of the underlying economics. In November 2000, Sun received zero orders from telecoms. After that came layoff after layoff, as companies built for endless good times now faced pessimism as extreme as the earlier optimism.
The bust: less money, more focus, and ZFS and DTrace
As a software engineer, Cantrill watched many people leave. He cites the statistic that U-Hauls leaving the Bay Area outnumbered those arriving 10 to 1. The people who had come because they cared about the technology stayed, and he says they were honestly not badly hurt. Everyone's equity evaporated; Sun lost 98% of its value. His coping strategy was to tell himself he never really had that money. The bust also reminded him that a boom can make you care about things you don't actually care about, because everyone around you is financially driven.
Fewer resources, he says, forced more creativity. At Sun in roughly 2001–2005, the systems software group produced ZFS, DTrace, and the Service Management Facility. Cantrill had joined Sun to work with Jeff Bonwick, who had wanted to rethink file systems since the mid-1990s. In the early 2000s, Bonwick and Matt Ahrens finally started from a clean sheet and built ZFS. Cantrill had "a chip on my shoulder" about how systems are observed and debugged, and with two colleagues built DTrace for dynamically instrumenting running systems. He doesn't claim all of this was caused by the bust, only that the timing lined up. He does think the bust made Sun open to new approaches, including open-sourcing the operating system in 2005, which he says gave these technologies "eternal life." Ideally, he adds, the industry would just be economically normal, but high tech seems to be either on or off.
Linux, x86, and the cloud displace the integrated server
By the time the host started in software in the late 2000s, Solaris was no longer the default. Cantrill names three shifts.
The first was open source. Linux grew up because companies like IBM, SGI, and Data General "backed up the truck" and contributed technology. XFS, still widely used on Linux, came from SGI's IRIX. Google was built on Linux from the start, and the companies of the next boom depended economically on open source. Linux, the BSDs, and eventually open-sourced Solaris gave builders many options.
The second was that SPARC "bluntly lost to x86." In the 1990s the fastest microprocessors were RISC chips: SPARC, MIPS, Alpha. Because Sun ran Solaris on both SPARC and x86, Cantrill could see how fast x86 machines were becoming, while Sun's microelectronics people dismissed Intel. In his telling, Intel focused on the "memory wall" and used speculative execution to pass the RISC chips. By 2004–2005, a leading-edge microprocessor meant x86, and the normal path became a Dell or Supermicro box running Linux or FreeBSD.
The third was the cloud, starting in 2006 with S3 and especially EC2 over the next few years. The host remembered a small mid-2000s company with its own hot server room and admins whom developers had to befriend to get anything deployed. Cantrill agrees that elastic, API-driven infrastructure mattered enormously, which he calls "not a deep thought."
Oracle, the lawn mower, and a nervous conference
In 2006 Cantrill started a storage group inside Sun. It succeeded well enough to attract Oracle as a customer for the first time in a long while. He jokes about lingering "shame" at possibly having drawn the "marine apex predator" that ate the company. Oracle's acquisition of Sun closed in early 2010, and he left soon after.
In a 2011 talk he warned people not to anthropomorphize Larry Ellison: treat Ellison like a lawn mower, which will cut off your hand if you put it in, without any anger. When audience members asked whether he feared retribution, he said they misunderstood; the lawn mower isn't angry at you. Then every video from the conference was posted except his. Colleagues suspected an Oracle conspiracy. Cantrill didn't, and says it wasn't one. What he had underestimated was the organizers' own fear of offending Oracle. When the video finally appeared, it opened with a disclaimer that the talk didn't represent the association's views, and the disclaimer was also shown above his head for the whole talk. He notes that people on Hacker News still cite "minute 33" whenever Oracle or Ellison comes up. He doesn't consider it a rant, just a description of what everyone knows.
2010–2014: AWS's relentless execution, and Kubernetes as the escape hatch
Cantrill describes 2010–2014 as a period of "relentless execution" by AWS with little real competition. Azure was "drifting out there," and GCP, though technically around since about 2009, was in his words "a joke." Every re:Invent brought a price cut that competitors dreaded and a new service that partners dreaded, because it competed with what they sold.
He calls Jeff Bezos "the apex predator of capitalism." The masterstroke, in his view, was persuading everyone that cloud was a terrible business. AWS didn't break out its financials, and annual price cuts made the market look like a bloody red ocean. Joyent was competing head-to-head, running a public cloud and also selling its cloud software to customers who ran it on Dell, HP, or Supermicro hardware. So Joyent knew the margins were actually good. Cantrill's reading was that S3 was funding Amazon's war on big-box retail, paying for Prime shipping. Several of Joyent's biggest customers were retailers who didn't want their spending funding a competitor.
For a while it seemed that competing in cloud required EC2 API compatibility. Eucalyptus tried and, in Cantrill's words, "it was just a disaster," and people assumed GCP and Azure could never catch up for that reason. What changed, he believes, was Kubernetes in 2015. Many customers used only basic infrastructure, not Elastic Beanstalk, Greengrass, or Redshift, and Kubernetes gave them a layer on which they could deploy anywhere. He argues that multicloud didn't really exist before Kubernetes and that much of its early momentum came from customers who felt locked into AWS.
The host mentioned that Kat Cosgrove, a Kubernetes release lead, had speculated on an earlier episode that Google open-sourced Kubernetes to help Google Cloud by making workloads portable. Cantrill agrees that this was likely the argument Kubernetes advocates made inside Google, which he describes as bottom-up at the time; "nobody prevented it." He recalls Craig McLuckie, who pushed to create the CNCF, saying a foundation would help Kubernetes get marketing dollars. Cantrill found that funny given how much cash Google had. Calling Kubernetes a master stroke gives Google "slightly too much credit, but only slightly." GCP is now a large, important business, and while Kubernetes isn't the only reason, he thinks it played a real role.
Why the hyperscalers build their own machines
The hyperscalers "never were" meaningful buyers of Dell or HP servers, Cantrill says. Early Google assembled machines from parts bought at Fry's and famously held them together with Velcro, reasoning that distributed software made hardware quality irrelevant, even skipping ECC memory. Cantrill's point is that DIMMs failing outright is survivable, but memory silently returning wrong data is not, because corrupted values end up in a database. Google overcorrected, and once the business was established, built machines far better engineered than commodity servers, including DC bus bars and power planning across the whole data center, as described in its book on the warehouse-scale computer. Facebook/Meta, Microsoft, and Amazon each independently reached the same conclusion.
The reason was scale. Dell, HP, and Supermicro designed for a server room with a rack of six, then twelve, then maybe 24 machines. For someone buying thousands of servers for a public cloud, those vendors had no product. At every level their machines were personal computers stacked together, not infrastructure designed for scale.
Joyent was acquired by Samsung in 2016, because Samsung's cloud bill was enormous and there was no product to buy that would let them bring it in-house, so they bought a company. That also meant one fewer company for the next Samsung to buy. When Cantrill and his co-founders considered what to do next in 2019, their thesis had two parts. First, cloud computing, meaning elastic, API-driven infrastructure, is the future of all computing. Second, you shouldn't only be able to rent it. You should be able to buy it and run it in your own data center for risk management, security, or economics.
The host noted that mid-sized companies like Basecamp did this with off-the-shelf servers in colocation facilities. Cantrill agrees Basecamp became a poster child for the economic advantage, and credits DHH's outspokenness. Basecamp is smaller than Oxide's target, though, and he says the economics are even more compelling at larger scale. He enjoys it when VCs who passed on Oxide for lack of a market send him DHH's blog posts. Oxide couldn't predict exact trends, but believed companies born on the public cloud would outgrow its economics and want to go on-prem.
Designing a rack from first principles: power, blind-mated networking, and a custom switch
The Oxide rack holds 32 compute sleds. From the start, Cantrill says, the team refused to build it from Dell, HPE, or Supermicro parts and instead began from the problem itself, and found a great deal of technical debt in the PC ecosystem.
Power is one example. A conventional rack has AC power into every server, with two power supplies per 1U or 2U chassis. Each supply has its own fans, which are densely packed, fight high static pressure, and are often what wears out first. At scale, the practice is an AC bus bar feeding an efficient power shelf that rectifies AC to DC and distributes DC along the rack, with sleds blind-mating into power at the back. Oxide planned this from the start.
Starting from a clean sheet also produced opportunities Oxide hadn't anticipated. The team assumed it would copy the hyperscalers and put networking cables at the front, in the cold aisle. Connectivity vendors asked why, if Oxide was starting over, it didn't blind-mate the networking too. According to the vendors, the hyperscalers would all do that if they started fresh but were now too afraid to change. That was "catnip" to Oxide, and blind-mated networking became an early bet-the-company decision: if it failed, they had nothing.
In the Oxide rack, sleds slide into a cabled backplane wired at the factory. There are no cables for the operator to handle and no miscabling. Cantrill notes that each computer actually sits on three networks: a power/presence-detect network, a service processor network, and the high-speed data network. Cabling all of that by hand in a facility is very error-prone.
Blind-mated networking depended on an earlier bet: building Oxide's own switch. No investor on Sand Hill Road asked about it, even though the team worried about it most. Without its own switch, Cantrill says, Oxide would face a third-party integration nightmare and couldn't meet its goal: a rack that comes out of the crate, gets wheeled into place, gets power and networking, and runs with minimal operator involvement. When the host suggested a switch sounds simple, Cantrill said to keep that attitude as long as possible if you want to build one, because otherwise you never will.
Switching silicon comes from "one and a half providers," essentially Broadcom, which he calls very proprietary. Oxide chose Intel's Tofino, from Intel's acquisition of Barefoot, because it offered truly programmable networking. Intel later killed Tofino, so the relationship is "complicated," but Oxide bought enough parts to buy itself time to design its next-generation switch. Cantrill says the custom switch paid off in many unexpected ways, and owning both sides of the connection is what made blind-mated networking possible. His broader lesson is that a big risk forces real deliberation, and once you commit, you often find dividends you didn't expect.
What designing a computer involves, and finding fearless electrical engineers
Asked what designing a computer actually involves, Cantrill says it's "very involved," mainly because everything is so fast. Signal integrity for DDR5 memory and PCIe is extremely complicated. "Digital is like a lie" that electrical engineers let the rest of us believe; these are analog signals racing through a substrate. He compares the CPU to a large airliner that needs an airport and runway: it needs a surrounding system to handle power sequencing, the power distribution network, environmentals, and its connections to memory and I/O. It is "fractally complicated," so most designers take vendors' reference designs and iterate on them instead of innovating.
Oxide needed electrical engineers willing to work from first principles, and they were hard to find. Engineers from traditional server makers created friction. Cantrill describes people used to asking a vendor's field applications engineer for the voltage regulator design, with no way to judge whether it was right. His reaction: then hire that person instead. The EE team Oxide eventually built, which he calls extraordinary and "absolutely fearless," came from outside the server industry, including people who had worked on CT systems at GE Medical.
Uniform pay as a recruiting signal
What changed Oxide's hiring was compensation. The team was brainstorming how to reach people outside Cantrill's personal network. An engineer pointed out that when explaining Oxide's values to outsiders, people shrugged until they heard the pay was transparent and uniform. Cantrill had assumed pay simply wasn't talked about publicly. The March 2021 blog post on the policy, he says, made hiring "nonlinear."
At the time the salary was $207,000. It has since risen, and Cantrill has lost track of the exact number; one employee thought an unexpectedly larger paycheck was an error after missing the end of an all-hands where the raise was announced. Cantrill says the attraction wasn't the equal pay itself. It was that a company would be "so nuts" as to do it, which convinced people Oxide took its values seriously. Everyone, including electrical, software, and support engineers, earns the same base salary as Cantrill.
To the common question of whether support engineers get the same pay, the answer is yes, and he says the result is what he believes are the best support engineers in the business. The host compared this to Gumroad paying support staff like software engineers and getting people who could fix code. Cantrill describes support as a technical challenge with immediate gratitude, and says several Oxide support engineers told him their hearts were always in support but career paths had pushed them elsewhere. He makes the same argument for QA. At companies where QA sits at a lower pay grade, as the host recalled from Microsoft about 15 years ago, the message is that QA matters less. Paying the same attracts "the best of the best."
The software stack: Hubris, the control plane, and the hardest problem, updates
Oxide's entire stack is open source. Cantrill jokes that Oxide has "God's own revenue model": anyone can run the software on other hardware, but the best place to run it is Oxide's machines, which aren't free. For the service processor, Oxide wrote its own operating system from scratch in Rust, called Hubris "because we had the hubris to do it," with a debugger called Humility. On the host CPUs (then AMD Milan, now AMD Turin), Oxide built its own hypervisor and its own control plane, named Omicron before the COVID variant made the name briefly awkward.
When customers power on a rack, they get a console that Cantrill says looks like AWS "if AWS looked better," plus an API and CLI. The control plane decides where instances are placed and where virtual storage lives. Customers can use Terraform and run Kubernetes on top without needing to know those details.
Among several "thorny" problems, Cantrill singles out updates. Public clouds rely on automation but also on humans with runbooks who step in when an update goes wrong. Oxide ships a distributed system that may run behind an air gap in a secure facility, often for customers who bought it specifically to run it themselves. If something goes wrong, Oxide can't be there.
To show why this is harder than updating a phone app, Cantrill lists what must be updated: the service processor, the root of trust, drive firmware, the host OS, and every component of the distributed system that talks to the others. During an update the system is partly on the old version and partly on the new. How should it behave in that hybrid state? What happens when the database schema changes between versions, which it has? How is each component updated, and how does the process stay robust?
Oxide's approach was to ship first a "minimum viable update," called MUPdate, which required parking the control plane: take the rack offline, update it, bring it back. It was robust and it worked, but it isn't what cloud users want, since their instances need to stay up. It did provide the foundation for building live update step by step, lighting up different parts of the system and automating more over time. The first fully automatic update ran on Oxide's internal "dogfood" rack. Cantrill credits Dave Pacheco and team for delivering in about the time they expected, which he calls rare for software, by carefully trading scope against schedule while treating quality as the fixed constraint. He calls Pacheco's internal talk looking back on two years of update work one of the best single talks on software you'll ever see.
How Oxide uses AI, and why "intelligence is not enough"
Cantrill says Oxide adopted AI tools early, but "no part of the Oxide stack is vibe coded." People use them in different ways: for tedious tasks, for generating test cases, and above all, in his view, for document comprehension. Oxide has a writing-heavy culture built around RFDs (Requests for Discussion), which he says makes a company "LLM ready," not for producing documents but for consuming them. In 2020 he spent about three hours trying to build an RFD glossary and gave up because the job "spreads to the horizon." That's now something an LLM could produce, though he hasn't found time to do it.
He is "definitely not a doomer," but thinks many people are being reductive. He gave a talk called "Intelligence Is Not Enough" about problems in building the Oxide rack that an LLM could never have solved. A prominent AI doomer made a reaction video, which his 11-year-old daughter found hilarious. What frustrated him was that the video fast-forwarded past the concrete technical examples, which were the substance of the talk.
The example he gives here is from Oxide's first board bring-up, the first time a new board is powered on. The CPU wouldn't come out of reset; after 1.25 seconds it reset itself. The team suspected a marginal power network, but AMD looked at the numbers and said the margins were very good. They eliminated one hypothesis after another for weeks, and Cantrill says they felt the company was "dead." What would you even tell an LLM? That it's not working? A desperate engineer finally examined the protocol between the CPU and the voltage regulator and noticed that when the CPU requested a voltage, say 0.9 volts, the regulator set it but never sent the acknowledgement packet. AMD's test tool, the SDLE, didn't care about the acknowledgement; the CPU did, so it kept resetting and retrying. The cause was a firmware bug in the Renesas controller, fixed with a firmware update. The Renesas FAE, whom Cantrill praises, said Oxide should have reached out much sooner. "Building a board is not an IQ test," Cantrill says. Intelligence is necessary but not sufficient.
He also stresses teams. Sometimes someone joins a Google Meet just to follow along, asks what they call a dumb question, such as whether some addresses look like similar virtual addresses, and that opens a new line of investigation. The host mentioned Armin Ronacher, creator of Flask, who is building a startup with a co-founder and "an army of AI interns" but wants to hire people because they bring energy. Cantrill agrees, citing Richard Sutton, the reinforcement learning pioneer, on the distinction between LLMs and artificial intelligence: LLMs don't have goals. "A prompt is not a goal and guessing the next word is not a goal." A startup team trying together not to die is a goal, and LLMs can be tools in service of it.
Where LLMs help: small, not large, and almost never in hardware
Cantrill uses LLMs heavily as an editor. When a blog post of his reached Hacker News and someone called it LLM-written, his response was that it was LLM-edited: the only change he made on an LLM's advice was deleting a paragraph it said wasn't working. For Rust, especially for newcomers, he finds it valuable to ask whether a short snippet is idiomatic. His summary is that LLMs are "more valuable in the small than in the large." He tips his hat to people who want to spend their lives as "middle management for robots," but that's not for him. At Oxide people own their work: "the LLM broke my code" is not an acceptable excuse, because LLMs have no accountability.
Across the team, Claude Code is used "a bunch," and Cantrill encourages experimentation. For much of Oxide's work, though, including C code in the OS kernel, AI is helpful "as maybe a polishing tool but less as at the epicenter of its creation," with some exceptions. When the host suggested that little has changed despite executives' productivity claims, Cantrill called it a powerful tool, not the only one. He pushes people who refuse to use LLMs to try them, and cites Simon Willison's advice to run models locally, where they're slow and poor, to understand their limits.
For hardware, his first answer to whether AI helps was "No. Zero," which he then softened: an LLM can help interpret an I²C transaction waveform and spot non-compliant behavior, but only at the edges. The host called this 0.01. Cantrill adds that hardware already relies on software: EDA tools with automatic signal-integrity rule checks, plus extensive simulation. He finds it frustrating that because programming is such a good fit for LLMs, some programmers conclude that AI will replace every job. "Not even close," he says; they need to "get outside a little bit more."
An 85-person company of many disciplines, working mostly remotely
Oxide will soon be around 85 people. Applicants write detailed materials about their past work, what matters to them, and why they want to join. Cantrill says much of his own LLM use is reviewing these materials, and he now sees heavily LLM-written applications. He asks applicants to stop. The worst case is someone who writes everything themselves and then has an LLM answer "Why do you want to work at Oxide?" His conclusion is that such a person doesn't really want to work there. The process, he says, attracts people drawn to the culture, the problem, and the team.
He believes every company has something to teach, even when that means "scraping the bottom of the barrel." Challenged to find something positive about Oracle, he points to Larry Ellison approving every hire. He disagrees with how Ellison does it but thinks a CEO bears responsibility for every hire and should look at each one, and connects this to Paul Graham's "founder mode" essay about founders losing track of hiring. He also welcomes the question of what he didn't want to copy from Sun. Oxide is often seen as Sun's second coming, but there was much about Sun he didn't love.
The host was surprised that most of Oxide works remotely despite building hardware. Cantrill explains that a server, unlike a tractor or wind turbine, can be modeled in a basement, and much hardware work, including layout in tools like SolidWorks and Altium, is software that can be done anywhere. Bring-up happens at the manufacturer anyway, so it wouldn't be in an office either. Engineers from the electronics industry often ask whether they'll have to spend weeks in windowless offices in Taipei, Beijing, or Shenzhen. Oxide's assembly is done in the US, at Benchmark Electronics in Rochester, Minnesota, where a group of employees was working the week of the recording.
Scaling without losing the culture
With a large Series B and, which Cantrill considers more important, strong customer traction, including customers who bought one rack and now want many more, Oxide is growing quickly. His main concern is that companies often "take their eye off the ball" on hiring. Oxide intends to keep absolute discipline, and he says it has an advantage: every employee went through the same values-based process, so no one needs convincing of its importance. The company will get bigger, but "the bones aren't changing."
Only Cantrill and Steve now know everyone at the company. Cantrill compares all-hands gatherings to the parties he threw in college, which were great not because of him but because his roommates came from unrelated groups: a computer science student who played ultimate, an engineer on the water polo team, a history student in the chorus, plus the women's swim team, always invited. People loved meeting others they'd never have met otherwise, and Oxide gatherings produce the same delight when colleagues discover each other. He tells the team they have "lightning in a bottle," must not take it for granted, and must meet customer needs while protecting what got them here.
AI, anxiety, and advice for junior engineers
Looking ahead, Cantrill calls AI a revolution that will let everyone do more, but finds talk of AGI and replacing all jobs "distracting kind of nonsense." The focus should be on putting more powerful tools in human builders' hands. He adds that humanity is running many AI experiments that may not make economic sense, and that will have to be worked out.
The host raised the anxiety among junior and even experienced engineers, fed by stories of layoffs attributed to AI that often turn out to have other causes. Cantrill notes that the dot-com bust also eliminated many jobs. He acknowledges the current disruption feels broader and more permanent, and argues society should encourage new company formation, since small teams like Ronacher's can now do much more. Engineers find livelihood and meaning by building useful things, and the question to ask is: if you could build anything, what would you build? That's scarier than the old path of school, the right concentration, and "mama Google" hiring you and feeding you breakfast. There's less job security, he concedes, but more opportunity.
For a student aiming to work somewhere with a high bar like Oxide within five years, Cantrill's advice is a mindset shift from "how do I create as much as possible" to "how do I get better every day." He compares it to a strong high school player aiming for Major League Baseball: it's hard, uncertain, and requires realistic, daily focus on improvement. People should admit how much they don't understand. He jokes that he keeps waiting for the day he'll know how computers work, and says he learns new things daily, not only about computers but about delivering them to customers. LLMs should be seen not as taking your job but as a private tutor you can ask anything, while fact-checking its answers. It's easier than ever to get into a new domain, which is powerful and also scary.
Three books
Cantrill closes with three books he recommended to his son for a high school assignment. The Soul of a New Machine by Tracy Kidder, a Pulitzer winner about building a computer at Data General, is one he thinks every engineer should read and will see themselves in. Skunk Works by Ben Rich, about Lockheed's Skunk Works founded by Clarence "Kelly" Johnson, shows what engineers can do when they take on the impossible. Steve Jobs and the NeXT Big Thing by Randall Stross was written before Apple bought NeXT, at Jobs's lowest point; it is "not here to praise him" but "to bury him." Cantrill notes that NeXT gets only a few pages in the Isaacson biography, and believes Jobs's failures across the 13 years at NeXT were essential to Apple's resurrection, because the Jobs who returned behaved very differently from the one who was fired. He finds Jobs enigmatic, admires some of what Jobs did, strongly disagrees with other parts, and thinks people should examine Jobs critically instead of simply lionizing him.
Can you tell us about the dotcom boom? We did much more technically interesting work in the bust than we did in the boom. There's a degree to which innovation requires some level of desperation that good economic times are kind of hard to summon that desperation.
How have AI tools changed how you're working at Oxide? Certainly we're using Claude Code a bunch and people are doing that, but for a lot of the work that we're doing it is helpful as maybe a polishing tool but less as at the epicenter of its creation.
Can you tell me what it actually means to design or build a computer? Oh, it's very involved. Yeah, it's very involved. So, first of all,
How have servers and cloud infrastructure evolved since the late 1990s and what is next? Bryan Cantrill was a distinguished engineer at Sun Microsystems during the dotcom boom and dotcom bust, built a small competitor to AWS called Joyent, and is now the co-founder at Oxide. Today, we go into the history of servers and the cloud from the late 1990s to today, the challenges of building hardware like the Oxide computer from scratch, how the Oxide team uses AI and why they find it practically useless for hardware engineering challenges, why Oxide builds everything as open source and how they manage to work remotely as a hardware startup, and many more. If you'd like to understand more about how the cloud works, and learn how a nimble hardware plus software startup operates, this episode is for you.
This podcast episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes below to learn more about them and our other season sponsor.
So Bryan, welcome to the podcast. Oh, it's great to be with you. Thanks for having me.
I'd love to jump back in time a lot, back in the 1990s, because you're someone who's been around the block, and back then you worked at some interesting companies including at Sun, and if you could give us listeners and viewers a sense of what was it like in the '90s in terms of software, servers, what was the vibe like?
Yeah, it was an interesting inflection point because I was interviewing in 1995. I started in 1996. So I would say that the internet, and I mean HTTP had been developed in like '93, '94. We had kind of the first web browsers but it was still very, very, very new and the internet was just kind of primed for takeoff. Java had come out in maybe 1995. Java had kind of taken off immediately. So there was a lot of really exciting energy, but it was nowhere near what it would become even a couple years later. It became very frothy of course, and it was exciting. It was very clear to me, I went to school actually on the East Coast, but just coming out here to Silicon Valley, the energy was extraordinary and I really knew that I wanted to come out here for my career.
So at Sun those next couple of years, I mean I got very lucky really, because Sun was in the right place at the right time with the right technology, which you know sometimes you only appreciate in hindsight, because it was so explosive. And if you wanted to build a website as part of that dotcom boom, you were buying Sun servers, you were buying Cisco switches.
Now why was this the case? Because again, just taking myself back, just being a bit naive, I would assume that, hey, I'm in 1995, I want to build a website. Could I have not just used a PC and spun up a server? Did it not work like that, or how did it work?
I mean a PC, like maybe, but you didn't really have an operating system, right? Because Linux is very, very new. Linux is not... I'll back down. Oh yeah, definitely. Linux would be like Haiku today, which is an operating system you haven't heard of for a reason. It's kind of like a hobbyist operating system. You know what I mean? You'd be like, what? No, you wouldn't. And then you kind of had the BSDs. FreeBSD was certainly out there, also still very much under the shadow though of this lawsuit from AT&T. So there's not really open-source operating system options.
There was the, actually this is kind of funny, because so where was the GNU option? It was going to be the Hurd operating system. So Hurd was kind of like the Duke Nukem Forever of its time. It was the operating system that was constantly coming kind of next year and next year and next year, and it was going to be microkernel based. And so it's kind of amazing, but you really couldn't do it on PCs because of the lack of system software. And actually part of my attraction to Sun was I had used Solaris on SPARC, but I knew Solaris existed on x86 but I never used it. So I was excited to use Solaris on x86.
And so what did Sun build? You mentioned Solaris. That was the operating system. Solaris is the operating system. We built servers. So we built SPARC-based servers. We built desktop machines. Sun was a computer company. It was a systems company. So we built desktop machines, built some ill-advised laptops. So basically desktop machines, workstations. But then at that time in the '90s, what was really exploding were everything from those kind of workgroup-style servers up to really getting bigger and bigger servers, up to the very large machines, machines that are physically the same size as what Oxide makes today.
And I remember vividly in what would have been like '97, '98 maybe, Greg Papadopoulos, then the CTO of Sun, giving it to the entire company saying, here are the top three applications for Sun Microsystems: databases, databases and databases. So that gives you an idea of kind of how it was being used. And this is again in that knee up of that dotcom buildout where if you wanted to really build a web presence, you were going to use Java, you were going to do it on Solaris, you're going to do it on Sun servers. And it was kind of a wild time for sure.
And can you tell us about the dotcom boom? Because right now I know AI is pretty exciting and it feels like we're in a special time, but what was it like, especially working on it? It sounds like it was the epicenter of it.
You know what was funny is it was frenetic in a way that was not always positive. So one of the things that is just a point of fact, and one can take from it what one will: we did much more technically interesting work in the bust than we did in the boom, because I think that when you're in boom times, everyone kind of secretly believes that this is because of me, that it is because of the thing that I am working on. I once had one of the early technologists behind Java once tell me with a straight face, every server that Sun sells, they sell because of Java. And I'm like, you know what's most amazing? You believe that, is actually the more interesting fact. I mean it is obviously false, especially with databases, databases, databases being the top three applications, but that kind of reflects the zeitgeist of the time, that everyone believes that if I work on the microprocessor, it's because the microprocessor is perfect. If I work on the operating system, it's because, oh, this is the operating system that people are buying the machine for. And that doesn't really lend itself to real innovation.
I think there's a degree to which innovation requires some level of desperation that good economic times, it's kind of hard to summon that desperation sometimes. So I think that during the boom it was just frothy, and it felt like there was a period of time where I'm like, this obviously can't go on forever, and The Economist is having these very gloomy covers about how this is all going to end and it's going to be an apocalypse, which I believed, and then I just stopped believing it. And I'm like, well maybe The Economist is right, it just went on longer. And one of my early life lessons from the boom and bust is these things go on longer than you think possible.
But when they... growth? In terms of the boom, when you're in frothy times, that boom will go on longer than you think possible. Mhm. And when it switches, it will collapse faster than you can fathom.
In the boom, do I understand correctly that customers were just wanting to buy your servers? They were flying off the shelves, all these companies. And on day-to-day work, what did it mean for you?
So I'll tell you, day-to-day it meant, first of all, that traffic was terrible. You couldn't get housing, everything was in short supply. Customers are buying. We had a customer that was going to buy 19,000 servers, which is obviously a very big number. And these were these massive big servers, right? Yeah, well in that case those were actually 1U servers to build out a broadband initiative, and that actually was a company called Enron.
I remember vividly we were at a dinner here in the city at a restaurant called Aqua, which is a very kind of fancy restaurant, long since out of business. I don't think Aqua survived the bust. And we were at Aqua with a bank who was a customer of Sun's, and they were spending a galactic amount of money every year with Sun, and we were at a dinner, and it was the kind of 19th-century Gilded Age kind of dinner. People are ordering nine courses. What I remember is at the end of that having Château d'Yquem, which is a Sauternes. I know very little about wine. I know nothing about Sauternes. What I did know is there was someone who knew wine, and it's like, we are going to all drink the 1952 Château d'Yquem. And I remember being, I'm not much of a drinker, but I was too drunk at that point to really appreciate it. So I have had this Sauternes that oenophiles kind of live their life to drink, and I'm sad to inform you that there's one less bottle of this precious vintage because it was poured down the gullet of a 20-something dotcommer. And I just remember being back in my apartment, being literally drunk on Château d'Yquem, and remember thinking to myself, this can't last, this is not sustainable. And I swear the dotcom boom turned to a bust like that night.
That is September of 2000. So Pets.com had kind of busted out and the NASDAQ had busted out early in 2000. The traffic got lighter early in 2000. Anyone who was here would say the absolute spookiest thing is it went from gridlock to COVID-like traffic in the span of like a month. Without COVID happening. Without COVID happening, with only the NASDAQ collapsing. And you're like, "Okay, that's very odd." And then 2000 kind of muddled along, and that dinner was in September of 2000, and what really stopped was the telco buildout. There was a lot of telco buildout because people are like, the internet is the future.
And telco buildout meaning the towers, the servers? The servers, the infrastructure, and then all of the concomitant fiber. JDS Uniphase was a huge company. You had these companies, Global Crossing and MCI WorldCom, and all these companies were explosive, and everyone believed that the internet is the future, and this is an important thing, and they were right. They were right.
Bryan just said how an important lesson of the dotcom boom was that people who believed the internet would be the future were right. Today we're in a similar stage with AI. It's pretty likely that AI will be part of the software stack in the future, even if timing is harder to predict. The latest shift is how AI agents are becoming a lot more commonly used for development. And this is a great time to talk about our season sponsor, Linear, and how they think about collaborating with agents. Linear has taken an interesting approach here. Instead of building one proprietary AI assistant and locking you into it, they built an open API and SDK that lets any agent plug into your issue tracker. That means you don't need to wait for Linear to build the features that you need. You can connect the best coding agents on the market like Cursor, GitHub Copilot, OpenAI Codex, and Devin, or you can build your own agent for your team-specific workflow. It's a fundamentally different approach from most issue tracking and project management tools on the market. You get optionality, and the experience is surprisingly natural. You assign an issue to an agent the same way you'd assign it to a teammate, or you can simply mention the agent in an issue thread. Cursor then can pick up a bug, understand the context from the issue, and open a PR. Codex can explore a fix while you're focused on something else. Sentry can do root cause analysis when something breaks. It's pretty powerful what you can get these agents to do. And here's what I like: you, the human, stay the accountable owner. The agent works for you, not instead of you. You review the work. You decide when it's good and when it ships. If agents are going to be a part of the tool set of building software, and it feels to me they increasingly are, you'll want a system that's actually designed for them. Linear is a system like this. To learn more, head to linear.app/agents.
And with this, let's get back to the point where Bryan was saying how those believing the internet will be the future back in 2001 were right.
This is the other thing. It's like they're right. And so a very famous impact crater from the dotcom boom is Webvan, right? Webvan was delivering groceries, and many people today are going to get their groceries delivered, right? Instacart. It's like they weren't wrong, but their timing was off and they lost track of the underlying economics completely. And so when it busted out, in the fall of 2000, in November of 2000 in particular, there were zero orders from telecoms at Sun. It went to zero. Wow. And you're kind of used to ups and downs, but that's just off a cliff. And from that point, going to 2000 and then 2001, it was then very, very grim. I would say that the thing that happened through the bust was layoff after layoff after layoff, because companies had kind of built themselves and geared themselves around these fat times lasting forever, and now they were gone. And as frothy as expectations were during the boom, they were that much negative in the bust. People were like, it's the end of days.
And were you a software engineer back then? Yeah, software engineer. Yep.
And then as a software engineer, both you and also thinking about your colleagues back at the time or friends, how did it impact you? Were you kind of just chugging along, or?
So I would say that lots of people left, and you had the statistic of the U-Hauls being 10 to 1 out of the Bay Area. So people moved away, and the thing that I noticed is that the people that had moved out to Silicon Valley because they really had an interest in the technology all stayed and were not adversely affected, honestly. Yes, every one of us, if you had equity in your company, which of course you all did, you try not to overthink it, right? You try to remind yourself, I never had it to begin with, so it's hard to... but it's definitely gone. Sun lost 98% of its value, so it's definitely gone. And I think also a boom can get you to care about things that you actually don't care about, because in a boom everyone is so financially driven that it's hard not to become financially
driven. But it's like that's actually not why I got into this. And so during the bust, I'm definitely able to put a meal on my table and a roof over my head. But it was really a reminder about what's important and again because we did do better technical work in the bust than we did in the boom. And I think it's because in the bust it's like, okay, now we really have to focus. We have fewer resources, and the fewer resources actually force more creativity.
So all of the things that we did, certainly speaking at Sun and system software, so ZFS and DTrace and the Service Management Facility, all these things that were really revolutionary for the operating system all happened in the same kind of post-bust period of time. So all of those things happened from 2001 to, say, 2005.
And so what were these specific innovations?
So I'd gone to work at Sun to work with Jeff Bonwick, and as long as I had known Jeff, from the mid-'90s, Jeff had wanted to rethink file systems, and now finally in the early 2000s he and Matt Ahrens were able to really go take a clean sheet of paper for the file system, and that's ZFS. I had a chip on my shoulder about the way we understand and debug systems, the way we observe systems. So I, along with two other colleagues, did DTrace, which allowed us to dynamically instrument running systems. And you can kind of go down the line and there were a bunch of things like this where
we, and I don't know that all of this is related to the bust. It's just that the timing lined up such that it was all happening during the bust, and what we ended up with was a whole bunch of interesting technology coming together actually in a single version of the operating system. And then, very fortunate for us, and I do think this is a bit of a consequence of the bust because Sun was definitely open to new approaches, we open sourced all of the operating system. So that happened in 2005 and that was very important to give these kind of technologies eternal life.
But I think, you know, we can never predict the future, but to me it is pretty positive in this sense that even in the bust, hearing the stories, innovation did not stop. Sure, it sounds like it was probably harder to get jobs and there might have been fewer of them, but industry kept innovating. And what you said that I didn't expect to hear is that it was a bit easier to innovate.
It's just less manic. We were able to focus more. And so, I mean, not that one should necessarily pine for a bust, because busts are brutal, but there is a clarity that you get too. So I mean ideally you would like to have, can we just be normal economically? But nope. Apparently in high-tech we've got to be on or off.
So bust aside, in the early 2000s leading up to this internet boom, most companies went about buying Sun servers with Solaris installed, and hardware and software came together. It was beautiful. It worked well. Again, I heard from folks who did it. What happened then? 'Cause when I got into tech in 2000, I did not hear about Solaris and that was not how it
No. Right. What was the shift?
So the shift was, first of all, open source, right? So we said in the mid-'90s Linux was kind of still very much a hobby project. Not so by the 2000s, right? So it grew up, absolutely, and it grew up because you had a bunch of companies that really backed up the truck. At first IBM and SGI, Data General, some other companies, those companies were very important because they decided to contribute their technologies, like XFS, right? XFS, which many people still use today on Linux, that's from SGI. XFS was SGI on IRIX. That was happening in kind of the late '90s.
And then in the 2000s, I mean, Google was always built on Linux, right? And so the companies that became that next boom were all built on open source and indeed needed to be built on open source. They economically relied on open source to be able to build what they built. So then it became much more practical to certainly run Linux, and I think the other BSDs, or we open sourced Solaris. So there were a lot of options that were now available. So that shifted. I think the other thing that shifted is that, I mean, SPARC bluntly lost to x86. And SPARC is a Harvard architecture.
SPARC is a microprocessor? Yeah. Because there was a time in the '90s when if you wanted the fastest microprocessor, it was a RISC microprocessor. It was a SPARC microprocessor, or it was MIPS, or it was Alpha. And x86 was a commodity, and obviously available with a personal computer, but was not faster than those RISC microprocessors. That shifted in the late '90s, and because we ran the operating system, Solaris, on both SPARC and x86, we could see how fast these x86 machines were. And you talk to the microelectronics folks, they kind of dismissed x86 and dismissed Intel, and you shouldn't do that.
And in particular, Intel was very focused and architected their way around what was called the memory wall. And they were able, in part because they used speculative execution, to actually make these microprocessors that became much faster than the RISC microprocessors. So by the time you are in, say, 2004, 2005, if you want a leading-edge microprocessor, it's x86. So that was a big and important shift.
So by the time you're coming up, it's like, okay, yeah, if I want this, I'll just, I don't know, get a Dell box or a Supermicro box and then I'll put Linux on it, or maybe FreeBSD, and away I go.
Then the next kind of big and important shift that happened started in 2006. You could argue with S3, but then especially in those next kind of
'07, '08, '09, with the introduction of EC2, and now you have the cloud that starts to come into play, and now people were like, why would I even screw out a server at all? I mean, it was so great to be able to just spin up infrastructure.
Yeah. I remember one of my early companies, mid-2000s, we had a server room. We had server administrators. The server room was always hot. And this was a small company, mind you. This was not a big one. Every company needed to do that. It's kind of amazing to think that every single company, no matter if you were a website, you had your own server room.
And if you were a dev, you wanted to be friends with the server admin because when you wanted to deploy your stuff, they could do stuff for you.
They could do stuff for you. That's it, totally. And so I think that cloud computing was really important. This is not a deep thought, that elastic infrastructure was really important, but the ability to have API-driven infrastructure. And so for me personally, I was at Sun and then in 2006 I started a storage group inside of Sun, which was great. A really successful group, but so successful that it actually attracted Oracle as a customer for the first time in a long time. This is a little bit of residual shame that I have: did I attract the marine apex predator that ate the company?
'Cause Oracle later bought Sun, right? And then they bought Sun, and that closed in early 2010. I left shortly thereafter because I could see what Oracle was.
Well, I never heard a story of your potential role here.
That's right. So, yeah. And Oracle, and maybe a year later I gave a talk in 2011 with some rather unvarnished opinions about Oracle and Larry Ellison in particular. I cautioned people about anthropomorphizing Larry Ellison. You have to treat Larry Ellison as a machine, like a lawn mower. You stick your hand in the lawn mower, it'll chop it off.
All right, so I'm giving this talk in 2011. Again, this is after I've left what was then Oracle, and I was just saying things that I felt were obvious, but the audience is kind of gasping, and people are coming up to you after the talk like, do you think there's going to be retribution from Oracle? No, you're misunderstanding. The lawn mower is not angry at you. It's a machine. It doesn't have the mirror neurons to be, it would almost show me that I'm wrong for Oracle to resent what I'm saying about the... Anyway, all the videos for that conference go up and my video does not go up.
Oh, right. Okay. And so my colleagues were like, "This is an Oracle conspiracy." I'm like, "This is not an Oracle conspiracy," which it wasn't. It wasn't orchestrated by Oracle, but what I underestimated was the fear of the conference organizers. So they themselves were terrified of offending Oracle.
Yes. Even though it probably would have been fine.
No. So the talk did finally go up. Before the talk starts, there is a disclaimer: the views in this talk do not represent the views of the USENIX Association. And you're like, "All right, I get it. Never seen this disclaimer before, but fine." Then during the talk, the format of the talk is you've got a slide and then you've got a little blank strip and then you've got this talking head in the lower right corner. There's kind of this dead space above the speaker. They took this disclaimer and they rejustified it and they put it above my head the entire time I'm speaking.
And maybe in this regard they were prescient, because to this day, if Ellison is mentioned on Hacker News or Oracle is mentioned on Hacker News, someone will immediately cite minute 33 of this talk, which is when I go on this kind of Oracle... again, I don't view it as a rant. I view it as just me describing what is obviously true, that we all know. But anyway, I left Oracle after they bought
So we're now around like 2010. So cloud has taken off. x86 architecture is everywhere. Linux is now winning both for small-time servers but also on the cloud. And then what happens? This was an interesting time when Google started to figure out that, hey, they could do something interesting on their cloud, right?
Yeah, that's right. So this is still a little bit before that. So this is kind of from, I would say, from 2010 to about 2014, a period of relentless execution from AWS. AWS is executing so extremely well. There are not really other public cloud options. There's kind of Azure drifting out there.
I think people forget that GCP on paper has been around from 2009, but up to like 2014 it was almost like a joke.
It was a joke. I would say it existed but it was a joke. And in particular, at every single re:Invent, Amazon would announce a new price cut. And if you were a competitor to AWS, you were dreading re:Invent because here comes another price cut. If you were a partner of AWS, you're dreading re:Invent because here comes the announcement of a new service that competes with what you're making.
I think people who have not been around have forgotten, but it really has happened, 'cause it's not been the norm the last, let's say, 5, 10 years or so.
Well, and in particular, they did a couple things where it's just like, man, you've got to tip your hat. I mean, Jeff Bezos is the apex predator of capitalism. Larry Ellison may be the lawn mower, but Bezos is ultimately the apex predator, because the thing that was so impressive is they were able to give people the idea that this was a terrible business. In particular, they did not break out their financials. So everyone's like, "Oh my god, what an awful business." They're cutting the price every year. You do not want to, this is a classic red ocean. It's bloody. You don't want to compete. And so we were at Joyent. We were actually competing head-to-head with AWS. So you were offering
A public cloud. So we ran a public cloud, and then, unlike AWS, we took the software that we had used to run the public cloud and made it available for people that wanted to run a cloud on-prem on their own hardware. So people that would buy Dell or HP or Supermicro, they would buy our software and they would run it on there and get a cloud.
So we ran a public cloud and we knew what the economics of a public cloud were. Namely, pretty good. Margins were good. And so what we knew, that Amazon wasn't volunteering, is that AWS S3 was underwriting a war on big-box retail. S3 was paying for your Prime shipping. It was a genius move. And so
Also some insider information that you had because you did your own thing.
Well, we know that the margins are very good. And then of course, you will be unsurprised to learn that several of Joyent's most prominent customers were retailers. This was not lost on retailers. Retailers are like, "Gee, I wonder what's happening." Retailers are like, "If you think I'm going to take my dollars and spend them on AWS so Amazon can go to war with me, no thank you."
There was a period of time when I felt like in order to be in the cloud, you had to implement every AWS API. So there was this idea that you had to be API compatible with EC2. There's a company called Eucalyptus that tried to do this. It was just a disaster. And part of the reason it was thought that GCP and Azure could never compete with AWS is because they could never be API compatible. And so I am convinced that, because what changes in like 2015? What starts in 2015? Kubernetes. And I think that part of that initial attraction to Kubernetes is that people wanted to get some optionality around their cloud, and they felt locked into AWS. They're like, I'm not using all this stuff. I'm not using Elastic Beanstalk. I'm not using Greengrass. I'm not using Redshift. What I actually want is this kind of basic infrastructure, and Kubernetes now gives me this layer upon which I can deploy and get some sort of true cloud neutrality.
So multicloud didn't really exist, I would say, before Kubernetes, and I think a lot of that, especially early momentum behind Kubernetes, is around this idea of, I need to have some optionality in here. I want to actually be able to go to GCP. So I think, and I think it's giving Google slightly too much credit, but only slightly too much credit, to say it is a masterstroke.
On the podcast I had Kat Cosgrove, who was a release project manager on Kubernetes, and she's been in the project for a long time. She was never a Google employee, but I asked her, why do you think Google open sourced Kubernetes? They have Borg, which is amazing, and they kind of built honestly a better version for external use, and they just released it just like that. They put a lot of work into it, and to me it didn't really compute, like why would Google, what is the business reason? And she told me that she thought, again, speculation from the outside, that they probably thought that it would help Google Cloud
That's right.
to have a container which is now portable, and now you can give the promise that if you run this on Azure, especially AWS, you could come over. So it kind of makes sense. Is this your thinking?
Yeah, absolutely. But I think that is definitely the argument that Kubernetes proponents would make inside of Google
in terms of why they did it. Nobody prevented it, you know what I mean? They kind of open sourced it.
Google was a pretty cool place in the sense that it was very bottoms-up, as I understand, back then still.
Yeah. And then I think part of their
You know, it was Craig McLuckie who really pushed for the CNCF, the formation of the CNCF around Kubernetes, to give it kind of a foundation home. I do remember one conversation with Craig, as he's contemplating the CNCF, and he's like, well, I think this is going to allow Kubernetes to get the marketing dollars that it needs. I'm like, don't you work for the most profitable company on earth? Like, isn't it just gushing cash over there, and you can't get a couple million bucks for marketing for this thing? But no, apparently you can't.
So I think that the argument that people were making internally was about, we should be encouraging cloud neutrality because we are the ones that have something to win, and they're right. And they did, and GCP is now not an afterthought. GCP is very important. It's a very big business, and I think that they've got Kubernetes to thank? Solely, no, but I think it's played an important role for sure.
And where are we today in terms of the hardware and the software stack running, specifically thinking of these big clouds? What's happening inside the likes of Meta, these giants? As I understand, they're no longer just ordering servers from Dell or wherever.
Never were. Never were.
So it's kind of funny because all of these folks took a somewhat similar path. They never were, because in Google's earliest days they were assembling machines from Fry's, RIP Fry's, Fry's being an iconic electronics shop that has long since disappeared, but they were kind of famously velcroing machines together and finding
So they bought the processor, the different networking switch, whatever.
And they had this idea that it doesn't matter what junk we run on because our software is going to run as a distributed system. It actually doesn't matter. We don't need ECC-protected memory because it doesn't matter if your DIMMs fail. And so I think they learned, well, it does matter a little bit if your DIMMs have rampant data corruption. DIMMs failing, that's actually not a problem. Your memory returning the wrong thing, that is a problem. Next thing you know, your software inserts that as a row into a database, and now you've got
Yeah, correctness is a problem.
Yeah. Correctness is a problem. It's like, okay, overshot the mark. So by the time they're like, okay, we're not going to velcro machines together, by that point in time the business was established enough that they actually built the machines that were fit for scale. So they have a great book that was written in kind of the mid-2000s, the warehouse-scale computer, where they talk about all the things they did, DC bus bar, really thinking about power across the entire DC. So they went from being kind of too cheap for Dell or even Super Micro to then being much better engineered than those systems ever were. So they were never really meaningful customers. And ditto for Facebook, Meta. They were never really meaningful. They kicked them out very early and did their own stuff.
Bryan just talked about how Facebook built their own servers because off-the-shelf solutions didn't work at their scale. And what's interesting is that companies like Meta and Google didn't just build better hardware. They also built incredible internal tools. Tools for safe deployments, feature flagging, experimentation, debugging, analytics, the whole stack that lets teams ship fast and with confidence. Most companies never get access to this level of infrastructure. You either build it yourself, which takes years and large engineering teams, or you make do with scattered tools that don't talk to each other. That's exactly where Statsig comes in.
Statsig is our presenting partner for the season, and they give every engineering team access to the kind of tooling that only the biggest tech companies used to have internally. At its core, Statsig is a toolkit for safer deployments and experimentation. You ship a new feature to 10% of users behind a feature gate. You validate that it behaves correctly, watch the metrics, and expand to the remaining 90% only when you're confident. And if something goes wrong, you can turn it off instantly, long before it affects everyone. And safe deployments require visibility. Statsig includes analytics, both product analytics and infrastructure analytics. So you can actually see what your code is doing in production: errors, performance changes, funnels, user behavior, because you cannot ship safely if you can't see what's happening.
Companies like Microsoft and Notion run hundreds of experiments per quarter with Statsig, velocity that used to require entire platform teams to build and maintain. This used to be infrastructure available to maybe 10 or 15 tech giants. Now startups and mid-size teams use Statsig to ship quickly without breaking things. If you want to give your engineering team world-class tooling from day one, go to statsig.com/pragmatic. There's a generous free tier, a $50,000 starter program, and affordable enterprise plans. And now let's get back to the conversation about the history of computing and what might be coming next.
And this was independent. So both Google and Meta came to the conclusion of, we should just build our own stuff.
And Microsoft and Amazon all came to the independent conclusion, because the scale at which they needed to run was not at all the scale at which Super Micro and Dell and HP were geared. What they were geared to do was to run the servers in your server room where you needed to know the devs, right? Where it's like, I'm going to have a little rack. It's going to have six servers. Then maybe it's got 12 servers. Okay, maybe we grow to 24 servers. That's what they were designed to do. If you're like, "No, I want to buy servers by the thousands because I've got a public cloud business," if you want to buy servers by the thousands, there is no product from those companies for you. And in very, very basic ways, like the DC bus bar, at every juncture they've been designed to be a personal computer, and you happen to be slapping many personal computers together, but they're not designed to actually run infrastructure at scale.
So that was happening inside effectively all the hyperscalers. And Joyent, meanwhile, was bought by Samsung in 2016. Joyent was bought by Samsung because their cloud bill was off the charts, and
They bought you to
Bring it in house.
Yeah. And there was not a product they could go buy. So they went to go buy a company.
So you're like, "Wow." And it's like, "Wow, that's a big AWS bill." It's like, yes, very big AWS bill. But then that was not a product or company that was available for the next Samsung. What does the next Samsung do? Well, that's one less company available to buy. So when we were contemplating the next thing in 2019, one of the things that we had seen, and we earnestly believed: one, cloud computing is the future of all computing. Not a deep thought that elastic infrastructure, API-driven infrastructure, that is modernity. Two, you shouldn't be able to only rent that. You should be able to buy that, own it, run it in your own data center. Why would you want to do that? Well, you might want to do that for risk management, for security, or for economics, because if you're at a certain scale, you'd rather own it than rent it.
And I think before Oxide, or in 2019 or even in 2020, 2021, if you were a midsize company, not big enough to build out your own custom cloud and build everything that the hyperscalers did, you could buy some off-the-shelf HP or Dell, a bunch of them. I think that's what Basecamp did. I think they posted that they bought a bunch of these things. They rented space in one of these shared, or I think two different locations. They put in their boxes with all the memory, and then they kind of set it up and put it together. So I guess those were the two options, right?
Yeah, those are the two options, and I think that Basecamp ended up being a real poster child for the economic advantage, because DHH is obviously outspoken, and the economic advantage was really, really, really clear. They're also at a scale which is not the scale that we're targeting, right? The scale we're looking at is a much larger scale. And so the economic argument is actually even more compelling when you're at that larger scale. I love it when the VCs that passed on us because they felt there was no market would then send me the DHH blog post. It's like, why are you sending this to me? I should be sending this to you. I know this. We just knew the economics of it, and we couldn't predict exactly what the trends would look like, but believed that there would be folks that were born on the public cloud that would outgrow the economics of the public cloud and want to go on-prem.
Economics aside, what does it take to build one of these things? I saw one of these things. We'll put in a picture of it. It's a proper, like, my 9 ft tall rack. It's big. It feels like you're putting, I don't know, 16 or 32 of those Dell things in terms of size, just to get a sense.
Yeah. We have 32 compute sleds in there. That's right.
And what did it take to actually build it? What did you need to design in terms of hardware and then software?
Yeah. So, and we knew this too going into the company, we knew we were taking a clean sheet of paper, right? And so we were deliberately like, no, we're going to start with a problem. We're not going to build it out of Dell, HP, Super Micro. We're going to start with a problem, and how do you best solve the problem? And as it turns out, there's a lot of technical debt that had been accrued by this kind of PC ecosystem. So, God, where do you start? Just on the environmentals, like on power, right? The fact that you've got AC power in each of these Dell, HP, Super Micro.
Yeah. So if you put 16, you have 16 separate AC
Times two, because you have two power supplies per 1U/2U chassis. Two power supplies. By the way, there are two fans sitting on those power supplies, and those fans are actually what wear out. In terms of the whirring fans, it's not just coming from the computer, it's coming from the power supplies, because those power supplies are dense. They're packed with stuff. So they've got to overcome a huge amount of static pressure. So that's not the way anyone does it at scale. The way people do it at scale is you've got a DC bus bar, you've got a power shelf that is much more efficient
And that rectifies from AC to DC, and then you run DC up and down, and then you blind mate into that. So we knew we were going to do that.
That's a little electronics engineering right there.
Yeah. The power engineering for sure, and we knew we were going to do that. We also knew that by taking a clean sheet of paper, we would have opportunity made available to us that we weren't necessarily thinking of, and that manifested pretty early. So we blind mate into power, which is to say that when you feed a sled in, that power connector, you don't see it, it's at the back. You lock the sled in, it blind mates into power. And we had assumed that we were going to do what Facebook and Google and others, Amazon, have done, and have networking out the front in the cold aisle. But as we were taking a clean sheet of paper, talking to some connectivity vendors, they asked us, wait a minute, you guys are taking a clean sheet of paper, why are you putting cabling in the front? Why wouldn't you also blind mate the networking connection? And we were like, can you do that? They're like, oh, you can definitely do that. Well, why don't the hyperscalers do that? It's like, oh, they would all tell you that if they could start over today, they would blind mate the networking, and they're just too afraid to do it at this point. Which was like catnip for us, you know, like they're too afraid to do it? Okay, we've got to. And one of the very early, holy God, we're going to bet the company decisions was blind-mating networking, because if blind-mating networking doesn't work, you've got nothing. You don't have a problem.
And so what is the difference in blind-mating networking versus
It means there is no cabling in the system at all. So when you've got a sled, you are blind mating into a cabled backplane. It's cabled in the factory. So the operator
So when the box comes in, that's why I didn't see any cables. It runs inside.
It runs down the back. And so
Versus when I look at the pictures of a data center of, let's say, Google, you see they're very neatly organized. I love organization, so it's beautiful, but it's cables everywhere, and you can see.
So you don't have that.
We don't have that. And in particular, because there's no cabling, there's also no miscabling, right? Every computer is not actually on just one network. It actually needs to be on three. It's on a presence detect network. It is on a service processor network. And then it's on that high-speed network that you really care about, the actual network. In any facility, you need another network for power, environmentals, and so on. It's very easy to have miscabling; that's got to go to a different router. There's a bunch of just complexity that we eliminate because we do. And then part of that decision came out of an arguably earlier bet the company decision, which was we did our own switch. So in addition to doing our own compute sled, we did our own switch.
And last time you told me about this, and in our deep dive we did a little bit, at first you said we did our own switch, and I was like, yeah, okay, cool, you did your own switch. And then you told me that actually that is a second computer to build. Can you tell me why? And it's funny because when we went through Sand Hill initially raising money for the company, nobody asked us.
Sand Hill Road. Exactly. And we were definitely, so people would be like, I've got a technical question for you, and you're like, oh God, here comes the switch question. But then no, some other random question, like, all right, that's not a very good question. But nobody was asking us about the switch. And we were concerned about the switch because we'd already come to the conclusion that in order to make this thing really work, we had to do our own switch. And the reason we had to do our own switch: if we didn't do our own switch, it would be a third-party integration nightmare, and we wouldn't be able to actually solve the problem that we're trying to solve, which is when this thing shows up in your data center, we want this thing to come out of the crate. We want you to wheel it up. We want you to put in power and networking and go. We do not want you to have to cable anything. The level of operator involvement should be really minimal. So we'd already come to the conclusion that in order to make this thing operable and manageable, we needed to do our own switch.
And so you're saying that buying, because a switch to me sounds like a somewhat simple component, and you're going to tell me why it's not.
Oh yeah, it's definitely not. No, but that attitude is very important. If you want to go build your own switch, I encourage you to have that attitude as long as you possibly can, because otherwise you won't go do it.
So what is your switch, being obviously the networking switch? What does your networking switch do, or what made it so important for you to build it as opposed to going to one of the many suppliers and saying, you know, let's get your—
Not many suppliers. So if you actually go to the actual switching silicon, it's coming from like one and a half providers.
It's all Broadcom, and so what you're actually talking about is Broadcom silicon. What we discovered is this actually interesting piece of actually Intel silicon from a company they had bought called Barefoot, and we found Intel Tofino, which allowed us to have true programmable networking. So we use Intel Tofino. Intel later killed Tofino. So, complicated relationship with Intel over this. We fortunately have procured enough Tofino to be able to take — we bought ourselves the time we need to kind of design our next-gen switch. But that programmability was very, very important for us, and that we were not going to get from Broadcom. Broadcom is a very proprietary company. We were not going to get a bunch of the things that we needed in building that switch; we were not going to get them out of Broadcom. So it ended up being very important.
We were concerned. I mean, again, another one of these kind of bet-the-company decisions. Very, very concerned about having our own switch, integrating our own. And what we found is that was a win in so many dimensions. So many dimensions that we did not anticipate. And now you can't imagine the company without having it.
Sometimes you do stuff and you might get some wins.
Absolutely. Well, I think also, whenever you're deliberating something big like that, the fact that it is big kind of forces you to really deliberate. And then once you commit to taking that big risk, you often see unexpected dividends. Like, well, as long as we're going to do this, as long as we are taking a clean sheet of paper, as long as we're doing our own switch, we can blind-mate the networking. If we were not doing our own switch, we really couldn't blind-mate the networking. We really needed to be able to own both sides of that in order to be able to do our own switch or blind-mate.
A lot of us, you know, listeners, viewers, are software engineers, so we don't know as much about hardware. Obviously we know how the things work, but can you tell me a bit on what it actually means to design or build a computer? Because, you know, I'll give you the novice approach, which is obviously going to be wrong. But the novice approach is like, oh, here's a processor, here's a few chips, here's a mainboard, I'll just put it on there and I'm done. But when I was in your lab at Oxide, you told me that one of the first engineers turned out to be a radio frequency engineer. You told me how this is great because of all the FDA approvals and all these things, and I was like, okay, this is way more involved than I ever imagined.
Yeah, it's very involved.
How do you build a new computer?
First of all, I mean, it would be a lot easier if it were all slower, right? The problem is it's very fast. It's high speed. So the connection to memory via now DDR5, double data rate memory 5, is ridiculously high throughput, and from a signal integrity perspective really complicated. These boards, by the way — ultimately this is all analog. We think of it as digital, and it is digital, but digital is like a lie that double Es allow us to tell ourselves. You are actually talking about signals that are racing through a substrate, and with PCIe or DDR5, those signals are very complicated to lay out. That's complicated.
The actual, like, how does the computer start? This computer is like a triple seven, right? Or, you know, a 747 used to be my favorite jet to kind of pick on, but now the 747 is retired, so I've got to pick something else, and I'm not going to pick another boring aircraft. An A380, I guess, right? I should pick an Airbus. But you think about it: okay, an Airbus doesn't just come by itself. It needs an airport. It needs a runway. It needs all the infrastructure to feed it. Well, so too for a microprocessor. Just the power sequencing for those things is very complicated. It needs another surround that manages the power distribution network, that actually manages its power-on sequencing, that manages all of its environmentals, that manages its connection to memory, to IO. So it is just fractally complicated, to the point that people often just take reference designs and iterate on them. They don't actually really innovate on this stuff because it takes so long.
And you told me this, which was really interesting last time: as I understand, reference design means — correct me if I'm wrong — that you're an electronics engineer or hardware engineer and you want to build new hardware, and you take an existing reference that has been tested, measured out, so it doesn't create accidentally all sorts of radio frequency things, and then you implement that. But you told me that this is not what you did. You also told me that it's pretty hard to find electronics engineers who are used to not doing reference design, but who are brave enough to—
Who are brave. Yes. I would say that in computer design in particular, the high-speed designs are so hard, people got very accustomed to taking the reference designs, and it was harder to find folks that were willing to take a clean sheet of paper. And we ultimately found them. I mean, we've got a double E team that is extraordinary—
And double E is electronics engineer, right?
Yeah, and absolutely fearless. And in part because they didn't spend their careers at Dell and HPE. No, they're coming from like GE Medical, where they worked on CT systems.
Wow. How did that happen?
How did they come to Oxide?
It's not — but it feels like such a different field. I would have assumed naively that, you know, if you're building a computer, you'll try to get electronics engineers who have built computers.
You would think. And that was probably our thought as well, and then we discovered that we were not getting along with those engineers. Well, we didn't hire them, but we were just finding there's a lot of friction because there wasn't a real first-principles approach from those folks. And this is where you get to talk to folks that have been at Dell for a generation, and for any design they're used to calling what's called the FAE, which is the field applications engineer for, you know, the voltage regulator. It's like, well, the FAE gives me the design. It's like, all right, well, how do you know that it's the right design? Well, no, he's there. So it's like, all right, so let's go hire that person then, let's forget you.
And we were really struggling. I was struggling to get outside of my own personal network to find the right engineers. And we were kind of brainstorming, like, how can we get people to see the company who wouldn't otherwise see it?
And specifically for hardware engineers, like we're talking about.
Yeah. And just in general, but for EEs specifically, yeah, it was feeling especially acute. One of the things — we were kind of brainstorming as a team, and one of our engineers said, you know, the values are very important to us at Oxide, which they are, and I relay Oxide's values and our principles to people outside of Oxide and they're like, that's just [bleep]. And I explain that, like, you know, normally I would agree with you, but it's when I get to the compensation that people's heads turn, because our compensation is transparent and uniform, and people are like, "Wait, what?" And I'm like, "I could write a blog entry on it." Like, "Yeah, that would be great." I'm like, "Okay."
And so, up until that point, we had not talked about it at all. We had not talked about it publicly at all. I just came up with the idea that compensation is just private. It's just not something you talk about with people, you know.
And you go to Levels.fyi or some of the forums where you're anonymously asking, people are anonymously sharing that. That's how you get information.
That's how you get information. And so I kind of had this idea that it just is not something that you — and so we wrote this blog entry in March of 2021, and it sent our hiring nonlinear. And it wasn't that people were like, oh my god, I want to work for a company where everyone's paid the same, like that's like—
Yeah, because your compensation was both the same and you also put the number specifically. I think it was something like $200,000 back then.
Yeah, it was a little bit less back then, but now a bit more than that. Now we just got another raise, so now I've lost track. It was 207, but now it's more than that. I actually don't know, because the one thing is, when compensation is uniform, you don't keep total track of it. Literally people were like, wait a minute, there's an error in my paycheck, I just got paid more. People like, "No, no, we got a raise." Like, "When was that?" Like, "No, it was at the last all-hands." Like, "Oh, you know, I did have to go to the bathroom at the end of the last all-hands. I didn't listen to the recording. I guess I missed my raise." Like, "Yeah, yeah, you've got to pay attention around here."
But it was more that what drew attention was that people — engineers in particular, but just in general — people were drawn to a company that would be so nuts as to do that. And ultimately that engineer that made the suggestion was absolutely right. It was the compensation that convinced people that we take our values really seriously, that we're a really principled company.
Which is, you're paying everyone the same base salary. Exactly the same. Yeah. They're making the same as you, the electronics engineer, software engineer, whatever other role you might have.
That's right. And I don't know if you should just go ahead and say it if you want to, but many people are like, would you pay support engineers the same amount? It's like, why do people always pick on support? They would ask.
Exactly. The answer to that is yes, and the answer to that is, if you do that, you find supportive support engineers. And so we have got, I think, the best support engineers in the business. We've got really, really phenomenal folks in support.
I heard a small company called Gumroad do this, where they paid their support staff really high, again about the same as software engineers, and then they got support staff who were software engineers, and they could fix the code or write tools for themselves.
And you get people for whom — because, I mean, you know, there's a certain thrill in being in support, because you've got someone with a problem. It's technical. You get to be technical, get to go solve a hard problem, and then immediately you get such gratitude, you know, and that's a rush. And there are people that are really drawn to that: I love helping other people, I love that feeling that I get when I resolve a problem for someone, that immediacy. So one of the things that we've heard repeatedly from several of our support engineers is, my heart was always in support, but my career path was forcing me into a different career path, and I love the fact that I get back to where my heart is.
Yeah, that's nice, because now you're not going to make more by doing something that you're not as into. I love that. So, going back to where we were, which is you build the hardware, you build this really complicated piece, and you went through electronics engineering, putting it together. Let's put the software on, because that's super exciting. What does it take to build software for this? Let's talk from the low level. Did you start from scratch from the operating system? Did you have to, or could you use—
There's kind of different answers at different levels of the stack. So on our service processor, we did start from scratch. We did our own de novo operating system in Rust, appropriately called Hubris, because we had the hubris to do it. The debugger, by the way, for Hubris is called Humility, which feels appropriate for a debugger. So that was de novo.
And this is open source, right?
Open source. Yeah, the entire stack is open source. Everything we've done is open source.
We can go on GitHub and check it out.
You know, go on GitHub and check it out. And yeah, I mean, we've got God's own revenue model, because you're like, well, what if somebody can download it, run it on a different computer? It's like, knock yourself out, because, you know, we think the best way to run this is on the machines that we make, and those are not free. An Oxide machine is not free, downloadable, but it's all open source.
So that was for the service processor. For the host CPU, we really were kind of at a quandary, like, what do we want to do on the host CPU? And that is to say, on the actual, what was then AMD Milan, now AMD Turin silicon, we knew that in the product we would do our own hypervisor and our own control plane. So this is not something that you run—
The control plane — is that controlling the whole thing? Like, you have a bunch of processors and memory and all that, and the control plane controls all that.
You plug this thing in, you power it on, you put in networking. What you get is a console that looks a lot like what AWS would look like if AWS looked better. I mean, it's a console. I mean, look, not to disparage AWS, but we know that design is not really the strong suit.
We agree with that.
Yeah. Exactly. So it looks gorgeous, of course. And you've also got your API, you've got your CLI, and you're provisioning instances. Where are those instances provisioned? It's the control plane that makes those decisions. You are attaching virtual storage to those instances. Where does that storage live? It's the control plane that makes that decision. So just like with AWS, you don't need to know that stuff. That's just happening. You're using Terraform to spin up your cluster. You're running Kubernetes on it. You're knocking yourself out.
So we are delivering all of the software, from that lowest layer, that service processor, to the operating system that's running on the host CPU, and then, very importantly, that distributed system, which we called Omicron before the Omicron variant of COVID, which was feeling very ill-timed. For a very brief period of time it was feeling ill-timed, and now I feel like the Omicron variant of COVID has just been forgotten, and now it's a good name again. So it's like, you know, we just—
It was really short-lived.
Yeah. So we lived longer than the Omicron variant of COVID. And that is our control plane, and that is a very sophisticated body of software. Because it's not enough to just provision an instance, right? You need to do that robustly, you need to do that via API, CLI, and so on. But then all the software that does that and keeps track of your instance and so on — it's very important that you can actually update that software, that whole distributed system. You need to be able to update to a new version of the software, and this gets really thorny, right? Because in a public cloud, you do that with a runbook, right? I mean, even — we don't feature it prominently, but even in GCP and AWS, yes, there's a lot of automation, but there are also humans involved, and there are humans that are
taking the responsibility for actually updating software. For sure. Really? Yeah. I mean, again, for the most part.
Yeah. I mean, there's a lot of automation involved, but in particular, if something goes wrong in an update, you know, you've got DevOps that can hop in and figure out what's going on and get it rectified. We are shipping a distributed system across an air gap in an Oxide rack that's potentially running in a secure facility. We cannot be there if it goes wrong. So, we need when
Especially because a lot of your customers are buying it because they want to do it themselves.
They want to do it themselves. So in many ways the thorniest software problem for us, we had actually several thorny problems, couldn't pick between them because they're all thorny for different reasons. One of the very, very thorny problems was how do we ship a distributed system that we can then update, and one of the things we did that was important was like, okay, because it's very easy to paint a roadmap that is very complicated for update. You'll never ship anything. So what we needed to ship in that first product that we shipped when you were back in Emeryville 2 years ago, we needed the minimum viable update. We needed an update where the software could be updated even if it was painful.
So what we did is we have this thing called mupdate, which is the minimum update, and mupdate in particular required the control plane to be parked. So we're going to take this rack that's running instances, take it offline, we're going to update it and then bring it back online. And that was robust. It was great and we got that working. That's great. That is great, and that you can update it. But that's actually not what you want in a cloud, right? You're like, I, sorry, I'm like using this thing 24/7. Like I actually, I want to, these instances need to remain up while I update it.
But that gave us the platform to go build that update functionality into the software. Extraordinarily sophisticated and really an extraordinary body of work. And actually just recently we had at our internal meetup the engineer who led the charge on that, Dave Pacheco, gave a presentation on looking back at two years of update. And I got to tell you, I think this is one of the best single talks on software you'll ever see.
And we will link this, but can you give me just a short overview of like why this update is so difficult? Because like some listeners will be used to just building applications, for example on the iPhone, and an update there, what it means, obviously I know this is way more complicated, but an update is there's a new binary version and it replaces the old binary version. Now, of course, you know, you're saying this is an operating system update or, you know, like with a car, and of course you might think like, well, you know, you could just replace the old version with the new version and there's some downtime, but where is the complexity that actually like puts all this thorn? Because I'm sensing this is like I am missing something, something very obvious.
So, because it's a distributed system. When you've got an app on an iPhone, it's not a distributed system.
Oh, and distributed system, meaning that you've got a bunch of different nodes,
components that are going to speak to one another. And it's like
those might need updating as well.
Oh, they definitely need updating.
Oh, they all need updating.
Yeah, the whole thing needs to be updated. You've got to be able to update all of the software in the rack.
Oh,
this is not just updating the operating system. This is updating absolutely everything.
So, you might need to update some parts or all parts.
You need to update the service processor, the root of trust, the drive firmware, the host operating system, and then all of the components that speak to one another.
Okay. And then it's like, okay, so I mean this challenge is fractally complicated. I mean one of the very basic ways it's complicated is like, so when we're updating, we are moving the system from one version to another version. In between, it's going to kind of be in both versions. Like what does that mean, to have the system that's operable while you've got some new components and some old components? What if you change your database schema from one version to the next version, which we definitely have? Like you have to have a method of doing that. What if you, and for every one of these components, how is it updatable? How do we got to reason about the system when it's in this hybrid state, and then it needs to be done in a way that's very, very robust.
So first and foremost we had to develop the foundation that allowed us to do this absolutely robustly. And so the way Dave and team did this is, you know, with that foundation and then very slowly lighting up different aspects of the system and making it more and more automatic over time, and, you know, first started running that on what we call our dogfood rack and did our first automatic update on the dogfood rack. It was a really great feeling for that team because this has been a very long software road and it has been one that has been very deliberate. And ultimately, like, and you know, full credit to Dave and team, took us about the amount of time that we thought it would, which is kind of very rare for software, because I think software is so fractally complicated, but that's only because they've been very carefully managing scope versus schedule, and because quality has got to be the constraint, and Dave's talk goes into that in detail in a way that I think is just extraordinary.
So I'd like to talk about the topic that is, you know, on a lot of people's minds, which is AI specifically and AI tools.
Yeah.
How have AI tools changed how you're working at Oxide specifically? Think about software engineering, maybe even hardware. Are you using these tools? Are you experimenting with them?
For sure. We've been early on in terms of using them, and yeah. I mean, you use them for different, and people are using them in different ways. I mean, no part of the Oxide stack is vibe coded. I think that is safe to say, but we are using it, and again, different people are using it different ways. We are, you know, using it to do things that are tedious. We're using it to generate test cases, you know, generate the. I use it because I think the thing that it is just unmatched at is document comprehension. We've got a very writing-intensive culture. We've got a lot of documents. It is great.
You always had that.
Yeah, always had that, and if you've got a writing-intensive culture, you're LLM-ready, not to generate those documents but to consume them.
And, you know, one of the things that I've always wanted to do, and it's still like, now it's possible, I haven't quite found the time to do it. Early on I wanted to make an RFD glossary. So RFDs are requests for discussion. We've got a lot of technical terms. I wanted to make a glossary. I tried to do that for like 3 hours, this is like in 2020, and I'm like, this spreads to the horizon. Just making a glossary is so complicated. A glossary is something that an LLM could just turn out. And so there are lots of things that we're doing to use LLMs.
It is clearly a very real, very, very big shift in lots of different aspects of software engineering. I think that, you know, but of course there are people that are being kind of reductive about it. I am definitely not a doomer. There are a lot of doomers that are out there, and, you know, I tried to give this talk about building the Oxide rack itself, and in particular the problems that we had along the way that an LLM was never going to be of any assistance on. And the title of the talk was "Intelligence is not enough," and one of the prominent doomers actually did a reaction video to my talk. It's like the only time I've ever had someone, and my daughter, who was then like 11, just thought it was hilarious that someone had held their own time in such low regard that they would spend it recording a reaction video to my talk. And so she was like, I want to watch this. I'm like, oh god, I do not want to watch this again.
Ultimately, the thing that was really frustrating is this person obviously disagrees with what I was saying, but then when I was giving these very concrete examples of here are the specific technical problems that required more than intelligence to resolve, that an LLM was not going to be able to resolve, he literally fast-forwarded through those parts. He's like, we just don't need this. You're like, bro, this is the talk. Like you can't do this. Like you're fast-forwarding over the actual meat of the talk.
Can you give an example of like a problem which you felt was this, like even, you know, if we fast forward to like
the arbitrary future. Yeah. Yeah. So yeah, super simple. I mean, we've had many, many scary problems, but we had the CPU when we did our first bring-up of our first machine.
And what does a bring-up mean?
A bring-up means taking a board and powering it up and trying to get it to work for the first time.
I think you mentioned that the term smoke test comes from electronics engineers.
Oh, I mean a smoke test, I always think of a smoke test more from aerodynamic, aeronautical engineers, but yes, I mean smoke is definitely a possibility. That's very bad. You do not want smoke, that is bad. No smoke, please, in bring-up.
So the bring-up
But we are doing bring-up and we are unable to get the CPU out of reset, and after 1.25 seconds the CPU resets itself. What's going on? Is the power network bad? And like when you have something like that happen, it's like, well, what's happening? I mean, it's just not working. I mean, what do you tell your LLM? It's not working. I mean, they can maybe give you some suggestions, but in this case it wouldn't. So we are going deep into understanding, like, maybe the power network is marginal. No, no, we resolved that. And actually we were working with AMD at the time and AMD's like, "No, these power numbers are amazing. Your margin is very good."
You're measuring it out. You're like eliminating that one.
We're eliminating that one. You're going through eliminating, eliminating, eliminating. And this was weeks, and you're like, we don't have a company, like we're dead. We are absolutely dead.
And I feel like this is the kind of thing that, you know, you get desperate. You're like, we're going to try kind of anything. And the engineer who was working on this actually looked at the protocol between the CPU and the voltage regulator. So there's a protocol that goes back and forth that says, hey, I need this voltage, and, you know, this is the voltage, and one of the things he notices is that there is no acknowledgement packet from the regulator. So the CPU asks for a voltage to be set to a certain level, and he's noticing that there's no acknowledgement packet back from the regulator,
which should come
which should come, and the test that they've got, something called SDLE, which is this great test goober that you, you take the CPU off, you put on the SDLE and it will measure the power for you. Well, the SDLE didn't care whether it got an acknowledgement packet or not. The CPU definitely did. So the CPU says, I want you to go to 0.9 volts. It never gets an acknowledgement back. And meanwhile it's sitting at 0.9 volts, and it's just like, well, I never got an acknowledgement, so we're going to reset and I'll do it again. And that was due to a firmware bug on the Renesas controller. And so we got a firmware update from Renesas and done. And I mean, to be fair, the Renesas FAE was great. He was like, well, you guys really should have reached out a lot sooner. Like, yeah, I know. We really wanted to make sure that we got everything. And that's the kind of problem. And there were many, many problems like this where it's not merely intelligence. Building a board is not an IQ test. I mean, you need to be intelligent to do it, but intelligence is not enough. You need these other kinds of characteristics.
Then I feel you also need a team in this case, right?
You absolutely need a team. 100%, you need a team.
Like you're going to solve these problems with, you know, you had that engineer who just thought of measuring this out,
Right. Well, an engineer who was desperate, you know, because we were all getting desperate. And, you know, again, we've had many of these over the history of the company. And you're right, you absolutely need a team. And you see also the value when you have a team. People have different ways of approaching a problem. That diversity is really important, because actually this has happened more than once with the company, where somebody is just kind of walking through the problem, and someone's like, hey, I'm just joining, you know, at a remote company anyone joins the Google Meet, I'm just joining because, you know, I'm following along. And someone will be like, hey, I've got like a dumb question. Are those virtual addresses? Those look like similar virtual addresses. And you need someone to kind of come and make that observation that is maybe less grounded in it, and people are like, oh wait a minute, that's actually something to go check. And so you need that different kind of approach that a team kind of uniquely summons.
And, you know, I think you might have alluded to it, but on the previous podcast Armin Ronacher mentioned to me, he's the creator of Flask, he's been around the block for quite a while and he's now doing a startup, and he said that right now it's just him and his co-founder and he's got an army of AI interns right now. He's prototyping. But he told me, "I'd like to start to hire people soon because people bring energy, and you need energy for a company to live and thrive." And I'm kind of sensing the same thing.
Oh, for sure. No, for sure. And I, you know, just listened to this great piece with Richard Sutton, who was the inventor of reinforcement learning, and I think rightfully, I agree with him. It's like, you guys are conflating an LLM with artificial intelligence. It doesn't have goals. This is really important. So a prompt is not a goal, and guessing the next word is not a goal. But us together as a startup, and wanting to make it together, not wanting to die here together, that's a goal. And so we can use that creativity. Maybe we, you know, use an LLM certainly as a tool to help us achieve our goal, but I do think that that's a very important distinction.
And can you tell me what kind of tools you use and what are the areas that you find it helpful? I understand you're experimenting with stuff and, you know, this is all work in progress, but where are areas that, and you mentioned like the summarizing was one example, of glossaries.
Yeah. Oh yeah. I mean, I use LLMs as an editor all the time. I find it to be really, I mean, actually it was funny. I had a blog entry that went on Hacker News and someone was like, "Oh, this is LLM written." I'm like, "Actually, it is LLM edited, but the only thing that I did based on the LLM is I deleted an entire paragraph." So there's a paragraph that wasn't working, and the LLM was like, "This paragraph is not working." And I'm like, "You know what? I'm just going to delete the paragraph." So I was like, I don't know, you want to say that's LLM edited? Because every word there is written by me, but there were some words
that there was written by me that an LLM... social I deleted there, which I deleted. So I mean I use it in writing for sure. I mean I also like to use, and this is like a stupid reason, stupid thing, but when you're writing Rust, and we write a lot of Rust there, especially when you're new to Rust, you wonder, like, the way I just phrased this, is this idiomatic? Is there a better way to do this? That's a great little problem for, like, I got this small little snippet of code. Is this an idiomatic way of doing this? Is there a better way of doing this? And that's a great thing for an LLM to make a suggestion or not, or tell you that, like, nope, that's an idiomatic way of doing it, maybe I would make this small adjustment. So I find LLMs to be more valuable in the small than in the large.
So, like, again, my, you know, hats off to people who want to spend their lives acting as middle management for robots, but that's not necessarily for me. Certainly at Oxide, our belief is that people take responsibility for their own work. So, if you want to have an LLM help you out on that, that's fine. But ultimately, if there's a bug in this, you can't blame the LLM. "The LLM broke my code" is not interesting. LLMs don't have accountability.
And so, one thing that is starting to spread across, I think, a lot of engineering is engineers using LLMs either inside your IDE with autocomplete, and also kicking off agents now. Now, there's more advanced ones like Claude Code and Codex where it can actually run command prompts and run your tests. Are you seeing engineers use some of these tools? And there's a little bit of back and forth as well. You know, it's very clear that when you're doing kind of more boilerplate things that are so-called on distribution, which is they've learned, like, React, TypeScript, it can spit out a bunch of stuff, but you strike me as someone who's doing a lot more nuanced things.
Yeah. I mean, you're writing a bunch of C code in the operating system kernel. It is less valuable.
Yeah. But so what are you seeing across the team in terms of...
You know, I encourage people to experiment, and I would say we're seeing a wide variety of experimentation. Certainly we're using Claude Code a bunch and people are doing that. But broadly speaking, for a lot of the work that we're doing, it is helpful as maybe a polishing tool, but less as the epicenter of its creation. It's not true of everything. There's some software for...
No, but that's also nice to hear, because I'm kind of asking you more to put on your CTO hat, who's also very, you know, you're very hands-on and you know what's going on with the industry, because a lot of non-hands-on executives are kind of licking their finger and thinking, oh, we must be 10 or 20 or 30% more productive. But what I'm hearing is things are kind of the same as before, right?
Yeah. I mean, my big belief is it's a tool. It's a powerful tool. I will say that occasionally I get people who are like, well, I don't want to use it at all. And I'm like, you should. So, like...
You should try, right?
Yeah. Like, let me get you off of that position. We had Simon Willison on our podcast. Simon's delightful. And, you know, one of the lines that he has that I really love is people should run these LLMs on their own laptop, where they run slowly and poorly, so they can see the bad output that they generate, so they can understand what some of the limitations are. So I definitely love that. I do think that people should use them enough to know where they are valuable. It's a very important tool in the toolbox. You want to be aware of it, but it's definitely reductive to think it's the only tool in the toolbox, because it isn't.
Now, you're in such an interesting company because, you know, you don't just do software, but you do a lot of hardware.
Yeah.
Have you found any use?
No.
No.
No. Zero. I mean, okay, zero is a bit reductive. I have found it to be useful when, for example, you've got a waveform of an I²C transaction. Actually, amazingly, you can send that to an LLM and have it interpret this, like, hey, what am I seeing? Is this I²C-compliant behavior? And it can help you out on that a little bit, but it's absolutely at the edges.
Okay, so that's a 0.01.
Also, I think people don't realize there are already tools for that. That's what EDA is. You spend a lot of money on it. We're not laying stuff out by hand with graph paper. When you do layout for a board, there are a bunch of rules that are automatically checked for SI. We do a bunch of simulation work. We're not doing that by hand. We're using software.
Yeah. I saw you have those machines in there. I saw that. I think it's a bit reassuring to hear, because I think it's very clear, maybe we don't realize as software engineers, but programming is such a great use case for LLMs. It's a simple grammar, you can validate it, and I think it's sometimes nice to just, you know, touch sand of an area that is very, very different.
Yes.
But it's cool that you're checking and you're seeing if it changes over time. I guess you always keep checking.
Yeah. For sure. And I think it is frustrating to me because programming is such a good use case for certain kinds of programs. So as a result, you end up with certain kinds of programmers who, in part because of their own self-centric view of the universe, believe that, oh, this is just going to replace every job. And it's like, no, not even close. Not even close. You need to get outside a little bit more.
Yeah. So speaking of getting outside and meeting different people, what I noticed when I went to Oxide is it was great. We had double Es, as you say, software engineers, people who used to work on virtual reality at Oculus, all in the same room. Can you tell me about how big the team is? What's the composition?
Yeah, so we've got some more offers going out tonight. So I think we'll be on the order of like 85. I should keep better mental track of it. We've got like 85, plus or minus. And we've been very blessed. We've really put a beacon out there. We've got a lot of people rooting for the company, and as a result we've got a lot of people who want to work for the company. So, as we talked about last time, we really put a lot on folks to describe the work they've done, what's important to them, why they want to work for Oxide.
I mean, a lot of my LLM use is I will look at someone's materials. As you can imagine, we've started to see materials that are heavily LLM-authored. Potential applicants to Oxide, please do not do this. We get people who human-author their entire materials, and then they get to the last question: Why do you want to work for Oxide? Why do you want to work in this role? And they have an LLM spit that out, and you're like, do you think you want to work here? Let's leave aside whether this is right or wrong or cheating or not. It's fine, I guess, but I don't think you want to work here. You're not going to get a job here, because I don't think you actually want to work here. Put it in your own words.
But that process really has allowed us to attract people who themselves are attracted to the company and attracted to the culture, the problem, the team, and it's just extraordinary. I just feel so lucky to be with such an unbelievable group of people across more and more and more disciplines.
I mean, the great thing about our approach is it brings people in who are like, God, I love this approach. We talked about support engineering. People who are like, God, I love this approach, finally QA can stand on its own two feet. I feel that QA has been kind of subjugated by these other disciplines. Now QA is really thought to be as important as anything else in the company, and it is, because from some monetary perspective it is as important as anything else.
Yeah, but I remember when I worked at Microsoft, like 15 years ago or so, the QAs were just on a lower pay grade. The senior QA was at the same level as, I think, software engineer 2 or something, which just kind of implied...
Yeah. You're less important.
You're less important. You're just less important. And so if you tell the world that we think it's as important, do you know who you get? You get people who are extraordinary at QA. You get the best of the best. And so that has been really exciting. And now we've got people coming from, I do love how many different companies, because my belief is that every company has something to teach us, that there is something positive you can take from every company. Now, there are some companies where it's like, you're really scraping the bottom of the barrel.
Maybe not an Ronaldo. They did buy some.
Yeah. That's right. Even Oracle, you can find... that may be a bit of a challenge. Let's not do that one. But you know what, at the time I thought this was a negative, but now I see it: Larry Ellison makes every hiring decision at Oracle.
So what's positive about that?
Exactly. I really think that the Paul Graham essay on founder mode is talking about founders that lost track of their own hiring. Now, I don't like the way Ellison does it. I think you want to trust a team to make a decision, but ultimately I believe that the CEO of a company bears responsibility for every single hire and should be looking at every single hire coming into the company. To me, that is a very important check on these kinds of companies. So there you go, something positive I take from it. And it's telling that your immediate reaction is like, wait, what's positive about that?
Yeah. I'm not sure you undid that take on Oracle.
Yeah. Fair enough. Fair enough. Exactly. There's more from some companies than others, but I love having all of these different experiences present at Oxide, because I do think that there's so much to learn, and you want to take all the positive things. Because I also think that every company, including... Actually, one of the questions I loved that I got once is, what do you not want to emulate from Sun? I'm like, "Oh, thank God," because people think of Oxide as kind of the second coming of Sun Microsystems. There are lots of things I loved about Sun. There are lots of things I did not love about Sun that I did not want to emulate. And so I think for any company there are things we want to leave behind, and when you've got a big, diverse team, you get to go do that.
And one thing that really surprised me last time I was at your office is that it turns out most people were not in the office. They work remote. I would understand that for software, but how do you make that work for hardware development, where physically you do need to be at the hardware sometimes? I understand you need to measure stuff. I saw a lot of units. Sometimes you need to go check on manufacturing. How does that part work?
Yeah, so a lot in people's basements. Fortunately, this is the advantage of making a server and not making, like, a tractor or, I don't know, a wind turbine or something. This is something that people can actually model in their basement, so that helps. But then a lot of even hardware engineering is using these software tools, using EDA tools. You're using SolidWorks, you're using Altium, you're kind of putting this thing together. When you're doing layout, for example, which is a very important task when you're laying out a board, all of that can be done anywhere. That's all just software.
And so there are things where that physicality is very important. And then when you're doing bring-up, you actually need to be at your manufacturer when you do that. So that is also not in an office.
You would need to travel anyway.
Yeah. You need to travel anyway. And anyone coming from the electronics industry is like, "Okay, I'm interested in Oxide, but please tell me I never have to go spend any time in Taipei or Beijing," or Shenzhen or wherever, because you go out there and you're out there for two weeks in a windowless office trying to get this thing brought up. All of our assembly is done here in the United States, in Minnesota. In fact, we've got a bunch of folks out there this week at Benchmark Electronics in Rochester. So this is wonderful.
And one thing that you told me is one of the things that's top of mind for you right now as Oxide is growing: you still have this culture of the same compensation, full remote. It's kind of been the same since the start. What will be the challenge in maintaining it? Because you've worked at large companies. You've seen how it goes. It can get tricky. What are the things that you're seeing, and what are the things that you're trying to do to keep this kind of startup vibe even as you get bigger?
Yeah. So I think the thing that is top of mind right now for me, especially because we raised a big Series B, which is great. I think much more importantly, we're seeing a lot of customer traction, which is great. So we've seen it paying off. Yeah, it really is. It's really great, and we kind of knew that was going to happen in the abstract, but it's fun to actually see it happen, and fun to actually see the customers that have, you know, bought one rack, and now they want to buy a lot more racks. I love what I'm seeing. Very, very, very exciting stuff. That means we're growing the company a bunch.
And one of the things that's very important to me, because I've seen this happen so many times, is companies take their eye off the ball when it comes to hiring in particular. And it is very important to me that we continue to have absolute discipline in the way we hire. And we're doing that. Fortunately, the nice thing about our hiring process is every single Oxide employee has gone through it. So I'm not having to persuade anyone about the importance of our process, because everybody has gone through it. And the thing that we've got overwhelmingly in our favor is, because we've used our values as a lens for that hiring, Oxide's culture is important to every single person at Oxide. That's what it takes to really preserve that. And it doesn't mean that it won't change at all, but the bones aren't changing.
What will change is it will be bigger. And I love the fact that even at 85, we're already so big that Steve and I know everybody at the company, but very few other people know everybody at the company. So when we get everyone together, it's like the best party you've ever been to. Because in college I used to throw the best parties, and the reason I threw the best parties in college is not because of me. It was because of the
roommates that I had. So I was a computer science student who played ultimate. My roommate was an engineer who was on the water polo team. My other roommate was a history student who was in the chorus. That's six different demographics that don't normally overlap. And then very importantly, we made sure that the women's swim team was always invited. The women's swim team, they were like the foundation water player.
Yeah, exactly. Waterfall. You always check their calendar to make sure they can make it. And people loved the parties we had. Why? Because they would meet people that they never met before who were really interesting. And what I love about Oxide is when we get the whole team together, people get all these delightful surprises. So people take me aside and be like, "God, you know, Ry is awesome." I'm like, "Yeah, I know. I know. I know. You know, too now. That's great." Or whomever it is. It's just really exhilarating. And I think that also serves to reinforce how important what we've got is. I tell the team, we have lightning in a bottle. And we cannot take it for granted. And that means that every single one of us needs to rise to the moment. We need to do what our customers need us to do, but we need to do it in a way that protects and preserves what got us here.
So thinking a little bit ahead, let's assume that these AI tools will just get better eventually. They'll be able to help more even on your kind of low-level things. You've been in the industry for quite a while. You've seen a lot of shifts. What do you think are some of the things, both in software engineering or in hardware engineering or just in general engineering, that will probably not change even if we predict these things being more capable?
Yeah, I think that it's certainly a revolution. I think it's going to allow us all to do more. I do think that we are going to hit a point where people understand that this is a tool, because there's a little bit where we still have this tension of like, oh, is this going to be AGI? Is this going to replace all jobs? And this is nonsense as far as I'm concerned. And it's distracting kind of nonsense. And we actually need to get back to putting the tools in the toolbox of the human that's building it. Now these tools have become much more powerful, and I think that's extraordinary. I think it's important. I think that also, we've got a lot of experiments right now, we humanity, that I'm not sure are going to make economic sense. So we'll be figuring that out as well.
But I think that one of the things I am a little bit worried about is a little bit of despair from younger software engineers in particular, who are like, what's the point? Like, an AI can do all this.
Well, and there's also the news, even from more experienced software engineers, in the mainstream media, there's this news that company X is laying off half their workforce because of AI. And by the way, when we look closer, it's not because of AI, but it is coming across, and it does give not just younger people a lot of anxiety, tons, even mid-level folks or even some more experienced. It does give a sense of, I think it's the first time in computer history that most of us remember that there is this thing that could threaten my job, and I think we've just never had to deal with this. I think there are industries that might have been a bit more used to it.
Yeah, I would say that there have been busts before. The dot-com bust was a bust, like a lot of jobs did disappear, right? But the bust has really come in what feels to be a broader and more permanent way. My view is, this is an opportunity for, I mean, I think one of the things we should be, society, really encouraging is new company formation. Because now, just like you're talking to Armen about how just a small group, just Armen and his co-founder, were able to do so much together, right? We should be really encouraging that. And what are some of the gaps that we can all go fill? Because ultimately we all need to find a livelihood, we need to find meaning, and the way we do that as engineers is we build useful things. And so we can now build many more useful things. What would we go build? If you could build anything, what would you go build? And that's kind of the question that people need to ask themselves.
It's scarier. It's scarier than like, go to this school, concentrate in this, and then mama Google will hire you and take care of you and feed you breakfast. It's like, no, that's not what's going to happen. And it feels a lot scarier because it feels like there's at some level less security, less job security. And yeah, that's true, and that's scarier, but there's also a lot more opportunity.
And for a college student or someone in school or with little experience who says, look, my goal would be one day in like 5 years' time to be so good that I could get a job at a place like Oxide. It doesn't need to be Oxide, but again, a place that has a high bar. They often hire experienced people, but I want to get there, and yeah, there's all this AI stuff happening. What would you advise them in terms of what to focus on, what areas to study, what things to do, or how to think about it? They have the goal there. What advice would you have them part with?
Yeah, so I think that you need to have a different mindset, and that mindset needs to be not around how do I create as much as possible, but rather how do I get better? How am I getting better every day? And I think LLMs are a great tool to get better. How can I learn about something new? Go deeper. Go into something that I wouldn't go into before. Get over that kind of fear. And especially if you're in school now, you want to work at a place like Oxide, you kind of have to view it as, all right, you want to play Major League Baseball, that's great. You're a great high school player. You want to play Major League Baseball. It's really hard. Got to get better every single day, and you're going to need to be really focused on getting better, and you need to be really realistic about what I need to go do to get better. And it's hard, and it's chancy because you might not get there, but you could get there, and you're certainly not going to get there if you don't focus on that kind of self-improvement.
So I really think that there is a shift in mindset that needs to happen, or that one needs to have, I would put it that way. You really got to have a mindset towards getting better, understanding more. What do you not understand? There is lots that you don't understand. I think one of the challenges of modernity is that we delude ourselves into thinking that we understand it all. You don't. I don't. One of the things that I've learned, I've joked at Oxide that I keep waiting for the day that I know how computers work, and it's like, it wasn't today, definitely wasn't yesterday. It's not like it's going to work.
You understand how
But I mean that earnestly, in that the amount of complexity, I mean, I knew but also didn't know. It's like every day I feel I'm still learning new facets, and not just of a computer, but actually delivering a computer to people. There's so much to learn out there. So many op... And now, the way you've got to view LLMs is not like this thing is coming for my job. You got to view it as, no, I've now got this private coach, tutor, what have you, that I can ask any question to. I got to fact-check its answers for sure, but now you've got the opportunity, and it is easier to get into this domain than it ever has been. And that's great and it's powerful, but it can also be scary.
And as closing, what's a book or two books that you would recommend to folks and why?
Oh, so many good books. I've got a 21-year-old, an 18-year-old, and a 13-year-old. And when the 18-year-old was in, he's now a freshman in college, when he was a high school senior, he got this assignment, great assignment from his English teacher, namely: go to someone that you know and ask them for three books that they would recommend that you read, and I'm going to assign you one of those three books to read, and you're going to read it, and then you're going to talk with them about that book. And I'm like, "Oh, I love this assignment." So he's like, "Dad, I'm coming to you." And I'm like, "Oh, thank you." And of course my wife was like, "Why didn't you come to me?" Like, "Hey, look, sorry." It was great. So yeah, I'll give you those three books that I gave to him, and I think that each of these is really terrific.
First is The Soul of a New Machine by Tracy Kidder. This one won the Pulitzer Prize in 1980 or 1981, about the building of a new computer at Data General, and it's extraordinarily well written. And even folks who are like, well, what do I have to do with a computer company in the late 70s and early 80s? Any engineer will see something of themselves in that book. It is just masterfully told. Tom West, who is kind of a complicated figure, but The Soul is still, I mean, it's literature for us. So I would absolutely say The Soul of a New Machine. Every engineer should read The Soul of a New Machine by Tracy Kidder.
For me personally, very influential was Skunk Works by Ben Rich. So about the history of Skunk Works. Clarence "Kelly" Thompson was kind of the originator of Skunk Works at Lockheed Martin. Extraordinary story about what engineers can do when they kind of task themselves on the impossible.
It's such a good book.
It's such a good book. Amazing book. And then the other one is Steve Jobs and the NeXT Big Thing by Randall Stross. So Steve Jobs is kind of lionized by the industry, but people forget about a very important chapter of his life, namely NeXT. And I believe it was just an anniversary, maybe it was the 30th anniversary, it must have been, or maybe the 40th anniversary, Jesus, of the announcement of the NeXT machine. So Steve Jobs left Apple, was fired from Apple, started a computer company called NeXT. Really interesting company in a lot of ways. He was at NeXT for a very long time. It's a 13-year journey before NeXT was bought by Apple. NeXT is bought by Apple, Steve Jobs returns to Apple when they buy NeXT. This book, Steve Jobs and the NeXT Big Thing, is written before Apple buys NeXT. And it is at Steve Jobs's lowest moment. It is not here to praise him. It is here to bury him. And it is very interesting about all the missteps at NeXT.
And the thing that we cannot know, because Jobs obviously died, but I believe, having read the book, which, NeXT gets essentially no treatment in the Isaacson biography. NeXT is like six pages of glory. It's like, that's not what it was. But Randall Stross's book is masterful, and in particular, I believe that Jobs's failures at NeXT were essential for the resurrection of Apple. Because you look at the way he handled himself coming back to Apple, it was very different from the Jobs that got fired from Apple. And I think that when people look at Jobs, they don't really take him apart. And I think you should, because I think he's a really interesting guy. He's enigmatic. He did things that I think are really fascinating and also things that I really strongly disagree with. So just to be clear, I'm not... but I think that he's indisputably an important figure, and that book is by far the best book. So Steve Jobs
No, I'm adding that. I actually want to read that now.
Oh, it's extraordinary. It's very good.
Well, Bryan, this was such a fun discussion.
Oh, my pleasure. I mean, we knew this was going to be long and wide-ranging, so hopefully it delivered, but I really appreciate it. We went from the '90s all the way to the future.
Awesome. Well, thank you so much for having me. It was terrific.
I've got to say Oxide is one of my favorite companies, and I say this as someone who has zero affiliation with them. It's just so rare to find a startup that built both hardware and software and are world class in doing both of these, and are so open about talking about exactly how they do it all. Honestly, the only downside I can think of about Oxide is how their server racks are built for pretty large companies and are definitely out of reach for hobbyist devs.
In this episode, I really appreciated how much of a straight shooter Bryan was, especially about the impact of AI tools. Yes, everyone at Oxide uses them, and they do find use cases for coding and working with documents, but it's eye-opening how it gives them basically zero help with hardware engineering. This is a good reminder that LLMs might be the single best fit for coding-related tasks. And as devs, we should know that these tools might be more specialized than many people think.
I hope you enjoyed the stories in this episode as much as I did. If you'd like to learn more about Oxide, I did a two-part deep dive about the company, and you can read it linked in the show notes below. If you enjoy this podcast, please do subscribe on your favorite podcast platform and on YouTube. This helps the podcast a lot. A special thank you if you also leave a rating on the show. Thanks, and I'll see you in the next one.
Article published
