Casey Muratori on Why Software Performance Matters and Why the Industry Keeps Ignoring It

Open on YouTube ↗
Overview

Casey Muratori is a programmer and game developer, founder of Molly Rocket, and author of the Substack Computer, Enhance, which focuses almost entirely on software performance. His argument, repeated throughout this conversation, is that most software runs 10 to 100 times slower than it should. That gap, he says, usually comes less from missed micro-optimizations than from architectural decisions made by people who never learned what the hardware can actually do. The conversation covers why the industry has neglected performance and why he thinks that is starting to change. It also covers how optimization should be done, why every programmer should learn to read assembly, his critiques of "clean code" and test-driven development, how game development has changed, and why he writes all his code by hand.

29 min read

From Digital Equipment Corporation to games on Windows

Muratori learned to program at seven, around 1982. His father was a programmer at Digital Equipment Corporation, the minicomputer maker behind machines like the PDP-11 and VAX. Parts of DEC were later absorbed by Intel and Compaq, and his father ended up at Intel without ever changing jobs. Because of that, the household always had computers at a time when few families did, and, more importantly in Muratori's telling, there was a programmer at home to teach him.

He got into games by chance, through an internship at Microsoft. At the time, Windows was not a gaming platform. It had little beyond Solitaire and Minesweeper. Muratori explains the technical reason. Games needed to fill pixels quickly with the CPU, draw to a back buffer, and flip it to the screen. Windows 3.x offered no fast way to do that. Programs had to hand Windows a bitmap that might not match the display format, and Windows then translated it. Games like Doom were not coming to Windows.

Chris Hecker, whom Muratori calls an unsung hero of Windows gaming, set out to change this with WinG, a library for fast blits to the screen. Hecker had no authority to build it. It was a skunkworks project inside Advanced Technology, an early precursor to Microsoft Research that was not supposed to ship core Windows libraries. Hecker's manager, Michael Edwards, provided cover. WinG shipped anyway, and Muratori describes it as the first step toward DirectX. He adds that the institutional push for DirectX came from three people, and one of them had been the tester on WinG, so the lineage was direct. Many programmers, such as Todd Laney, did essential core work on it.

Muratori arrived for his internship the week WinG "blew up" internally. Edwards, who was supposed to be his manager, had stormed out of the building and was not seen for about a week. Edwards later returned and was moved to another division. The upshot was that Muratori landed at "ground zero" of games on Windows. He met Hecker, who took him to Humongous Entertainment to meet Ron Gilbert, creator of the SCUMM engine and one of Muratori's childhood heroes. Gilbert gave him a Secret of Monkey Island mouse pad.

After that came an uneventful startup with Hecker, then Gas Powered Games (Dungeon Siege), then a long stretch at RAD Game Tools. There he built a character animation system that was widely used and, to his surprise, is still in some studios' pipelines through source licenses, even though he has not touched it since 2004. Since then he has worked independently through Molly Rocket: contract work, the Substack, and some game work, including the movement system for The Witness, which he did as a friend helping out on a big project. Molly Rocket also has an unannounced project. They are keeping it quiet because the Substack comes first and they cannot predict how much time they can give it.

Why the industry neglects performance and why that may be changing

About three years ago Muratori messaged the host to ask why the industry puts so little emphasis on performance when the evidence says it affects the bottom line. Looking back, he gives three explanations.

The first is the host's original answer, which Muratori accepts for part of the market: in enterprise software, the user is often not the buyer. Someone high up picks HR software based on cost, compliance terms, and legal liability. Whether each record takes 30 seconds to open never appears on that sheet. The people who suffer from slow software often have no say over it.

The second is monopoly and network effects. Muratori points to Bluesky and Threads trying to challenge X, and to the entrenched positions of Facebook, Instagram, and TikTok. Better responsiveness might be one part of a pitch against an incumbent, but without a plan for adoption and influence, it cannot sell a product into a monopoly space. Downloadable apps where users choose freely are a shrinking share of all software.

The third is more hopeful. Muratori thinks a decade of people, himself included, pointing at the problem has had an effect. He sees more discussion of performance, more published benchmarks, and new products that lead with performance against incumbents, naming File Pilot and the Blick video editor. The host adds Bun versus npm, where claims of being 10x or more faster got developers' attention, and Linear versus Jira, where Linear advertises a 300 millisecond budget per action.

Muratori says that 300 millisecond number shows how low the bar has fallen. Three hundred milliseconds is "an eternity in computing." Users often wait seconds for simple operations, even though network round trips to a physically distant data center can finish in under 10 milliseconds. When he says software is 10x to 100x slower than it should be, people don't believe it, but he argues there is plenty of proof.

Optimization starts from the theoretical maximum

The host brings up Simon Eskildsen's "napkin math" project, a list of baseline costs for operations such as transferring data between data centers or writing to an NVMe drive. According to the host, Eskildsen noticed that when Shopify teams compared database vendors with their own benchmarks, the results sometimes made no physical sense. A write that should take around 100 milliseconds at the theoretical limit was measuring 10 seconds, and it often turned out the benchmark was wrong.

Muratori says this is the core point of his Substack. The common idea of optimization is to run a profiler, find the big parts, change something, gather statistics, and keep the change if the numbers improve. He says that is "not how it is done" by any of the many excellent optimization people he has worked with. The right approach is to first ask what operations the system must perform and what the hardware can do at its theoretical peak. You then measure the gap between that peak and what you actually get, and your job is to shrink the gap until what remains can plausibly be explained. You often can't reach the theoretical number, which is why it is called theoretical.

Without that reference point, he argues, you are only finding a local minimum. He calls that "improvement," not optimization, because optimization means making something optimal. The theoretical baseline also helps you learn. Hardware keeps changing, and a big unexplained gap can point to something nobody knew about. He says this has happened while producing the Substack: they found a register-renaming behavior in recent Intel chips (he mentions a "RAT table") that nobody had documented and that now had to go into their performance model. Back-of-the-envelope estimates are how you discover such things.

Why you should learn to read assembly

Muratori's course teaches reading assembly language, and he explains why. Source code in Java, C, Haskell, OCaml, or Rust is only input to a compiler. It tells you nothing certain about what the CPU is asked to do. The assembly output tells you exactly. He stresses that the skill is reading, not writing. Writing assembly is rarely needed beyond occasional test cases where it is easier than coaxing a compiler into producing a specific sequence. Reading it, he says, is essential for optimization work.

It also opens up other knowledge. When a vendor presents a diagram of a new core, someone who reads assembly can see things like peak multiplication throughput right off the chart. Without that background the diagram looks like a meaningless flowchart.

The host notes that assembly is much simpler than high-level languages, and Muratori agrees strongly. To understand a modern website you need the JavaScript syntax and libraries, the DOM, CSS, and React. For assembly, he estimates you might need to learn 20 or 30 instructions, because compilers rarely emit most of the legacy x64 instruction set. And in performance work you usually look only at a small piece of code you already suspect. His summary: "If you can vertically center a div in HTML, then you can probably learn assembly language."

"Premature optimization" and the architecture you can't fix later

Playing devil's advocate, the host brings up "premature optimization is the root of all evil" and the usual practice of building first and optimizing only if needed. Muratori has a two-hour lecture on the history of that phrase from this year's Better Software Conference. Here he addresses the idea behind it, which he says is not entirely false.

Deferring optimization works when the slow code is local. If you wrote a naive loop, or a simple hash table where you could spend a week researching the fastest one, and you understand the problem well enough to know the surrounding architecture won't need to change, deferring is sound engineering. Maybe the naive version is always fast enough. If not, someone can target that one spot later.

It fails when you don't know whether your choices leave the code optimizable. His example is a codebase where every operation asks the server for something, waits, computes, and then asks for the next thing. Hundreds of thousands or millions of lines get written that way. When it turns out to be too slow, the performance experts who are called in can only say there is nothing they can do. The code is a serial dependency chain: each step waits for the one before it, so it can't be parallelized, multithreaded, widened, or amortized. Overall performance is generally bounded by the longest such chain. If the team had instead been told to request everything an operation might need at the start, and to chain requests only when truly unavoidable, the problem would not exist. Fixing it after the fact means rewriting nearly everything, if that is even possible.

His conclusion is that a codebase with a few hotspots that can be optimized later no longer happens by accident. It has to be designed that way up front. Everyone making architectural decisions must understand performance so that the people downstream inherit an architecture that can be optimized. Otherwise, "you're just rolling the dice."

As evidence, he points to the many company blog posts, from Facebook, Uber, and others, about rewriting entire systems for performance. If performance problems were always hotspots, full rewrites would never be needed. The host describes watching teams at Uber move from Python and Node.js to Go and Java. According to the host, OpenAI and Anthropic are now doing something similar, moving services first built in Python because their ML staff knew it toward Rust (or possibly Go) to get multithreading and more connections per machine. The host adds that public engineering blogs tend to describe these moves in flattering terms, since they help people get promoted and may be shaped by content teams. Muratori replies that the fact these rewrites keep happening is the signal. If the conventional wisdom were right, the only reasons to rewrite in another language would be preference or something like memory safety, not performance.

A learning path: much less than people expect

Asked how an engineer should get better at writing fast software, Muratori starts with good news. Heavy hotspot optimization, like hand-coding routines in assembly, is rarely needed today. Many libraries are already optimized, and modern CPUs are very good at running bad code quickly. The real task is avoiding the choices that make software 100x slower than it needs to be. If you avoid those, he says, you'll typically be within about 2x of optimal, "which is 50x better than the people who are 100x away."

So what you need is one solid round of experience: learn to read assembly, see how the CPU works, time some code, and experiment. He suggests "a month or two of nights." One of the first things his course shows is the assembly executed for a + b in Python. It is so long he has to skip most of it, while the C equivalent is a single add instruction.

The host asks whether a React or iOS developer would ever look at assembly in their daily work. Muratori says the point is to internalize the orders of magnitude. Once you know that an add in Python costs roughly a hundred times more instructions than in C, it makes sense why Python programs lean on libraries written in C, and why operations over large amounts of data can't be written in plain Python. With that knowledge you can decide whether a piece of code can afford a 100x slowdown, and if it can't, reach for a library or something like Cython and structure your code around it, so the 100x cost applies only to rare operations. Most programmers can make that call without being performance experts. The key is knowing the question exists. After that, a quick search, or asking an AI, will turn up the specifics.

What to understand about the CPU

Muratori clarifies that assembly is really a way to see what the CPU is doing, and the CPU is what matters. Modern cores such as Apple's M series, AMD's Zen, and Intel's Core line are not documented at the level of every internal mechanism, and they are too complex to study that way anyway. Treated as black boxes, though, they can be understood through a few categories.

The first is how data moves into and out of a core: load and store units, the cache hierarchy (L1, L2, L3, and now sometimes L0), and its granularity and policies. This matters because data layout and access patterns can change performance enormously, and those are architectural decisions that are hard to undo. The second is how instructions flow through the core, which covers concepts like branch misprediction and instruction cache misses. He says these are simple to understand conceptually even as predictors get more sophisticated. The third is execution scheduling: the raw throughput of floating-point multiplies, integer adds, divisions, and so on, and how assembly instructions become micro-ops distributed across execution units. With that background, he says, you can look at the diagram for a new core and roughly know its performance.

He doesn't think everyone needs to become an extreme optimizer. It is fun if you enjoy it, but not the important part. He argues this baseline should be standard for software engineers: "You go to school for four years to learn this. There's no reason you can't learn this in a few months."

The host brings up craftsmanship, and Muratori agrees that many programmers find high-level work unfulfilling and feel much more satisfied once they understand what is happening underneath, even if they keep working at the same level. He adds that it compounds. If library maintainers take performance seriously, everyone using those libraries gets faster, and APIs designed to be optimizable spread that benefit further: "The more people are doing performance, the less people need to do performance." The knowledge also stays current fairly easily, because CPU vendors present their changes and people running microbenchmarks publish the quirks they find.

How games used to get built

Muratori warns that his view of game development is dated. The biggest revenue earners today, including Fortnite, Roblox, GTA Online, and Minecraft, are live services that ship features incrementally, much like SaaS. He has friends at such companies but has not worked at one.

He can speak to the earlier era. Back then there were no licensable engines. Code reuse was informal, like borrowing a routine from a colleague at Atari. The first wider reuse came with id Software's Doom and Quake engines and Ken Silverman's Build engine, but games built on them tended to closely resemble the original game. A studio's existing codebase was a competitive asset. Blizzard could carry its Warcraft work into Warcraft II, while any competitor had to build pathfinding, level editors, and rendering from scratch.

Projects faced two big risks. One was engine risk: could the team build technology that did what the game needed, fast enough to develop on? Developers couldn't buy much faster hardware than consumers would have. Some bought SGI workstations for exactly that reason. There was no way to reduce this risk except grit. Some games failed on technology alone, and some studios survived on technical skill, like id with first-person engines and Bullfrog with the pseudo-3D engine behind Magic Carpet and Dungeon Keeper.

The other risk was whether the game would be any good. Without a working engine or level tools, it was very hard to know what you were making. Muratori cites Thief: The Dark Project from Looking Glass. From what he heard from the team, its core gameplay only came together near the end. As budgets grew into the millions, studios moved to building a vertical slice first: the whole studio builds one playable, hacky slice as fast as possible, proves it is engaging, and only then plans asset production. Stressing that he is not a game historian, he believes this was a major shift that made development much less "seat of the pants."

Licensable engines as the games industry's "AI transition"

The host suggests that widely available engines like Unreal, Unity, and Godot may hint at what AI will do to software in general. Muratori calls this a brilliant analysis and says he has told people the same thing: licensable engines "kind of was our AI transition already," and "the news is not probably that positive."

Early on, the effects were good. People who could never have assembled a technical team could now make games, which enabled new artistic work and some excellent titles. But the flood of releases followed. Muratori estimates Steam now sees tens of thousands, perhaps a hundred thousand, releases a year. It used to be that a fun game would be found through word of mouth or storefront exposure. Now a game will almost never be found organically. Small hits with no marketing still happen, but "the chances that you will be that game are like zero." Without a real marketing strategy, he considers it unwise to expect more than a few thousand sales. The host sums up that quality is now table stakes and distribution is the differentiator. Muratori agrees, and says he doesn't know whether it was a good trade.

New games also compete with old ones. The host mentions spending hours on the '90s game Death Rally. Muratori explains why this is getting worse. Old games used to look obviously dated, and in 1995 technical advances drove sales. A large share of today's market by revenue no longer cares much about graphics beyond the current level, so a 2017 game "just looks fine." Live-service games like Fortnite, Minecraft, League of Legends, and Dota take up many players' hours, and entertainment time is zero-sum, shared with Netflix and everything else.

Why GTA 6 is taking so long

Asked how GTA 6 can take more than a decade despite better tools, Muratori says to look at it as a business decision. GTA 6 is not only a new game for players. It replaces GTA 5, which he believes was by far the highest-grossing entertainment product ever, with the online portion earning billions, "kind of like Fortnite before Fortnite." Rockstar and Take-Two are replacing their most profitable product, which is still earning, and the worst outcome would be a successor that cannibalizes it and then earns less. He's sure many people on the team care about the single-player experience artistically, but he assumes a lot of planning has gone into the live component.

In his view, GTA 5's online success was probably a surprise to the company, and Red Dead Redemption 2's online mode didn't reach the same level. That makes GTA 6 the first release where Rockstar knows the audience is there. He compares it to relaunching Google Search: "if I was in charge of that project, I would be sweating bullets." He doesn't predict whether it will succeed.

Critique of "clean code"

The host brings up Muratori's video "Clean Code, Horrible Performance." In it, a polymorphism-based design in the style Robert Martin recommends ran about 1.5x to 15x slower than a plain switch or table version. Muratori says the reaction was more positive than he expected, though it was certainly controversial.

He separates two meanings of "clean code." If it just means code you consider well written, there is nothing to argue about. His target was the specific rules: keep functions under a certain length, avoid knowing types at runtime, always prefer polymorphism. He calls these "just bad programming practices" when combined, though some are fine alone. Many small functions are fine, for example, if the compiler can see them and safely inline and merge them.

He says the problem isn't the cost of a virtual function call itself, which depends on things like branch prediction and stack traffic. The real cost is that runtime dispatch stops the compiler from optimizing. When functions are known at compile time, a modern optimizing compiler can merge small functions, remove redundant work, and even widen code paths for SIMD, turning mediocre source into efficient machine code. When it must allow for any class being substituted at runtime, it can do none of this, and he says markers like final don't reliably fix that. He believes the video's example showed a milder slowdown than a large production codebase built this way would. His position is that code can be maintainable and readable without following those rules, while still leaving the compiler room to do its job.

Testing should inform design, not drive it

On test-driven development, Muratori calls his position pragmatic. Tests are worth writing when their cost to create and maintain is less than the time they save, for example by catching bugs that would be hard to find or expensive once shipped. At RAD, he kept a regression tester for the core routines he wrote so he could cover the cases customers would hit.

What he objects to is the "driven" part. Development shouldn't be test-driven by default. A team might sensibly decide to drive a particular project mainly through tests, but for another project that could be a costly mistake. The cost of tests includes what they do to the codebase. If changes mean rewriting tests, teams may avoid changes they should make. Some projects should have many tests and some very few, and he says knowing which is which is part of being a good engineer.

What good code and good engineers look like

For Muratori, good code does what the machine needs to do to solve the problem as directly as possible. It is split into well-chosen, well-named pieces that others can follow, especially your future self. It avoids redundancy; his example is writing a Euclidean distance function once rather than repeating the formula everywhere. And its pieces can be recombined by the compiler into efficient code. He has never understood the idea that well-architected code and fast code are in conflict. In his experience, properly architected code is usually the fast code. Readability only suffers when you push very close to the theoretical maximum on specific hardware, which he says is rare. Code only becomes both hard to modify and slow "once you think you need to have 27 factories and 8,000 microservices." He wants good code to leave the path to theoretical performance open for anyone who needs to go further later.

On good engineers, he uses a baseball analogy: are you asking about a pitcher or a designated hitter? Some great engineers are utility infielders who quickly learn a messy codebase and patch what needs fixing. Others spend eight months exhausting one problem, sometimes to the point of inventing new algorithms. His advice to anyone building a team is to think about roles, not a generic "great software engineer."

Pressed for traits they share, he gives two. He can't think of a great engineer who couldn't read assembly, so curiosity about how things work seems common, though how much they use it varies. Some work at that level constantly. Others just avoid architectural mistakes because they know, for instance, that certain work has to be batched. The second trait is not being dogmatic about ideas they haven't tested themselves. He calls much received programming wisdom "nonsense" that no one has actually tested. Its concrete downsides can sometimes be demonstrated while its upsides often can't. He values engineers who look at what works in practice and can be measured repeatably, rather than following a talk because "someone at Google says always call memset or never use if statements."

Why Muratori doesn't code with AI

Asked how Molly Rocket uses AI tools on its unannounced project, Muratori says, "We are not using them at all." He presents it as a philosophical choice, not a productivity judgment. The project exists because they want to program it: "If I just wanted an AI to program them, I'd just go get the Unreal Engine." He says he did not reject AI over time savings, copyright, or ethics, though he calls those fair grounds for evaluating it. It simply doesn't serve the project's goals.

More broadly, he says nobody can predict what AI will look like in ten years. But he expects a handcraft tradition to survive, as it has with other automated work. IKEA exists, and someone in an industrial district still welds iron-and-wood tables that some people want. People still knit hats they could easily buy, some even raising the sheep themselves. He doesn't claim this has more or less value, only that humans do it. Given the choice, he would be "the organic farming guy" and has "no interest in managing a division at IKEA." He wants to be among the people who keep programming by hand alive, and he says he has little to add about AI coding workflows.

AI's effects so far: too early to judge

On AI in the games industry, Muratori says it is too early to judge. Some people have been saying AI writes human-quality code for two years, but the people whose opinions he trusts didn't consider it very usable until much more recently. He suggests waiting at least another six months to a year for best practices to settle. He knows many people in games are using AI, but he hasn't seen visible results, like Fortnite shipping weekly without bugs. In his view, there are several possible explanations. Current tools still need humans to set them up and fit them into their work, so adoption takes time. The models may need to improve further. Or real gains may already exist but be too small to see from outside. A 10% improvement across the board would be impressive but almost invisible. If AI turns out to be truly transformative, he says, it should become obvious, "five people are now shipping Fortnite instead of 5,000."

The host describes hearing from many engineers who feel AI fatigue or burnout. They are prompting instead of coding, under pressure to produce more, and wondering why they are there. Muratori says he has heard of such cases but hasn't encountered them directly, likely because most people he talks to have a lot of say in how they work. The host connects this to something Armin Ronacher told them: people with autonomy tend to welcome AI, while people handed tickets tend to see it as a threat. Muratori finds this logical. If you have autonomy, you use AI for tasks you didn't want to do, so even when it fails, the downside is small. If you are told to use AI for the work you wanted to do, maybe right after layoffs, the psychological effect is completely different. He suggests a question: "Are you using an AI to do your job, or is an AI using you to do your job?" He notes that reports about Meta, from the host and others, suggest employees felt they were mainly generating training data to replace themselves. He agrees that people who have options might weigh autonomy when choosing a job.

Closing advice: read papers

Asked for book recommendations, Muratori instead recommends reading research papers. Before working in an unfamiliar area, he reads many papers, follows their references, and uses survey papers to find more. He learns historical context and, often, techniques he had never heard of. He doesn't name a specific paper. He suggests searching Google Scholar for a topic in your own domain, reading one paper, and following the references. Since he doesn't use AI, he can't say from experience, but he guesses AI tools might be good at suggesting papers based on your interests.