Chris Lattner on LLVM, Swift, and Mojo: Building Blank-Canvas Systems
The Pragmatic EngineerChris Lattner created the LLVM compiler infrastructure and the Swift programming language, worked on AI infrastructure at Tesla and Google, and now leads Modular, the company behind the Mojo language. In this conversation with Gergely Orosz of The Pragmatic Engineer, Lattner explains how each project began and what it took to get adopted inside large organizations. He also describes the mistakes he carried forward into the next project. His central argument is that the tooling he has spent his career on exists to widen access: Swift brought people into iOS development, and Mojo aims to open up GPU and accelerator programming, which Lattner sees as gatekept by aging, vendor-specific tools.
The compiler landscape before LLVM
Lattner started with Linux in the mid-1990s, when GCC was the compiler that standardized much of the software ecosystem. Before GCC, he explains, every hardware vendor built its own compiler. In the late 1980s and early 1990s there was far more diversity in instruction sets, with HP, Intel, and the new RISC designs all competing. The C standard existed, but like most specifications it was incomplete. Each vendor's compiler had its own bugs, misfeatures, and missing capabilities. Software of that era relied on tools like autoconf, which Lattner calls "this really weird macro processing thingy," to work around the differences, and he describes the result as "a gigantic nightmare."
GCC cleaned that up. Chipmakers adopted it, it was free software, and Linux built on it. Lattner says open source owes GCC "a debt of gratitude." By around 2000, though, GCC was over 20 years old and its architecture was old-school. It was built to do one thing: take C in and emit assembly for a given chip. It was not modular, so it could not do just-in-time (JIT) compilation. At the time it also could not optimize across files, so a function in one C file could not be inlined into another.
LLVM as a university project nobody expected to succeed
Lattner began LLVM around 2000 while studying compilers at university. It started as a code generation system that could target multiple architectures, but he frames it as a way to learn by building. He stresses that successful systems look obvious only in hindsight. Most university research projects go nowhere, and the default assumption was that LLVM would go nowhere too.
His advisor, Vikram Adve, encouraged the group to keep building it, and it was used in a couple of classes. At that stage LLVM was only an optimizer and code generator. It plugged into GCC and used GCC's parser for C and C++, so it was useful to compiler researchers but not to application developers. When it was open-sourced, it attracted "two or three people," mostly other compiler enthusiasts.
Lattner's approach to building a community was to treat outside contributors like members of the research group. That meant open development, answering questions, and treating people with respect. The community grew very slowly, drawing people interested in esoteric languages, in performance, and in other niches.
LLVM also differed from GCC in being written in C++. That was controversial because GCC was written in C and Richard Stallman opposed C++. Lattner describes a "parallel universe": after joining Apple in 2005, he proposed to the GCC community that the two projects merge, arguing that LLVM could lift GCC's architecture. The proposal went nowhere, partly because of C++ and partly because of "not invented here." Experienced GCC developers, as he puts it, responded along the lines of "okay kid, why would we listen to you." He says he is glad the merge didn't happen, because LLVM had to grow up on its own.
Making LLVM relevant to Apple
By graduation LLVM was five years old, with releases every six months. Lattner says he deliberately advocated for the project and showed momentum. That caught the attention of an Apple VP and of a few Apple engineers who were adding a PowerPC backend to LLVM. Apple was frustrated with GCC, which was the foundation of its compiler technology, including Objective-C. GCC was hard to extend, performance was lacking, and the GCC community was annoyed with Apple for various reasons. Management told him to come work on LLVM.
He started with one other engineer and little guidance, and spent his early time on PowerPC support and Mac OS integration. Then his manager warned him that the work eventually had to matter to Apple. LLVM wasn't in any product, and it was worse than GCC in several ways. Lattner got the message that if nothing shipped within about a year, he would be asked to work on something else. With his manager's help, which he calls "amazing," he looked for near-term business impact.
The first use turned out to be graphics. OpenGL did just-in-time compilation, and LLVM could handle a small piece of that work. The value was modest, "at least non-zero," but it changed LLVM's standing from science project to something useful. That foothold funded the next set of features, then the next.
Over about five years, the team replaced Apple's developer tools piece by piece: compiler, code generation, and debugger technology. They built Clang, a new C++ parser and Objective-C front end. Lattner points to the first 64-bit iPhone, the iPhone 5S, as a milestone the LLVM work made possible. According to Lattner, the industry thought 64-bit phones were pointless, asking why a phone would ever need more than 4 GB of RAM. He says Apple shipped the iPhone 5S before ARM had even received its own 64-bit chips back for testing. Apple had built the entire toolchain and OS support internally while collaborating with ARM, and ARM "had no idea how far ahead we were."
Asked how a side project reached the core of Apple, Lattner credits a combination of forces. There was top-down support from people frustrated with GCC. There was bottom-up success, with something new working every six months. And he built a reputation for getting things done, which brought more scope: he went from engineer to manager, second-line manager, and eventually senior director. Because the work happened in the open, other organizations adopted LLVM as well. Cray used it for a supercomputer, Google adopted it around 2010, and companies such as Intel and ARM eventually cancelled their internal compilers and moved to LLVM. Today, he says, LLVM has a community of thousands.
On motivation, Lattner says he doesn't start from a picture of eventual success. What he loves is doing the work and "figuring out that really obscure thing" that makes an architecture compose and scale. As a result, he says, most of the time he works on things people don't understand, and he has "normalized to this." Projects click for others only when they map onto a value system people can measure.
Asked how he got Apple to accept open source, Lattner says Clang was easy: "we didn't really ask too many people for permission. We just kind of did it." The hard one was Swift.
Swift's origin: nights and weekends
By WWDC 2010, Lattner's team had shipped a full-stack replacement of Apple's C, C++, and Objective-C toolchain. Building a complete C++ compiler was a major technical achievement, but Lattner found it demotivating because C++ is "a beast of a language." Coming out of that project, he wondered whether something better was possible, and with the infrastructure now in place, the team could build almost anything.
He started Swift on nights and weekends without asking permission. For the first year and a half he was the only person working on it, while also holding a day job managing a team of about 40 people. His first step was a survey of other languages: Java and C++, functional languages like OCaml and Haskell, esoteric languages, and newer ones such as TypeScript, Dart, and Go.
He defined success as a language that was easy to use and scalable. JavaScript shaped that thinking. It started as a simple hack for onclick handlers, gained adoption, and developers then carried it into other domains until people were running web servers on it. His conclusion was that a successful language will be pulled into domains it wasn't designed for, so it should be designed for generality from the start. For Swift, that meant scaling down to embedded systems and up to something close to scripting.
Objective-C also shaped the design. It combined a Smalltalk-inspired object model with C for performance, which Lattner calls a beautiful combination because it offered both high-level APIs and a path to the metal. His insight was that "they didn't have to be two languages." A single language could be easier to teach, more memory-safe, and more modern, with good type inference and cleaner syntax. He notes that Python today has the same "two world problem," with a pleasant object layer on top of C, C++, and Rust.
How to approach a blank canvas
Asked how he designs a language from nothing, Lattner says most of his major projects have been blank-canvas work: LLVM, MLIR, Swift, and Mojo. His process starts with reflection on what success should look like, followed by a survey of what exists. He believes that even systems he dislikes contain good ideas that can be separated from the rest. His example is Java. He didn't like garbage collection or finalizers, but JIT compilation was powerful, and he calls Java a big step forward overall.
After that, he expects to design, redesign, and iterate. He validates specific assumptions one at a time, keeping most parts deliberately simple. He compares this to the "draw two circles, then draw the rest of the owl" meme, where the circles stay circles for now. He invests in proving one piece at a time, pins down its interfaces, and then moves on. Throughout, he says, you have to be "unafraid to go massively change and flip over the table and try again."
Four years of research and internal persuasion
After a year and a half, Lattner told management about Swift. He says their response was roughly "why would we want a new language? Objective-C is what made the iPhone successful." He was allowed to tell one or two people on his team, but no resources were committed. Because of his track record, he says, he was given "a little bit more rope" than others would get, and he started building demos.
About three years in, most of the feedback was still: if you don't like Objective-C, make Objective-C better, since you own it. Lattner calls this a sensible business reaction, because improving the language the whole community already uses is low-risk. So he did improve it, but in a way that moved Objective-C toward Swift. Objective-C memory management relied on manual retain and release, which he calls "a nightmare." His team invented ARC (automatic reference counting), which made the language safer and easier to teach. They then added modules and literals. Each feature moved everyday Objective-C programming closer to what Swift would be, so that Swift would feel like less of an abrupt leap.
He describes the four years as discovery, research, and "socialization process with executive leadership," not a fully staffed build. About three years in, by which time he was running Xcode and the developer tools organization of roughly 250 people, a large executive review took place. Lattner asked that bringing Swift to market become the department's focus. At that point Swift could not yet talk to iOS, lacked many object features, and had no apps built with it. The final year involved finishing the language and integrating it with the debugger, Xcode, code formatting, and the iOS SDK.
The 1.0 mistake and the break-your-code promise
Lattner says it was a mistake to call the 2014 release 1.0. Apple does not launch 0.5 releases, so it had to be 1.0. The team also needed real-world use, because at launch only about 250 people at Apple knew Swift existed, and a language can't be designed well in a vacuum behind an NDA. His lesson is not to call something 1.0 until it has been validated through real usage, and he says Modular applies that lesson to Mojo.
What went well, in his view, was communication. Apple told developers they could ship Swift apps to the App Store, but that their source code would break as the language changed. The host recalls Uber's 2016 app rewrite, done on Swift 1.2. The platform team took a significant risk on it, trusting that Apple was committed, and the host credits open-sourcing Swift with building trust. Swift 1 and Swift 2 broke source, and only Swift 3 declared stability. Lattner is glad of that: otherwise the language would still carry all of its early mistakes.
Asked why Xcode Playgrounds and a REPL shipped on day one, Lattner says every group wanted something from the new language. The documentation team wanted a book, so Swift launched with one. Management and marketing wanted interactive programming, which led to Playgrounds. Lattner wanted a REPL built on the debugger because a modern language should have one. The vision was right, he says, but the team was "massively overextended." Many of these features were still unbaked at 1.0 and 1.2, and waiting for a later version might have been better. He adds that such judgments are easy in hindsight.
Why experts resist new tools
Many developers rejected Swift outright. The host recalls a survey of Uber's iOS engineers in early 2016, two years after launch, that split roughly 50/50 between Objective-C and Swift, with the two camps "violently disagreeing."
Lattner's explanation is that experts don't like change. An expert Objective-C programmer has spent five or ten years learning the language's tricks and edge cases and is the person others go to for help. A new language resets that person to day one, the same as everyone else. Protecting that prior investment, he says, is "sensible human behavior." Some people are intellectually flexible or simply like new things, but many need a compelling reason to switch. For some of Swift's long-tail holdouts, that reason was SwiftUI, which arrived years later. He describes adoption as technology diffusion along an S-curve, from early adopters to people who have to be dragged along. He sees the same pattern playing out with Mojo and with AI.
The part of the Swift story Lattner is proudest of is how it widened access. After launch, he says, people stopped him to thank him. They had tried building apps in Objective-C but couldn't manage the pointers and square brackets, and their apps crashed. Swift made app development easy enough that they became app developers, and in some cases the experts at their companies. He sees GPUs today in the same position: CUDA, C++, and related tools, in his view, gatekeep many developers from this kind of computing. "I believe in the power of programmers," he says, and that belief is what drives his work.
Swift's technical debt, in Lattner's words
Lattner says he can't know whether Swift would be better or worse if he had stayed at Apple, so he talks only about his own mistakes. Swift was the first language he had built. Clang had implemented a language with a spec; Swift did not have one. "I didn't know what I was doing," he says. He adds that he is not "a math guy." He takes responsibility for Swift's "expression too complex to type check in reasonable amount of time" error, which he says comes from specific features and is very hard to fix.
His second mistake was adding features faster than the compiler's architecture could absorb them. Features worked 80 or 90 percent of the way but didn't fit together. As more were added, the mismatches compounded into complexity that developers can feel, which is why people now joke about Swift having too many keywords. He attributes this to prioritizing a fast path to app development over refactoring and getting the core right. He says C++ has the same problem.
He also gives an example Swift got right. In C++, int and float are hardcoded into the language, while complex numbers are library templates. That's because C++ inherited C's built-in types before templates existed and can't go back. Swift made int and float library types, which makes the whole system more uniform. His general rule: "make new mistakes, not old mistakes," fix mistakes where possible, and avoid painting yourself into a corner.
From Apple to AI: Tesla, Google, and SiFive
Lattner calls his Apple years "1.0" of his career: developer tools, CPUs, OpenCL, and GPU compilers. Around 2016 he "fell in love with AI" when the Photos app began telling cats from dogs. He owned developer tools and knew programming well, yet had no idea how to write an algorithm that detects a cat.
At Tesla, working on self-driving, he learned Caffe and TensorFlow. He concluded that TensorFlow could be much better, which led him to Google. There he owned the lower layers of TensorFlow (CPU, GPU, TPU, and several internal ASICs) and scaled the TPU software platform. His main lesson was how hard it is to build AI software without CUDA, Nvidia's roughly 20-year-old programming toolkit, which he says the entire AI ecosystem is heavily biased toward. A new accelerator like the TPU starts with no software at all, and TensorFlow and PyTorch had to be taught to use it. He credits his time at Google with teaching him the algorithms and frontier applications, noting that the "Attention Is All You Need" work was done on TPUs during that period. He built components he describes as now widespread, including MLIR, a compiler framework for domain-specific chips that he loosely calls "LLVM 2.0." He says it could not have been built without the mistakes learned from LLVM, and that it has been adopted by roughly every AI chip company.
At SiFive, which builds RISC-V hardware, he wanted to work on the hardware side of the boundary. He learned about physical design, verification, and the IP business model. When SiFive set out to build AI IP, he ran into the same gap he kept finding: "where's the software?" No end-to-end AI stack existed that a chipmaker could simply plug into.
Modular's thesis: an LLVM for AI
Lattner and his co-founder, whom he met at Google, started Modular to build "something like LLVM, but for AI." He argues that AI software today looks the way compilers looked before GCC, with every chipmaker building its own vertical stack: XLA for Google TPUs, Metal and MLX for Apple, ROCm for AMD. These stacks share very little code. The host notes that Anthropic is hiring separate CUDA, TPU, and Trainium kernel engineers. Lattner says companies end up writing the same models three times, and he cites an Anthropic engineering blog post describing the resulting bugs, quality problems, and outages when three parallel implementations drift apart.
Lattner says no existing project, including PyTorch and ONNX Runtime, was on the right path. At Google, he says, every year brought a new chip, while product teams such as Search, Ads, and YouTube needed the best engineers to fight fires on the current one. Nothing fundamentally new could be done if it took longer than a 6- or 12-month performance review cycle. In his view, research progress came from brilliant engineers "hacking the daylights out of" every layer of the stack, not from anything designed to scale, compose, or bring up new hardware quickly. Matching CUDA's 20 years of investment requires years of work and a very specialized team. A venture-funded startup, he says, was the only way to recruit people away from Apple, Nvidia, Meta, and Google.
Why kernels and compilers both fell short
Lattner describes two earlier approaches. TensorFlow and PyTorch, both dating to around 2015, put Python APIs on top of hand-written CUDA and Intel MKL kernels. Kernels were so easy to write that thousands accumulated, and all of them had to be rewritten for each new chip.
For TPUs, Google instead used compilers to generate kernels automatically, with benefits such as automatic fusion of adjacent kernels. The problem, according to Lattner, is that compiler engineers are scarce. When a new algorithm such as FlashAttention, a new sparsity pattern, or a new float format appeared, it couldn't be supported until compiler engineers found time for it. He says this bottleneck held back TPUs, and that the arrival of generative AI "really invalidated that technology approach."
Modular therefore bet on a two-level stack. Mojo, a new programming language, sits at the base, designed to provide full control over the hardware and "all of the performance," not a simplified model that gets most of it. Compiler techniques are then built into the language so that Mojo programmers get "the power of compilers without having to be a compiler engineer." Many more people can write code than can write compilers, and Lattner sees this as the way to address the talent shortage.
Predictability over "sufficiently smart" compilers
Lattner says compiler engineers, including himself around 2005, like to show off clever optimizations. That makes sense for benchmarks, where the source code is fixed and all the cleverness has to live in the compiler. His example of how this fails users is loop vectorization. It can make code four times faster using SIMD, but only when pattern matching and pointer analysis line up. Fix a bug in a way that breaks the pattern, and you get a sudden 4x slowdown that only a compiler engineer can diagnose. Report it, and you may be told you're "holding it wrong." The "sufficiently smart compiler," he says, "never works"; it becomes a leaky abstraction and makes the programming model unpredictable. In AI, where GPUs are expensive and frontier labs need peak performance, that unpredictability is unacceptable.
Mojo instead uses a simpler, more predictable compiler, which Lattner describes as more modern than what Swift and Rust are built on, and moves control into libraries. Instead of hoping for auto-vectorization, you call a library feature that vectorizes explicitly. SIMD, bfloat16, and compressed floating-point formats are native to the language. For productivity, Mojo has powerful compile-time metaprogramming, which Lattner says was "roughly stole[n]" from Zig and improved, with thanks to the Zig community. A vectorize function takes a higher-order function and stamps it out across vector lanes.
He contrasts this with C++, where templates and constexpr form a separate meta-language. It is nearly impossible to debug, can barely accept a string, and certainly can't take a binary tree as a template argument. Mojo, like Zig, unifies the program and the meta-program: ordinary code runs at compile time, so it can be debugged by running the same code at runtime. Mojo goes further by allowing heap-allocating structures such as lists and dictionaries to be built at compile time and embedded in the program. Lattner describes the resulting high-level algorithms as "mini compilers" that specialize on dimensions and constants.
Getting started: from slow Python to GPUs
Lattner deliberately avoids "the rabbit hole of really weird technical things only Chris understands," such as linear types. His pitch is that Mojo belongs to the Python family, so it's easy to learn, and AI coding tools make learning a new language easier than ever. The simplest entry point is speeding up an existing Python module. The usual route, rewriting hot code in Rust or C++, requires bindings that are mechanical and error-prone. Mojo, he says, extends Python without bindings, much as Swift and Objective-C interoperated natively, and it integrates with pip packages and build systems. Mojo also runs on CPUs, which he says are complex machines worth optimizing, so no special hardware is required.
His favorite example is a community member who introduced themselves as "merely a geneticist." This person moved Python DNA-sequencing code to Mojo and, after learning threads and vectorization, got roughly a 100x speedup. Having read that Mojo worked on GPUs, and without any GPU experience, they got the code running on a GPU in an afternoon and reported it was "a million times faster." Lattner contrasts this with CUDA, which he respects but describes as 20-year-old C++ that wasn't designed for modern architectures and "can't even acknowledge" tensor cores.
Mojo currently runs on AMD, Nvidia, and Apple GPUs. Modular's website offers "GPU puzzles" for learning, and Lattner invites open-source contributions. He argues that accelerators matter well beyond AI, citing bioinformatics, chemistry, oil and gas exploration, and HPC generally, but that the surrounding software is "really weird" and inconsistent across vendors. Making it consistent and teachable could bring a new generation into high-paying work.
Modular today
Modular turns four in January. Lattner says it is an unusually large and expensive startup, with just over 140 people. It supports seven architectures from three vendors: Nvidia Ampere, Hopper, and Blackwell; AMD's 300, 325, and 355 series plus consumer parts; and Apple GPUs, which are in beta. He hints that more is coming. Applying the Swift lesson, he wants Mojo 1.0 to be meaningful and to provide the stability Swift 1 didn't. He expects it in early summer of next year, with detailed dates still being scoped.
Commercially, Modular makes money mainly from a cloud platform, not from Mojo, which Lattner describes as something they had to build to scale across hardware. Enterprise customers want something that is easy to use, reliable, and scalable, and many want the option to buy the best chip for each workload from multiple vendors without rewriting their code three times. He predicts a lot of new silicon from many vendors in 2026 and says the company is still "at the beginning."
How Modular uses AI coding tools
Modular encourages AI tools such as Claude Code and Cursor. Lattner still writes code and uses Cursor as his daily driver. He estimates about a 10% productivity gain for him, mostly on mechanical rewrites. Even if the gain were smaller, he says, it increases his enjoyment. For prototypes and for PMs building wireframes, he calls AI "transformative," potentially a 10x gain. For production code he is unsure: he has seen agents "grind and grind" through huge numbers of tokens on tasks a person could have finished sooner, and he doesn't know how the gains and losses net out.
He insists that engineers not "turn their brains off." AI should assist humans, not replace them, and code must be reviewed and its architecture understood. Vibe coding in production "terrifies" him, less because of jobs than because of what happens six months later when the architecture needs to change and no one understands how anything works. He has seen AI tools duplicate logic in several places, which leads to bugs when only some copies are updated. The tools, he says, "still need adult supervision."
Hiring and the question of LLM-first languages
Modular hires deep specialists, such as compiler experts and GPU matrix-multiplication experts with ten years of experience, and also new graduates, who "haven't learned all the bad things yet." For early-career candidates, Lattner looks for hunger, intellectual curiosity, willingness to work hard, and fearlessness in a fast-changing field, instead of freezing up. Open-source contributions are his favorite signal because they show a person can work with a team. Because interviews make people nervous, Modular lets candidates use their normal tools, including AI for mechanical coding, since a whiteboard-only interview would be "very strange."
Asked whether languages should be designed for LLMs, Lattner rejects the idea. He hears the question more bluntly as "if AI writes all the code, why build a language?" His answer is that code has always been read more than it is written, and AI makes writing even cheaper. What matters is the intersection of expressivity (can you express the full power of the hardware? JavaScript, for example, never will for GPU kernels) and readability, the ability to understand code and build scalable abstractions. Assembly has full expressivity but not readability. Mojo keeps Python's widely known syntax and replaces "basically the entire implementation." He expects LLMs to keep improving at handling unfamiliar languages and finds them an excellent way to learn one. On making error messages better for agents, he asks what would be better for an agent that isn't also better for a human. The most important thing for AI-assisted coding, in his view, is a large open-source corpus. Modular has open-sourced roughly 700,000 lines of Mojo, with full history, which tools can index.
Why learn compilers
Lattner closes with why he loves compilers. In his university course, each project built on the last: a lexer, then a parser built on the lexer, then a type checker built on both. Mistakes had to be fixed because everything above depended on them. Unlike classes where you "build a thing, turn it in, throw it away," this resembled real software development. For newcomers, he points to LLVM's Kaleidoscope tutorial, which he wrote about 15 years ago; to the Rust community, which has many compiler enthusiasts and compiler projects; and to books and courses. He doesn't think everyone should become a compiler engineer, but he believes compilers "don't get the credit they deserve," that there are good jobs in the field, and that it needs more people.
And so I just started again nights and weekends. Didn't ask permission. Just started fiddling around seeing what could be done that would be better. The first year and a half it was literally just me working in nights and weekends and I had a day job and I was managing a big team and a lot of stuff going on.
Hold on. So at this point, yeah, you were already second level manager or something and had a team of 40 people running a lot of very interesting and very fun stuff.
You're good at juggling stuff.
This is a passion project. I'm also a fairly good programmer, so that helps.
How do you go from this blank canvas to like, okay, here's what I think the language will be?
I start from a lot of blank canvases across my career, right? So, LLVM blank canvas, MLIR later blank canvas, Mojo, Swift, many of the systems I build are blank canvas projects.
Do you think there is any logic in having a new language that is designed to make it easy for LLMs to code with? So, I often get asked given the AI is writing all the code, why are you building a programming language? Which is another way of asking the same question, maybe a little bit more aggro. And so to me, I don't think that optimizing for the LLM is the right thing to do at all because the important thing
Chris Lattner created some of the most influential programming language and compiler technologies of the past 20 years. He's the creator of LLVM, used by languages like Swift, Rust, and C++. Created the Swift programming language, worked on TensorFlow, and now works on the Mojo programming language. In this conversation, we cover the original story of LLVM and how Chris managed to convince Apple to move all major Apple dev tools over to support this new technology. How Chris created Swift at Apple, including how he worked on this new language in secrecy for a year and a half, and why Mojo is a language Chris expects to help build efficient AI programs easier and faster, how Chris uses AI tools, and what productivity improvements he sees as a very experienced programmer, and many more.
If you'd like to understand how a truly standout software engineer like Chris thinks and gets things done, and how he's designing a language that could be a very important part of AI engineering, then this episode is for you.
This podcast episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes to learn more about them and our other season sponsor.
So, Chris, welcome to the podcast.
Well, thank you for having me. It is so nice to meet you in person because I feel like the work that you did has had an impact personally on me at Uber. We migrated to Swift and of course a lot of the software we use runs on things that you've built. So can we rewind time back all the way to the early 2000s and for some of us who have not really been around there can you describe like what was the industry like in terms of compilers, languages and what was the status quo back then?
Yeah. So I mean I was involved in the kind of early days of Linux and so I started, maybe not the really early days, but I started working with Linux in the mid '90s and this is when Java was coming on the scene and a lot of things were changing in the industry as that played out. GCC was the compiler that standardized a lot of software and a lot of people were using it to build Linux-based software and this is how I got to know it.
GCC is a great thing. Before GCC, a lot of people don't know this, every hardware vendor was making their own compiler and it was a gigantic mess for even just C code because none of these compilers were compatible with each other.
And the reason they did that, just so I understand, is like you have C code, it needed to translate to assembly for a specific hardware and then the hardware vendor, you know, they did the mapping and figuring out like what custom instructions or something like that. Was that a reason?
Yeah. So basically back in particularly like the late 80s and then the early 90s there was a lot more diversity in terms of like instruction sets and chip makers were innovating and you had HP building things and you had Intel was making a lot of different systems back then and there's all kinds of different stuff going on and so you had RISC computers that were being invented.
And so the challenge at the time was that everybody had to build their own compiler and so because everybody had to build their own compiler they all wanted C and some C++ was emerging and C was a standard and so there was a spec that said this is what the code is supposed to look like but as we know for many standards and many specifications they're never complete and so what ended up happening is each of those compilers was a gigantic mess because they had different bugs, they had different misfeatures, they lacked certain capabilities and so software of the day was built with systems like autoconf which was this really weird macro processing thingy that tried to work around some of the limitations and it was a gigantic nightmare. So GCC came onto the scene kind of in the '90s really and cleaned up that mess. It became the standardized thing that people could plug into and a lot of chipmakers embraced it and it was free software and so that led to the rise of Linux. Linux adopted it and that really kind of brought the world forward.
Now GCC was a good thing. It really did help standardize the world and I think a lot of free software and open source pays a debt of gratitude towards GCC but also it was a really old thing and so I made fun of it at the time. It was over 20 years old and its architecture was, you know, very very old school in many different ways. It was also not built to be a modular design. It was designed to do one thing, take C code in and then put out assembly code for a given chip. And so there's a lot of things you couldn't do like JIT compilation and you couldn't at the time even optimize across files within one
And JIT is just in time compilation, right?
Just in time compilation. Exactly. And so there's a lot of things that people wanted to do that you couldn't do with GCC. And so in around 2000 that's when I was in university and I was studying compilers and I said oh okay well wouldn't it be interesting to build something new in the space and so I started working on LLVM. LLVM started out as a code generation system so you could target multiple different architectures but really for me it was a learning process. It was about saying okay, compilers are cool. Don't let anybody tell you otherwise, compilers are cool.
Even today, right?
Even today. And so but I didn't really understand anything about it and so I wanted to learn by building and so across my university project I built this thing up more and more and more. We then open sourced it. It got a little bit of a community. Later I went to Apple which really helped foster and fund a lot of the development and that early starting point was coming from, you know, GCC and open source technologies are really good but there's a lot of things we can't do. And so that's where I kind of fell down this rabbit hole.
And like when you started and open sourced, what did you open source? What could it do? And then, you know, what was the reaction to this early version of LLVM?
Oh, so LLVM back in the day, and this is super funny because many people look back on successful systems and they assume that everything was obvious when actually every step along the way is challenging and you have to earn any success you get. Very very rarely do you like luck into it. And so when I first started working on LLVM again it was a research project at a university. There's a lot of research projects at a university that don't go anywhere. And so
I'd argue most of them probably won't go anywhere.
Exactly. And the default assumption is that LLVM also would not go anywhere. That's a safe assumption for any vers.
And so we used it internally. And so my adviser Vikram Adve encouraged us to continue building it and then we taught it to a couple of classes. And so we had a few people using it internally and we got some like use cases with it. But at that point it was really just optimization and code generation. And so it plugged into GCC, used the parser to actually parse C code for example, and so it was just about code generation and it was useful for compiler people trying to learn things but it wasn't very useful for general application developers really.
When we open sourced it then we got I don't know two or three people that were interested in it and it was mostly other compiler nerds that, you know, are delightful, but there was no major community. There was no major reason for people to contribute and what I did was I said okay well actually just treat the world like the rest of the research group and just have open development, encourage people if they come by and they want to help and do something awesome. If they have questions, I'll answer them and just treat people with respect. And what happened over time is slowly the community grew. We got individual people that were interested in different things. Some people were interested in esoteric programming languages and they were interested in compiler technology for that reason. Other people were interested in performance and different people had different interests and so very very very slowly it kind of grew.
And what was the big difference between GCC and LLVM, right? Like of course modularity was one thing but also you mentioned capabilities that you wanted to do that GCC either couldn't do or was really difficult to do. What were those capabilities?
Yeah so originally it was just in time compilation, so runtime code generation.
That was a thing GCC
couldn't do? They could do only compilation
They could only do a kind of batch upfront traditional Unix-style code generation. At the time, GCC could not optimize across files in your project. And so if you had a function in one C file, it could not inline it into another C file. So things like that. LLVM was also written in C++, which was highly controversial at the time because GCC was all C and Richard Stallman was very opposed to C++.
And so actually there's a parallel universe because in 2005 I had joined Apple. We decided okay well let's see if maybe LLVM and GCC should merge and actually proposed to the GCC community, hey we've got this interesting technology, this seems very complementary, it could lift the architecture of GCC, maybe we should do this, and it didn't go anywhere because primarily it was written in C++ but it was also not invented here. There's a bunch of other, you know, very serious, very experienced people and they're like okay kid, why would we listen to you and so it did not go anywhere and thankfully so. It meant that LLVM had to grow up on its own. Yeah.
And then did LLVM really start to take off when you got to move to Apple to now work on it full-time? I assume there must have been at least a small team investing in this. Is that how it went or something?
Yeah. So the story of getting to Apple. So when I was graduating, I was looking for jobs in compilers and jobs that ideally would let me work on LLVM and continue the work because it was still an advanced research project. It was 5 years in, the community had grown quite a bit. I was doing regular releases every six months and so was really putting a lot of energy into advocating for the work and showing momentum and velocity. I caught the eye of a VP at Apple and there was a couple of people that were working on adding a PowerPC back end to LLVM because that's what Apple was doing. And so I got to know them and they said, "Hey, come, like we are frustrated with GCC. That is the foundation of all of our compiler technology"
because GCC compiled Objective-C for example, right.
That's right. Exactly. And it was again an older architecture. It was very difficult to work with. They weren't getting the performance they wanted. It was very difficult to add features. The community was also kind of annoyed with Apple for a variety of reasons. And so basically management said, "Come work on LLVM." And I said, "Yes, sign me up." And so when I started they did give me one other engineer to work with, but really they didn't give me a whole lot of guidance. And so I was like, "Okay,
You pretty much
cool. I will go implement PowerPC support and integrate with Mac OS back in the day and things like this." And then there became a time when they said, "Okay, cool. You're working on this." And my manager said, "It's great that you're working on this cool technology, but at some point we need to make sure that it's actually relevant to Apple."
Was it not relevant?
Well, it wasn't used in any products yet.
Oh, I see, you added support, but it wasn't used just yet.
Just yet. Yeah, exactly. And it wasn't as good as GCC in a number of different ways. And so I basically got the vibe that it's like, okay, well after a year if Apple's not using it in some product, then we'll ask you to start doing something else, because we're not just going to pay you to work on an open source project, it has to actually have impact on our company. And so when I got that memo, suddenly, and my manager helped me a lot, he was amazing, went around shopping for, okay, well, what is the near-term impact that we could have on the business?
And so from there, the first use case was actually for graphics. And so OpenGL was doing just in time compilation to do some graphics stuff. And so we were able to do something very small that actually had value. And suddenly, aha, this is not just a science project. This is actually interesting. Now it's still missing tons of features. But that enabled the development of the next set of features, the next set of features, the next set of features.
And it was now in a product that Apple actually used and shipped and it was, you know, generating business value if you want to put it like that.
A small amount but at least non-zero.
Non-zero.
Exactly. And so what I did over the course of many years across Apple is then say okay cool, we'll keep investing in the technology, keep building an open technology and an open team, so community, but make sure to deliver more and more and more and more value to Apple and as that happened, what ended up happening is I ended up replacing all of the developer tools, compiler, code generation, debugger technologies within Apple and that
One step at a time.
Exactly. And so it took about 5 years to the point where we built a new C++ parser, Clang, and Objective-C front end and all this kind of stuff. We got to about 2010, that's when Apple was releasing the first 64-bit iPhone. And suddenly this was made possible by LLVM, by all the technology we had built and all this kind of stuff. And it was an epic moment in the industry because everybody was convinced that 64-bit phones were stupid. That was not a thing. And 64 bits doesn't make sense. Why are you going to have more than 4 GB of RAM in a phone? Right. Hilarious.
Wow. Back in the day. But that's how people thought, right? Like
That's how they thought. And they're like, "It's inefficient. It can never work." And so actually the first 64-bit iPhone, it was the iPhone 5S, shipped in production before ARM got their chips back to test. And so we had built the entire software stack. We enabled the entire operating system, the entire tool chain, all this kind of stuff internally. And we were collaborating with ARM. But they had no idea how far ahead we were. And then suddenly we're shipping it in production and the whole world's heads explode because of how good the performance was and all the capabilities we enabled and it was quite fun.
So I'm having just trouble putting two things together and you can help me out with what I'm missing. On one end, you know, you made it seem like okay, like, you know, initially when you got to Apple for a year it was an experimental project, then
You just went slowly with one team and after the other, but now, next thing we know, fast forward to four or five years later, it's actually powering the core of Apple. How did it go from the small projects to actually getting into the core of the business? Did it help that you started with developer tools and more developers were convinced? How did you get through to probably the highest level architects and decision makers who, in the end, I assume, had to take a risk on technology that was somewhat new compared to GCC, which had been around for like 25 years at this point?
Well, so it's a combination of having top-down support, so people that wanted me to be successful, and that's because they were frustrated with GCC and so they wanted a new technology. So that was very helpful. A combination of bottom-up success, and it wasn't one thing. It was like every 6 months we'd have a new thing that worked well and go. I became known as somebody who'd get things done, and so I got more scope and responsibility, and so I went from being an engineer to a manager to a second-line manager to eventually a senior director over the course of a number of years. And so it wasn't one thing. It was just a lot of hard work, and it was a lot of fun because we were able to, you know, a great thing about Apple back in that time was that you could have a lot of impact because the team was growing, but also if you could get stuff done, you had an amazing platform in which to work because the iPhone came into existence and a whole bunch of other technologies were made possible.
And Objective-C was very Apple-specific, but we could move it. That was huge. We added a whole bunch of features to Objective-C. And so as we built into this, there's a lot of opportunity, but it was hard won, right? In any given moment it wasn't guaranteed. But as my team at Apple was making progress, we were doing all this in the open, then suddenly other people started seeing value. Cray built a supercomputer and used LLVM for it, and Google eventually adopted it in 2010 and started using it within Google for some small things. And then kind of a parallel project within Google, another hero within Google built a team and had a lot of impact at Google, therefore justifying more people to work on LLVM within Google, and different tooling and different projects and different things all happened. Eventually, you know, Intel or ARM or these companies ended up canceling their internal compilers and switching to LLVM.
And so then suddenly the vast majority of their engineering teams are all working on a thing, and then what you get is value that compounds, and today LLVM has a community of thousands of people that come together. And so it's amazing to see that.
It's incredible. It feels a little bit like when you're starting to roll snowballs and you're starting to roll this small thing, and then you do it for long enough, and most people don't do it, but then you could have a snowman and you're like, oh, where did you get all that snow from? Just incredible.
Well, so I mean it's funny because people always imagine success, right? And so you start from a starting point and you say, I want to be the president of the United States or something, right? And some people get very motivated by that, and it's very powerful. What I really love is doing all the work, taking each of the steps, doing the thankless thing, figuring out that really obscure thing that, if you get the architecture right, makes the whole thing compound and scale and compose the right way, and it allows you to solve problems that nobody else can solve. But what that means is most of the time I'm working on things people don't understand. And so I'm fine with that. I've normalized to this. This is what I expect at this point. I don't expect people to understand me. But what happens is that these projects over time grow, and they get to the point where suddenly it clicks, and suddenly people start to understand, and suddenly you can explain it because it maps into their value system in a way that they can measure.
Basically, Chris just mentioned how things click a lot more for people if it maps to their value system in a way that they can measure it. In my experience, measuring things is helpful all around, not just when trying to convince people but also when figuring out if a feature you built works as you expect. Here's a challenge most teams face. They ship features but cannot easily measure if they're working. Did that new checkout flow improve conversion? Is the feature helping retention or hurting it? Without measurement, you're just hoping things will work.
That's where Statsig comes in. Statsig is our presenting partner for the season, and they've challenged the status quo of how product teams work, much like how Chris challenged how compilers and programming languages are supposed to work. Most teams accept fragmented tools: feature flags in one system, analytics in another, experiments somewhere else. You're stitching together point solutions, running ETL jobs to sync user segments, hoping timestamps align across different systems. That's the status quo, but it doesn't have to be.
Statsig built a unified platform that gives you everything in one place: feature flags, experimentation, analytics, and session replay, all using the same user assignments and event tracking. So you ship a feature to 10% of users and the other 90% automatically become your control group. You can immediately see if conversion changed, drill down to where users drop off in your funnel, then watch session recordings to understand what went wrong, all with the same data. This is how you measure impact at the pace you ship.
Companies like Graphite, Notion, and Brex rely on this approach to make data-driven decisions without waiting for data teams or manually correlating information across tools. Statsig has a generous free tier to get started, and Pro pricing for teams starts at $150 per month. To learn more and get a 30-day enterprise trial, go to statsig.com/pragmatic.
With this, let's get back to the conversation with Chris.
One thing I feel that is probably pretty applicable for anyone is the fact that you were inside a company that was pretty big and successful at that time, Apple, and you still managed to get so many things done, just as you said, by solving hard technical problems continuously and figuring out how to help the business. And there was this interesting part where Clang got open sourced, and at Apple back then LLVM was open source, so that kind of, I understand,
Grandfathered in. Yeah.
But I want to ask you how you managed to convince a company that back then was pretty closed, and even today is pretty closed outside of a few things, to be okay open sourcing a new project. And I assume your reputation or hard work or getting things done might have been there, but what else did you use for leadership to green-light this? Because I'm sure top-down they needed to say yes, right?
Yeah. Well, so Clang was actually the easy one because we didn't really ask too many people for permission. We just kind of did it, and so that was the easy one. The hard one was Swift.
Let's get to—
So I think we'll get there. But that was the actual hard one.
And with that, let's get to Swift. The story of Swift, as you've shared many times, is how you started to work on it in secret. Can you tell me what led you to start experimenting with a new language? What you saw the problems being with Objective-C, especially because now we'd solved the compiler problem, the compilation was probably as good as it could have gone. And how did you pull off this kind of—what does secret mean, right?
Yeah. Well, so the backstory on Swift is that I joined Apple in 2005, and I built and basically replumbed their entire C++ and C and Objective-C toolchain, and by 2010 at WWDC, which is their summer conference, we had launched a full-stack replacement and it was working. And building a full-stack, fully integrated C++ compiler is no small feat. C++ even then was a very complicated language, and it was a very big technical challenge, but it was also very demotivating, at least to me, because C++ is just a beast of a language. And so you couldn't come out of that and not wonder, could we do something better, right? And as you said, we'd built a lot of this fundamental technology and could build pretty much anything we wanted at that point. And so I just started, again, nights and weekends, didn't ask permission, just started fiddling around, seeing what could be done that would be better.
And to me, what I did was I actually said, okay, let's go look at a lot of existing languages. So let's go look at both the new ones, let's look at Java, let's look at C++ things, but let's also look at functional languages like OCaml and Haskell and things like this. Let's also look at esoteric languages, the really weird other things. Let's look at TypeScript and Dart and other things that were kind of on the scene and happening, and Go and languages like that. And so it was really about saying, well, what do I like and how do I see success? And what I defined success as with Swift was making it easy to use and scalable.
And so I saw a lot of languages, JavaScript is probably the best example of this, where you start with something really simple, a weekend hack with JavaScript, and then it gets adoption, and then you get developers, and then developers bring their tools to other domains. And so JavaScript was designed for onclick handlers, but now people are running web servers in it.
Yep.
We see it all too much.
Exactly. And so my observation was that success meant that you would be naturally brought into other domains, and so think beyond just the first win and think towards more generality and scale. And what I realized is, if you take something like Objective-C, which I was deeply familiar with, for example, there was this dichotomy between objects and C. So Objective-C had this Smalltalk-inspired object model and then it had C for performance. It was actually a really beautiful combination because you can get all the way down to the metal with C, but you had high-level APIs. And what I realized is that they didn't have to be two languages.
So you could take the good parts
and pull it together into one system that could actually scale and would be way easier to teach, way more memory safe, way more modern, good type inference, fix some of the syntax problems, because Objective-C was a little bit of a mashup between two different worlds.
And the funny thing about that is that you fast forward all the way to today, that's exactly what Python is. Python is a beautiful language for objects built on top of C and C++ and Rust, and it's got this two-world problem going on within it. And so it's very interesting to me, again, how history rhymes with itself, and there's so many different things going on. But the initial ideas of Swift were really about saying, how do we build a scalable language and how do we pull together the best ideas, shamelessly borrowing ideas from different places.
And by scalable, as you meant, that means that it can be used for your specific task, let's say for iOS, but then it naturally lends itself to other domains.
Yeah, and specifically scaling all the way down to embedded systems development, all the way up to something that kind of approaches scripting, right? So super easy to use for very high-level programming.
How did you design a language, in the sense that as developers we use a lot of languages, right? I think all of us have tried and built with many different ones, especially with AI tools, it's so easy to try out. For myself, TypeScript, Rust, Go, Haskell, try out this and that, Prolog, and I'm used to using them, right? And I kind of discover it and figure out what is there. But when you have a blank canvas, or when you were thinking, how do you go about it? Do you kind of write down what you'd love to see? You also of course built compilers, so you start to immediately think of how that would work. Can you give us a little bit of a view of how you go from this blank canvas to, okay, here's what I think the language will be?
Well, so that's a good question for me in general, because I start from a lot of blank canvases across my career, right? So LLVM, blank canvas, right? MLIR later, blank canvas. Mojo, Swift, many of the systems I build are blank canvas projects. And so for me it's a couple of things. First of all, doing that kind of analysis, the reflection, deciding what success ultimately looks like. Then do the survey. Go look at what's out there. What's good? What's bad?
One of the things I really strongly believe is that you can take a system that you don't like, but buried within a system, even if you don't like it, there are often good ideas. And so you can take the ideas that are good and ignore all the rest. So, for example, Java: I didn't like garbage collection and finalizers, and there's all this stuff that came with Java, but being able to be JIT compiled was very powerful, right? And so there's a lot of things you can look at, and Java was a huge step forward for the technology space for a variety of reasons. But you can look at that and you can say, okay, well, I like that one aspect of it. All the rest of these things are separable, and maybe I don't like how they work. And do that survey.
And then when I start building a new system, I expect to design and redesign and iterate and learn and grow. And what I try to do is go through proof points and validate specific assumptions and say, okay, I'm not going to solve all the world's problems. I'll keep this part and this part and this part super simple. It's like the two circles that make your owl, just leave them as circles, and then invest in proving this piece. And if I can get this piece done and I can understand that it works and I can gain validation and confidence in that, I can know the interfaces going in and coming out of it well, now I can go out into the next system and build that out and build this thing out so that we can understand it and learn as we go. But every step along the way, being unafraid to go massively change and flip over the table and try again.
And so you started to experiment with Swift on weekends and on the side in 2010-ish. Why did it take until 2014 for a first release? What really goes into building out a language? Because again, I think we can all kind of empathize: this is cool, you have these ideas, of course you have a lot of experience, you try out things, but then, as a software engineer, I shouldn't ask this, but I'm going to ask: what took so long?
Oh, well, so it turns out building programming languages is not easy. It takes about four years, I guess.
Can you give us a bit of empathy for what makes it so hard?
Yeah. Well, so let me just frame up how it went, just so you know, because the first year and a half it was literally just me working on it.
Oh,
Okay. So nights and weekends, and I had a day job and I was managing a big team, and a lot of stuff going on.
Hold on. So at this point, yeah, you were already pretty—
Second-level manager or something, and had a team of 40 people and was running, yeah, a lot of very interesting and very fun stuff.
Well, okay. You're good at juggling stuff.
This is a passion project. Yeah. I'm also a fairly good programmer, so that helps. But after a year and a half, I got to the point where I told management
about it. I said, "Hey, I'm working on this thing. What do you think?" They're like, "Yeah. No. Why would we want a new language? Objective-C is what made the iPhone successful."
Right.
Right. And so I'm like, "No, what?" I'm like, "Well, how about I just have like one or two of the people in my team like work on this part-time?" They're like, "I guess you can tell them about it, but we can't commit resources really, you know?" [laughter] And so, but you know, again, I was respected and doing good things. And so I was given a little bit more rope than some other people would. And I got to the point of saying like, okay, let's build up some demos. And then it was about 3 years in and most of the feedback I was given was, okay, well, if you don't like Objective-C, go make Objective-C better. You own Objective-C. Also,
It's a natural reaction. I mean, as much as an engineer, I don't want to hear it. If I put a little bit of a business hat on, it's what you would
Makes total sense. That's what the entire community is using. It's very low risk. It makes your existing investment even bigger. And so what I did was I said, "Okay, well, across that four-year journey was not just building the thing full out with like a team of people that knew what they're doing. That's the discovery, the research, and also the socialization process with executive leadership." And so, what I realized, I said, "Okay, well, I want to have memory safety. Objective-C does not have memory safety."
That was one of the biggest things,
Retain and release, and it was completely manual, and it was just kind of a nightmare. And so we invented ARC, so a way of doing reference counting automatically. That made Objective-C way more teachable and safer. It also pulled it much closer to Swift. We then did modules. We then did literals. In Objective-C, we added all these features that what it was doing was pulling the lived experience of doing Objective-C programming closer to what Swift would be. So that ultimately when Swift came out, it would be less of an abrupt leap, and it was still a pretty big leap. But we're doing that. But meanwhile, you know, management's like, "Okay, cool. Why are you doing this?" and they're having their own issues. And we got to about 3 years in and said, "Okay, well, we think we should do this in a serious way." At that point, I had a team of 250 people-ish, something like that. I was running Xcode and the whole developer tools team
And still on the weekend, you were still doing this.
Well, yeah. You can look at my GitHub history, like you can go back to 2014 and see what I was doing. Yes. But at that time, right, they said there was this big executive review and a whole bunch of internal discussion. And I said basically, look, I want this to be the focus of the department to bring this into market.
Yeah.
And so, at that point, we didn't really have it talking to iOS, right? It was still missing a lot of the object features and things like this. We had no apps that were built with it. And so that last year was a massive amount of work in terms of finishing the language, but also integrating with the debugger, into Xcode, into code formatting and tons of different features, building more integration with the iOS SDK. That was a big deal and we had not done that. And so a lot of that made it possible to finally launch it four years later. And I would say when we launched it, so a mistake we made is we called it 1.0.
I was about to come to this because that's where I have like firsthand experience with 1.2. But yes, 1.0, you know, came out and, you know, the feedback, I can just talk from, you know, what I remember from there. There's two things. People were super excited that Apple is releasing a new language, and iPhone was huge. 2014 was the time where like you knew that you could build a big business on iPhone. Uber, I think Snap, all of these companies were just going up. It was a status symbol already and not everyone had iPhone just yet. However, when you tried it out, it was just not good. Xcode didn't have refactoring support. Debugging was a nightmare.
We should launch. [laughter]
So how did it go? Why did you decide at the time, like, you know, I'm just interested in your thought process at the time on how we went about it, and, you know, what were some learnings now that you look back?
Yeah. So it had to be 1.0 to launch it because Apple doesn't launch 0.5s, basically. So if we wanted to launch, it had to be called 1.0. So that was one aspect of it. But the other aspect is we wanted to get it out there. And at the point that we launched it, only about 250 people at Apple knew about it at all. Most of those people were in my team and then some executives in marketing and other people like this. And so we knew we couldn't build a good programming language in a vacuum with an NDA that you had to sign to get access to it, right? You have to actually get it out there. You have to get usage experience. And so the learning for me is don't call something 1.0 until it's actually ready and it's validated and you have a lot of usage experience and things like this. We bring a lot of these lessons to Mojo, by the way.
Chris Lattner and the team launched Swift back in 2014 and it's since become the de facto way of building native iOS apps, which brings us nicely to our season sponsor Linear. Linear's iOS app is also built on Swift. It's a fully native app. It's not built with React Native or using a web wrapper. And there's good reason for this. Linear's entire philosophy is about speed and performance. The desktop app is famously fast, but that same obsession about speed carries through to mobile. Linear have just redesigned their mobile app with something pretty interesting: their own take on Apple's Liquid Glass design language. But instead of just using Apple's API under the hood, they rebuilt the UI from scratch. The Linear team wanted more control over the navigation and customization than Apple's system would allow. Classic engineering mindset. If the existing solution doesn't give you what you need, build your own.
Here's my favorite detail about the app, though. Linear has this feature called polls. It's basically a feed that gives you an update on all your development work, project updates, what's on track, what's blocked, etc. On mobile, you can actually listen to it as an audio briefing. The Linear team told me how they get raving feedback from engineering leaders who listen to their polls update on their commute. It's like a personalized podcast about the team's work. One VP shared how he gets through his entire standup prep while driving to the office using polls. This is why I love partnering with engineering teams like Linear's. They're not just porting desktop features to mobile. They're actually thinking about how people use their phones differently. If you want to see what they're building, check out linear.app/pragmatic. You can use the mobile app to manage issue tracking on the go, and that Liquid Glass implementation is genuinely beautiful engineering work. And now let's get back to Swift 1.0 and how the language was pretty painful to work with on launch back in 2014.
But also it's about making sure that we communicate clearly with the community. And so one of the things I think we did well is we said, look, internally to Apple, we said, look, we need to launch this, we know it won't be perfect, and so we're going to tell the community that you can adopt it and you can build apps and submit them to the store, but we will break your source code.
Yep.
And this is because we know that we have to change the language when we get usage experience. And so we told people, like, we will help you, but expect code breakage. And I'm glad we did that.
I remember that that was really clear. People were upset about it, but they knew. Like when we had the discussion of should we move over to Swift or not at Uber, the rewrite happened in 2016 where Swift was at 1.2, and the platform team made the decision after evaluation that they will take a risk, and it was a big risk, and there's now stories about it. I have a blog post, there are some other people with their experiences, but they did it because they knew that Apple is committed. I think the open source part also helped build a lot of trust, and I almost feel that that was one of the last times that I saw the Apple community really unite, like iOS developers specifically, just go like, you know what, we'll be in the pain together. We know it's going to be painful, but all mobile engineers I know knew that this was pain. They told their management and all that. They said this is the right thing to do. And of course I think it really helped that Apple had this trajectory and it was like, okay, well, they didn't want to be left behind. They didn't want to risk not having access maybe in the future one day to APIs. Not at that point.
One thing I was surprised by when Swift was launched, I don't know if this was right on launch day, but Xcode Playgrounds and the REPL, the read-eval-print loop, was launch day, was also part of launch day. And, you know, for those that don't know Xcode Playgrounds, it was like you could just try Swift inside Xcode, and with the REPL, you could also just really easily play. Why did you decide that that needed to be there on launch?
Well, so part of that came from that year-long process of getting it to launch. And so part of it was, okay, if we're going to do a new language, what are we going to get out of it? What is the developer experience? Everybody wanted to get something, right? And so, for example, the documentation team, who I love, said this is our chance to write a book, which was awesome because Swift had a book on launch and it was fantastic. The Xcode team said, okay, well, and, you know, management and marketing people said, okay, well, can we get interactive programming? This is something we've never had, and so we want that. And so that's where Playgrounds came from. And so the REPL came from me saying, okay, well, modern languages should have a REPL. This is a thing that builds into the debugger, and there's a bunch of different capabilities we could use, and it would make a lot of sense to make it much more modern feeling. And if we're going to do this, let's do it right. And so there's a lot of that: okay, if you're going to do it, let's get the advantages from it.
Now the flip side of it is we were massively overextended. And so a lot of the vision was right, but particularly at 1.0 and even 1.2, many of these things were still unbaked and it took a long time. And I think that we underestimated how hard some of these things were, and waiting for 2.0 or 3.0, you know, maybe it would have been better for some of this. But anyways, it's easy to be negative and judgmental in hindsight, right? At the time, we were all working really hard and trying to make something amazing. And, you know, we were learning as we went. And I was really thankful for the community. But one of the other things I'll say is that while you may have been on the "oh wow, awesome thing, and let's adopt and let's get ahead of the curve" side, there are just as many people that were the "what are you talking about? Objective-C is great. If it ain't broke, don't fix it." And they're like shaking their fists and saying like, I will never use a new thing, and blah blah blah, and there's all these bad things which they write about with Swift. And so it took quite a long time for the community center of gravity to move, and
Well, this was a fun fact. In 2016, I remember the middle of 2016, or maybe a little bit earlier, Uber was one of the big mobile-first companies, built on iOS, grew up on iOS, lived and breathed iOS, and of course Android. And the mobile platform team ran a survey sometime in early 2016, 2 years after Swift was out. This is Swift 1.2, saying we are planning to rewrite the app. The rewrite was a thing
Just because they needed to for architecture.
We needed to redo every single screen, every single interaction, which kind of meant a rewrite of all the business logic. So it was going to be a rewrite no matter what. And they ran a survey: would you prefer, you know, as an iOS engineer, Objective-C or Swift? And the results were pretty much 50/50. And these were again very experienced engineers. They knew, and the two camps were violently disagreeing with one another. And I still don't know why Uber went with Swift. Someone was a tiebreaker in the end. But back then this is how foot on foot it was. And of course, you know, there was grumbling, there were problems, so the Objective-C people said, like, told you so, we could have just done
There was at some point a thinking of like maybe we should have gone, but as you say, it was not trivial. But I guess this is what you don't see, like, in hindsight.
Well, so, but I've learned a lot, and again you reflect forward because this impacts exactly what I'm working on now. Same thing, which is that
With Mojo.
Experts don't like change.
Right, what I got
This is counterintuitive.
Well, I mean, if you think about it, if you're an expert Objective-C programmer, I don't know if you were, but if you're an expert Objective-C programmer, you've been doing it for five or 10 years, you've learned all the baggage, tricks, all the weird cases, you know all the arcane things. You are the person that people turn to for help. Swift comes out, now your expertise is invalidated. You're on day one just like everybody else. You are not the expert. You don't really want a new thing. You want to be the king of the hill of the old thing. And you don't want the new thing to be successful. And so now some people are intellectually flexible and will switch over, and there are early adopters and things like this, or people that like new things. And so it's not true of everybody, but it is actually sensible human behavior to protect your prior investment. And for a lot of people, you have to have a really good reason to switch over. And so in the case of Swift, you know, SwiftUI was like a reason that moved some of the long-tail people,
Which was years later,
Years later, right? So there's different inflection points. And so what you look at is it's almost technology diffusion, right? And so there are early adopters and they'll get on board quickly just because it's a shiny new thing.
Yeah.
And then there are some people you have to drag Objective-C out of their dead hands, right? And so everybody else gets mapped somewhere along this curve, and it kind of looks like an S-curve.
But this is interesting because, you know, we're going to get to some of the AI and LLMs, but I feel this playing out.
That's right. That's exactly the same thing, and it's part of human behavior. It's not about Objective-C. Now let me tell you the thing that I'm most proud of in this journey. So when I started in 2010, it was just a toy project, playing around with in spare time, not expecting it to go anywhere. I was kind of implicitly coming from the assumption that we could do something that was better than C++ and Objective-C and that was more easy to use and things like this, but didn't really know how it would go. After we launched, I would have people that stopped me in the streets and said, "Thank you, Chris. Because of Swift, and not just me, but the whole team working on this, I was able to become an app developer. Before Swift, I tried building apps with Objective-C, but I could never figure out the pointers and the square brackets and it would crash, and I'm not smart enough to build an app."
"But now with Swift, it became easy enough that even I could do it. And now I am an app developer instead of a web developer. And I've become the expert in my company. And now I have learned and grown in my career." And so now, thanks to this new technology, I progressed. And if not for this enabling technology,
it never would have happened. And that's the exact same thing happening with GPUs. Right now we have the exact same thing with CUDA and C++ and all these tools that are gatekeeping so many developers away from this emerging kind of compute and kind of technology. And it's the same exact thing.
And so this is what really makes me passionate about working on this kind of technology, because I believe in the power of programmers. I believe in the human potential of people that want to create things. And that's fundamentally why I love software, is that you can create anything that you can imagine. And to me, that's what's so powerful about working in this space.
I really feel it and see it. And just looking back on Swift, because you created it. It's your baby. It started as a project. You were a guardian of it or shepherd or however you want to call it for a while, and of course now you stepped away. It has its own life. Looking back at it, how do you feel on how the language has evolved? What are parts that you're really kind of happy on how it lived up? What are parts where you're thinking that maybe if you would have stayed longer, maybe you would have tried to have it in different ways, if there's any?
So, it's very hard for me to say if I'd stayed at Apple and if I had kept my fists on and controlled it, if it would have been better or worse, but I can talk to the mistakes that I made, right? Because they're learnings that I could bring forward, and that way it's my fault and it's not me saying other people weren't able to do it. Right.
Fair. I like that.
But early Swift, I'll just say it bluntly, I didn't know what I was doing, right? So I'd never done it before. I'd built a C++ compiler which had a spec.
Well, it was the first language you ever built.
It was the first language I had ever built, right? And so a lot of the ideas going into it were new. A lot of the type-checking algorithms were exciting, but I didn't really know what I was doing. I'm also not a math guy, and so if I actually understood math, then maybe it would have been obvious. But things in Swift like "expression too complex to type check in reasonable amount of time", that's my fault, right? And there's specific features that cause that to happen, and it's very difficult to fix. And so that's one mistake.
Another mistake is that we were very ambitious in terms of adding features that we wanted to see in the language, and we added them faster than the architecture could keep up. And so particularly if you talk about Swift 1, but even today, what ended up happening is the internals of the compiler and the design didn't really line up the right way. And so you get a lot of things that work kind of 80% or 90%, and they didn't quite fit together, and then as you start adding more and more and more stuff, it doesn't really line up right and it just gets more and more complicated. C++ has the same thing going, by the way. Swift's not unique this way.
But as a consequence of that, you get an emergent complexity that then developers can feel, and Swift today, people make fun of too many keywords and things like this. Well, it's because the origin story was really not, in my opinion, didn't put enough energy into really refactoring, really redesigning, really making sure the core was great. Instead, it was okay, let's get to app development as quickly as possible. And so, as a consequence of that, what happened is then, I and many other good people put more stuff onto it and more stuff onto it and more stuff onto it, and the language got more complicated. And so, partially that's maybe in any individual decision, but also it's because some of those original things didn't line up the right way, and so the complexity kind of compounded.
Wow. I'm just hearing we have a lot of tech debt inside a language and it builds up.
Yeah. Well, I mean, if you
Well, I mean, this is just one way of putting it. Of course, you put it differently, but it just feels this all repeats itself. It's whatever project you do.
Absolutely. Well, let me talk about a simple example in C++. In C++, int is hardcoded into the language.
Yep.
Float is hardcoded into the language. Complex numbers are not. They're templates. Why is that? Why don't you just get rid of these concepts, have one fewer thing, and just make int and float part of the library, and then everything's more self-similar, and everything will work the same way, etc. And so, Swift fixed that problem, right? So int and float are now part of the library, and all the stuff is much simpler in that way, right? But C++, you can't go back and fix it, because it comes from C, and so they added templates after C, and so all the C stuff had to be hardcoded in. They couldn't fix that. And that's the origin story of why C++ works the way it does.
When I was talking with Steve McConnell on the podcast, he told me that when he worked at Microsoft for a while, the best developer he met there did this thing where he wrote everything three times. First, he wrote a first version in like a week and a half, then a second version in 3 days, and a third version in a day. And when I think back of the history of Python that I just heard, there's a documentary where we can listen to it, it was also the second or third iteration of a language. And what I'm also hearing now is also, in Swift, you saw some mistakes in C++, but you couldn't for some of the things, but now we're going to get to Mojo as well.
Well, and that's always my goal, right? The goal is, so you're always going to make mistakes in life, just so there's things you can do. You can decide to make new mistakes. So learn from the past and make new mistakes, not old mistakes. And then you can make the decision to fix mistakes wherever possible. And then you try not to paint yourself into a corner.
And so again, a thing that I think Swift did well is it said, "Okay, well, the initial release was 1.0. That was predestined, but we're going to break the language." And so Swift 1, Swift 2, ultimately to Swift 3 broke the language to try and make it better. And then only at Swift 3 did we say, "Okay, now it's stable," right? And I think that was really good, because can you imagine if it was still Swift 1,
But we'd be left with all these mistakes,
All the mistakes, right? And so it was painful, but I think that was a good set of decisions that we made, and it led to a much better result.
Yeah, there's a lot of trade-offs. So, as we get to what you're doing today, after Apple, you worked at a few interesting places: Tesla, Google Brain, SiFive. Can you talk about what you've learned at each of these places and kind of how it's, because I think it'll be interesting to look back on how it just all is leading to what you're doing now.
Yeah, totally. So the way I frame it up is that 1.0 of my career was at Apple and building developer tools and understanding CPUs, and I built OpenCL, which is a GPU thing also, and compilers for the GPU and all this kind of stuff. And so learning a lot about that hardware-software boundary and developers and developer tools and Xcode and that whole thing, right?
2016 I fell in love with AI. And the reason was that about 2016 the Photos app had the ability to tell you, is this a cat or is this a dog? And it was Inception v1 or some really ancient AI model in there. I'm like, I own developer tools. I know a lot about programming. I have no idea how to write an algorithm to detect a cat in a picture. How the heck does that work? So I fell down this rabbit hole in 2016, and this is what leads me to 2.0 of my journey, which is a very AI-focused, go figure all this stuff out.
And so Tesla is doing applied with self-driving cars. That's where I learned Caffe and TensorFlow, and I realized TensorFlow was a good thing, but it could be a way better thing, and so that led me to joining Google and helping scale TPUs there. I ended up owning the bottom part of TensorFlow, all the CPU, GPU, TPU, plus a bunch of other weird internal ASICs within TensorFlow, but notably I scaled the software platform for TPUs, and across years learned both, hey, it's actually really hard to make AI software without CUDA. TPUs don't have CUDA,
CUDA, which is Nvidia's way to program kernels, right?
Yep. So basically the entire AI ecosystem back then, but also today, ends up being very heavily biased by Nvidia and by their API and software and language toolkit called CUDA.
Mhm. Which is like 20 years old or so?
It's about 20 years old, and it's awesome and it's very mature and there's a lot of good things about it, but it's also 20 years old, and so there's time and space for new things. But when you're building an ASIC, so TPUs are a very large-scale data center training and inference accelerator, very fancy, very frontier, particularly back in 2017, you don't have software, and so you have to create everything from scratch, and then you have to get TensorFlow and PyTorch to talk to it, and nobody really understood how that worked.
And so across the years I learned so much, and I'm so thankful for my experience at Google, because I learned about the algorithms, I learned about AI, the frontier applications, like the "Attention Is All You Need" paper was invented in that time frame on TPUs, and so we
Wow.
Yes, and so tremendous amount of things going on. Large-scale embedding systems for ads and recommenders and stuff like this was going on. It was a brilliant time. And so across that, kind of an early version of what I'm doing now, said okay, well, we need to reinvest in the software technology and the platform technologies, and so I built components of what is now, I think, pretty widespread AI software. So there's this MLIR framework, new runtime technologies, a whole bunch of different things. So each of these things ended up being kind of 2.0 of the first leg of my journey, and so MLIR is a compiler framework. It's kind of like LLVM 2.0. And so it's a rough analogy, and we probably won't go into MLIR full details, but
But it's for ML, right?
Well, it's for domain-specific chips and specific chips. Yep.
Yep.
And so, but I could not have built that without the experience of building LLVM, right? And so having gone through that experience and being willing to do a 2.0, a clean slate, looking into the abyss, you're standing on a cliff and it's just blackness around you and you could do anything, but you have to decide where to start, right? That really helped, having gone through that and having made a bunch of mistakes with LLVM and being able to then fix them. And so we're able to build something that has now been widely adopted by roughly every AI company or every AI chip company that exists, which is very exciting.
Wow.
And through that journey, SiFive for example then said
And SiFive built specific, like built CPUs, right, on RISC, one of the best in the world as I understand.
Yes. So SiFive builds RISC-V hardware.
RISC-V hardware.
Yeah. And so I decided, okay, I've always been on just the software side of the hardware-software boundary. How about we go play on the other side? That sounds fun. And so I learned a tremendous amount about physical design and chip design and the business model of hardware and IPs and all the different technology for verification and things like this, and tremendous amount of fun. And one of the things we decided to do is, let's go build an AI chip or an AI IP. And so you get to that and then suddenly you get to, where's the software?
And so across each of these legs of my journey, I'm like, okay, where is the software? Where's the software? Where's the software? And in the case of AI, there's basically nothing there. There's nothing that you could take off the shelf that provides an end-to-end AI solution that you can then just go implement your chip. And so, I mean, the origin story of Modular really came down to my co-founder and I, who were buddies as we got to know each other previously at Google, but really saying, we need something like LLVM, but for AI,
But for AI,
Like, we need the thing that people can implement support for their chip into, and then get an AI software stack. Because if you're building CPUs 20 years ago, you didn't want to have to build a whole C++ compiler and C compiler and code generator and kernel and web browser. You just want to do a little bit of work and then get software. And that doesn't exist for AI right now.
And then for those of us who are software engineers but we don't build AI software, we just use the models, we might invoke the API or we might invoke the trained one. Can you explain to us, if we moved into the world of actually building AI applications, running them directly on GPUs, may that be Nvidia, or now there's AMD, or Google TPUs, what software would we use today? What is the status quo, or what was at least the status quo when you started Modular, and then what is your vision, which I sense there's a bunch of similarities with GCC and LLVM here.
Yep. Absolutely. So before GCC, every chip maker had to build their own software stack. That's literally the shape of AI software today, because there is no thing to plug into besides Modular. Everybody has to build their own vertical software stack. And so when I was at Google, we built a thing called XLA. That's the thingy that we built, kind of like CUDA but for Google TPUs.
Google TPUs.
If you look at Apple, they have Metal and MLX and their whole stack. If you go look at AMD, they have ROCm, their whole stack. Everybody has to have their own verticalized stack, and they share very low code.
Oh, so this is why I was just looking at, I talked with some folks at Anthropic and looked at their job sites, and I saw that they're hiring both CUDA engineers but also Google TPU kernel engineers
And Trainium, which is the AWS
And Trainium,
And so they're writing the same models three times over and over again
depending on the hardware they do, and then I guess when it's a new version. Okay, I get it. So this is where we are today, or where we were at least.
Well, that, and so you mentioned Anthropic, they have an amazing engineering blog post from a few weeks ago. I can remind you and I can dig up the link.
Yeah, we'll put it in the show notes.
It talks about the problem of reimplementing the same models three times over.
Turns out they don't work right. You have bugs. You have quality problems that then users can feel in your product. You have outages, because doing things three times in parallel with three different teams, of course, they don't line up.
So where Modular came from was saying, okay, well, I've lived this experience with Google with TPUs, and a number of other chips by the way, but TPUs are the public one. I have seen this at SiFive building into the hardware ecosystem. At that point, this was about four years ago at this point, the AI hardware ecosystem wasn't mature, but suddenly I had built up this wealth of knowledge of how to do this stuff. And when I looked around at all these projects, I looked at PyTorch, I looked at ONNX Runtime, and there's a million of these things out there, I didn't think anybody was on the right path. Nobody was willing to do the work.
Because what I saw is I saw the problems I faced at Google. I love my experience at Google. I learned a tremendous amount; the people are amazingly brilliant. But the problem that I had, and that we had as a team, was that every year there'd be a new chip, and every year you're scrambling to get the next chip to work. At the same time, all the product teams are using the current chip. And so all your best people are getting assigned to put out the burning fire for search or ads or YouTube, whatever the
workload is. And so you could never actually get the mindset or the bandwidth to be able to do something fundamentally new, particularly if it took more than a 6 or 12 month performance review cycle.
Oh yeah, Google is infamous for that.
And so what I saw is I saw a lot of these very short-term projects and a lot of improvements and a lot of micro optimizations and a lot of great work, a lot of massive innovation in models and the architectures, but the progress, the research progress was happening because Google had infinitely smart engineers that had built every level of the stack and could hack the daylights out of it just to make simple things work. It wasn't scalable. It wasn't beautiful. It wasn't built to compose and it certainly wasn't built to bring up new hardware quickly. And so, as you know, I've been working in compilers and dev tools and languages and stuff for quite a long time. And so, I said, "Okay, well, I think it's finally time to solve this problem."
But to solve it, you can't solve it in a weekend with like three people. You need years of time and a deep, very high-tech team to be able to invest in this technology because CUDA, as you mentioned, is 20 years old. It's had thousands of people work on it. Huge investment. Any one of these chips and these algorithms is extremely complicated. The AI research landscape moves really fast. You need a huge amount of effort and the only way to tackle this was with a VC-funded startup because that's the only way we could pull people out of Apple and Nvidia and Meta and Google and all these companies to be able to go build this.
And you said something interesting in a previous podcast or interview. You said between kernel engineers, compiler engineers, and AI engineers, they just kind of think differently and there's not much overlap. Can you tell me a little bit about, like, because you have been all three. You of course know so many of them, but what is the difference in the end? You know, they all need to work together at a place like this, but you said that somehow they don't understand each other.
Yeah. So, I'll tell you the aha moment that came to our technology approach for Modular. Again, if you go back four years ago, you're talking 2021. Funny how that works. Okay, so late pandemic, a lot of amazing stuff had been done and a lot of people had built, first of all, TensorFlow and PyTorch. I consider them to be some of the very early production frameworks. TensorFlow, PyTorch both come from 2015, by the way. So they're nearly 10 years old, which is ancient in AI time, right? And so they were built under this idea of we will take CUDA kernels and we'll take Intel MKL kernels and then we'll build Python APIs on top of these.
And so one of the challenges with that generation system is that it was so easy to write kernels that people made thousands and thousands of kernels, but then they couldn't be ported and you had to rewrite thousands and thousands of kernels to be able to bring up a new chip. When we worked on TPUs we said, aha, you know what's cool? Compilers. And so instead of writing thousands and thousands of kernels, what we can do is we can use compilers to generate the code and automatically synthesize kernels on the fly. And this was a huge thing. And we said, hey, look, we can get performance. We can get things like autofusion where you merge two kernels together. And compilers are a lot of fun, by the way. I encourage people to learn them. And so that was great fun.
Now the problem with this is that where you can't scale kernels by writing thousands of kernels every time you bring up a new chip. Turns out AI moves really fast. Turns out compiler engineers are lovely. You know, I love many of them. But there's very few of them. There's very few of them. And so what you ended up with is you ended up with a bunch of these technologies where it was cool because you could support, you know, some compiler for some chip, but suddenly you couldn't do flash attention or some new algorithm comes on the scene. It's a massive breakthrough in research, but the compiler people couldn't be in the loop to implement the new sparsity or the new float format or the new this or that. And so having compiler engineers in the loop was a huge problem and it held back, for example, Google TPUs. A lot of what happened, particularly when gen AI came on the scene, really invalidated that technology approach.
So I believed, and so going into this, and also part of the learning from working on TPUs is that we needed programmability. We need full power over the hardware. We didn't need the watered down simple solution that made it so that you got most of the performance. What I believe in is getting all of the performance. And so
Otherwise engineers would just write kernels as they
Exactly. Exactly. And there's so many technologies, you see this across software, where it's like, okay, it works up until a point and then you have to switch to a different system or a different stack, and I don't like those things. Go back to Swift. I wanted to be able to scale from embedded all the way up to application level. Actually unify this. Don't make it so that you have fragmentation in different systems. And so what we bet on was we bet on a two-level stack. We bet on Mojo
A programming language,
A brand new programming language. And as we all know, making a new programming language is doomed and never works and you cannot get it adopted, all those things. Well,
You have a good track record though.
Yeah, it's actually just a hard problem. But you have to be smart about it and we can talk about that. And then on top of that saying, let's take this programming language and design it so it can solve these problems and scale across hardware, etc. But then take these compiler algorithms and put them into the language so that you as a Mojo programmer get the power of compilers without having to be a compiler engineer. And so a lot of what makes Mojo really special is it takes power out of the compiler and gives it to you as a developer. Because it turns out a lot of people can write code, they can't necessarily write compilers. And so what this does is it breaks open the talent problem. And so now we can have a lot more people building these very highly adaptive, highly scalable, highly composable systems in ways that they couldn't do before.
So can you tell me about both Mojo, but also how this language can actually break the compiler problem? That part, you know, the new language I think we can all understand, but how you can get through to compilers and get the performance out of it. How does the language do that, or how does the stuff around it or underneath it do it?
Well, also, again, this is my life's work, my journey, and so I've made a lot of mistakes along the way. And so a lot of it, as we talked about before, when you don't know what you're doing, go learn from other people and go decide what they've done right and wrong and what you like. For me, a lot of it was learn from what I had done wrong and make sure to do it better the second time. And so one of the things that classical compiler engineers and myself back in 2005 or something, which is now 20 years ago somehow. Yeah. Was is that compiler engineers want to show how awesome their compiler can be. Let me show you. I can implement this optimization. I can detect this pattern. In this case, I can make this code go, you know, two times faster or five times faster, 10 times faster.
The challenge with that approach, and this is actually really important for benchmarks, because if you think about it, if you're a compiler engineer, you generally can't change the source code. Like, you know, if you're optimizing for the SPEC benchmark suite or something. Yeah. The source code's fixed, and so all you can do is implement crazy stuff within the compiler, matching and transformations, and maybe change the memory layout of something using some heuristic. But there's a problem with that as a user. And so if you're a programmer and you're using one of these compilers and say you're working with a loop vectorizer, for example. Loop vectorization can make your code go four times faster by using SIMD to take advantage of the hardware, right? But it requires a lot of pattern matching on your code and it has analysis of how you use pointers and all kinds of other stuff that all has to line up for the optimization to work. So now you can have this wonderful moment where you pick up one of these magic compilers and then you get amazing performance and you're like, "Wow, this is cool. It goes four times faster. I love this thing. This is great." But then you go and fix a bug. If you fix a bug and you break the hack in the compiler, well, suddenly you get a four times slowdown.
And you're not sure why. And you cannot really debug it unless you're a compiler engineer.
Exactly. And so now you file a bug and again there's not enough compiler engineers in the world and so they may or may not look at it or they may tell you you're holding it wrong and too bad, it has to be this way. Right. But what that experience is, is that's caused by compilers that are trying to be sufficiently smart. And the sufficiently smart compiler never works. What it ends up being is an abstraction that's leaky. Leaky abstractions exist in lots of software and lots of different domains. But when they exist in a language, it means that you have an unpredictable programming model. So fast forward to today in AI. Today in AI, basically people want peak performance because GPUs are really expensive. As you said, if you build a system that doesn't deliver performance, they'll switch to a different system because they have to get that performance, right? Particularly for the really large scale frontier labs and things like this.
So they need to squeeze everything out of it.
They need to, it's required. The cost is just astronomical otherwise. Well, you can't have a system that is unpredictable or flaky or weird. You want something that has control. And so what Mojo does, unlike Swift and many LLVM languages, Rust, all these guys build on top of the same kind of technologies, we go for a much simpler compiler. It's way more modern. It's way fancier than what Swift or Rust or languages like that are built on, for a variety of reasons. But it's way more predictable and so it puts control in your hands. And so instead of, for example, vectorization, instead of hoping to magically do it for you and maybe it works out, maybe it doesn't, it gives you library features where you can say, "Hey, vectorize this for me."
And I saw, on Mojo, of course, there's a tutorial and getting started. I saw there's these special characters/keywords where if you want to, I mean, you can just write code that looks like Python. It's a subset of it and you're like, whatever, but then you can kind of take control, right? So when you're coding, when you decide you want a certain compiler feature that it supports, can you give examples of what I can do with it?
Yeah. So for example, take SIMD. SIMD, so vectors. I don't think this is even a really standardized feature today in most programming languages, but it's essential to get performance out of a CPU, right? Most of the operations and floating-point operations in a CPU are in SIMD because they allow parallelism. So Mojo makes that very native. Mojo also makes things like bfloat16 and compressed floating-point formats and stuff like that very native. And so it exposes that to you as a programmer. But then you have this problem of you want productivity. You don't want to have to write assembly code effectively, or very low-level code, to be able to do basic things. And so you want very powerful abstractions that you can build in the library. And so what Mojo does is it has very powerful metaprogramming capabilities. And by the way, I didn't invent these. It roughly stole them from Zig and made them better. So thank you to the Zig people. They have amazing ideas. But the idea is to say let's pull this compile time metaprogramming system up, make it much more powerful, embed it in the language. And so now you can build compile time transformations of your code. And so you can say, hey, I want a vectorized function. Well, what does it do? It takes a higher order function, which is, you know, all the operations you want to vectorize, and then it stamps it out across your vector lanes and it handles that transformation for you, giving you a very productive way of expressing these very high performance algorithms that give you extreme predictability.
And so this design point is very different. It's very important for accelerators because with accelerators you can't have the memory allocation that magically happens. You need control. But these features are a very different design point than what most languages are coming from. The other advantage, as you say, is it makes the language way simpler. And so if you take a language like C++, C++ has metaprogramming, has templates and constexpr and this whole weird system that is
It's a bit not the simplest to use. I mean, you learn it. So like
You learn it, and I've written way too much C++, and so, you know, even I can handle it. But what C++ does is it has the program and then the meta program and they're effectively different languages, and so this is really complicated, and writing something using templates, it's almost impossible to debug. You can't even pass a string to a template really. There are hacks, but it's just really difficult to do things. You certainly can't make a binary tree or something and pass it into a template. That's not a thing. And so in Mojo and Zig, what they do is they unify the program and the meta program. And so you can just write code and then run that at compile time. And so now if you want to debug it, cool. Just use the same code at runtime and step through it in a debugger. Makes it super simple. It means you have one thing to learn and it works in both places. It means you can have real types and general types. In Mojo it's very fancy because you can use a list or a dictionary kind of data structure that does heap allocations and you can build it all at compile time and then plot the output of that into your program. And so it will synthesize the mallocs for you for a bunch of things you compile-time evaluated, and it all composes the right way. And so what this does is it gives you a tremendous amount of power when you care about performance, when you care about building these higher level algorithms which end up being like mini compilers. People generally don't think about it that way, but they give you the ability to specialize based on, you know, dimensions and numbers and magic constants in your algorithms, and allows you to build these very high power abstractions very naturally.
And if I want to get started with Mojo outside of just the basics, so using some of these more advanced features, is the best way to actually rent, may that be a GPU or something, these days it's pretty easy, and go and try to implement an algorithm, a program that actually needs high performance, and start to play and understand these additional ways that I can control a compiler? Or what would your suggestion be for someone who is like, okay, Leon believes that this is the future and this is cool to experiment with. On their job they don't currently have a problem like this. How would you go about it on the weekends, right?
Yeah. So I can talk about what makes Mojo cool.
Yeah, let's do that.
Well, but see, that ends up in the rabbit hole of really weird technical things only Chris understands. Wait, let's not talk about linear types. Let's not talk, let's, but instead, let me tell you about why you might care. Okay, so Mojo is a
member of the Python family. So it's super familiar, easy to learn. And by the way, with the AI coding, it's the easiest time ever to learn a programming language. Just like you mentioned, one of the simplest, easy things to do with Mojo is to extend an existing Python module.
So if you've ever been coding up Python, building and building and building, it gets big and big and big and suddenly gets slow. Well, historically your option is to go rewrite a big chunk of it in Rust or C++ and then you have to use bindings and it's really kind of a pain in the butt to be able to do this kind of work. Mojo is the most beautiful way to extend Python you can imagine. No bindings, just like Swift and Objective-C. They natively spoke to each other.
Oh yeah. No bindings, because Rust is a popular choice to extend Python, but with a binding.
Yeah. But you need bindings, right? And so it's very mechanical and you can get it wrong. And so Mojo does that without the bindings, but also integrated with pip packages and the build systems, all integrated, and it's just super beautiful and awesome and you can stay on CPU. You don't need anything. It's all free, like go nuts and go do this. We have very simple
Because I guess one thing with Mojo that I don't think you mentioned, but I know that it's also not just for GPUs or TPUs but also CPUs.
CPUs. And today's CPUs are extremely fancy and complicated machines and getting performance out of them is also really interesting.
And so what I've seen people do, and to come back to the empowering stories that I love: we had a person come to one of our community meetings and they rocked up and it's like, oh okay, tell us what you're working on. And they said, well, I don't know anything about this fancy, you know, AI stuff, right? I am merely a geneticist. And I'm like, merely, right? He's like, I work with DNA sequencing and this stuff and I'm a Python programmer, and so I saw that Mojo makes Python go fast. Awesome. And so I said, okay, well, maybe I would try using it for some of my DNA sequencing stuff. And so I pulled over some code and wow, it was like a hundred times faster, because I pulled it over to Mojo and I could learn how to use threads and then I could vectorize, I could... it was very cool because
And then you can start to use the CPU's resources if you're running on a CPU, which, as you said, have a lot of capabilities, and you can play around with it if you learn how to control the
Exactly. And so Mojo makes it really easy to do that where you are on a CPU. And then the thing I loved about it was he said, and then I read that Mojo is really cool for GPUs and I'd never done that before. I don't know anything about GPUs, but I decided to give it a shot and in an afternoon I got my thing running on a GPU and now it's a million times faster and now I'm a GPU programmer, right? And so that to me is super empowering.
And this is where, when you look at technologies like CUDA and things like this, again, they're amazing systems and there's a lot of good energy and good work put into it, but it's 20 years old. It's C++. It wasn't designed for modern architectures. GPUs today have tensor cores and CUDA can't even acknowledge that that exists really, and so this is a huge, huge opportunity for us to get more people into the space and for people to learn about this for the first time.
I mean, you know, just as an engineer who is like, I don't program GPUs, but of course everything I write runs on CPUs or cloud or wherever. This reminds me a little bit of the story you previously shared of the Swift person, because, speaking for myself, I can't really imagine myself just saying all right, I'm gonna write CUDA, I don't really have a use case. But I can absolutely imagine myself just adding a bit of Mojo, optimizing it and then maybe just deploying it on a GPU to see how it is, if I see stuff that is parallelized, play around with it, is it faster, and then just like with the geneticist
Build and learn as you go.
Yeah, and then start to understand this other hardware, which of course will increasingly probably be useful, you know, we know that. And also they're now available. You can rent them on the clouds. Maybe not today, but I think today you can rent them by the hour.
Yeah, if not, it's going to happen for sure.
Yeah, you can totally do that. Well, and so we even go further. It turns out that lots of things have GPUs today. With Mojo you can run on AMD, Nvidia, and Apple GPUs. And so if you do want to learn about GPUs, you can go to our website, search for GPU puzzles, and we'll teach you how to do this. And actually, it's an amazing time. I love it
because what we're doing, and we're still early, we're still building into this and we'd love help. Please come join our open source community and help us build into this. But what we really truly believe is that GPUs and AI accelerators and all these chips, whether they're for AI or for DNA sequencing or chemistry or bioinformatics or oil and gas exploration, like there's all HPC, there's all these applications that can use these accelerators, but all the software around them is really weird, right? And if we can break that down, if we can make it easy to learn, if we can teach people, if we can make them consistent across the hardware vendors, because these hardware people don't always get along, by the way. But if we can make it more uniform, then we can get a bigger community.
In that community, sure, some of the people, some of the old dogs from the existing community will convert over and we love that. And you know, Mojo's way nicer than C++. And so it's way better from a user experience perspective and compile times. And there's so many things that are better about it than older technologies like CUDA. But also what I believe is that there's an entire new generation of people that should be programming these chips. And if we can get more people into this ecosystem, they can upskill. They can get better jobs. Like, these are high-paying jobs to do this kind of work. And so it's a better time than ever to learn how to do this. And so it's, you know, part of my life's work here to enable this and enable people to get into this ecosystem and break down the gatekeeping.
And really it feels, you know, you have just always been passionate about compiler technology and going towards here. Can you tell me where Modular is as a team? You started about 3 years ago. Can you tell me how much progress you made, how large the team is right now and where Mojo and the ecosystem is and, you know, where you're focusing next?
Yeah. So Modular will be four years old in January. It's hard to believe. Turns out things do take four years sometimes and so
Even with the full focus.
Yeah. And so we are an unusual startup and we went into this knowing this, but we are a very large team and very expensive, because we believe and our investors believe that this problem that we're tackling is very important. And so at this point we're just over 140 people, something like that, and we've enabled seven different architectures from three different vendors. So that's Ampere, Hopper, Blackwell from Nvidia, the 300, 325, 355 from AMD, plus all their consumer stuff, and now Apple GPUs in beta and coming online, which is very exciting. We're probably not done yet. We'll keep going, but can't share stuff before it's time to share.
But it's still very early days. And so what we're doing is we're building into this from a technology perspective. Mojo, for example, I want 1.0 to be meaningful. And so I think that we'll have 1.0 early summer next year. So that'll be pretty cool and it will be way better than Swift.
Yeah, but I was about to say, like, this 1.0 will be
So the team's scoping that out. We'll come up with detailed dates soon and we'll share that when it's the right time. And so as we're doing this, what we're doing is we're being really thoughtful and learning from my previous journey and from other people on the team's journeys, to make sure that we do this well. And so we want to provide stability for the ecosystem, which Swift 1 didn't do, by the way, and build into this. And each time we do pass through one of these milestones, you know, we can expect more growth and more use cases and more investment from different kinds of people. That's what I've seen.
So commercially we have customers and we're scaling into that. We have a cloud platform, which is the way we monetize most of this stuff. Mojo is not a product that we're trying to make tons of money on, by the way. It's a thing we had to build to be able to scale across lots of hardware. And so we see a dual world where we can actually catalyze and grow and teach people how to build and use all these accelerators, and then we can help enterprises build and manage AI into their ecosystem. And a lot of what these folks are looking for is they want something that's easy to use, that actually works, that's reliable, that they can scale. A lot of them want optionality on hardware. They want to be able to buy the best chip for their workload. They want to be able to multi-source from different vendors. They don't want to have to rewrite it three times.
Yeah. I mean, this all only makes sense especially when there's this dependency on one or two vendors, usually one.
Yep. And what I see and what I predict in 2026 is there's going to be a lot of silicon from a lot of different vendors and we want to make sure that people can scale across that effectively. And so, you know, for people that are interested, we have lots of job postings on our website, and we're growing very rapidly and we have a very big mission. It's not another six months and then we're done. We're just still at the beginning, and every year of our journey is just another epoch of new things to build and learn and grow and develop into.
And of course you're building, sounds like, the infrastructure under AI, or hopefully a lot of it could be. How is your team, the engineering team specifically, using AI tools? May that be, you know, IDEs, agents. You mentioned that Mojo is a premier language to write with AI, but I'm interested underneath the hood, are you seeing the impact of these tools? Is it helping engineers work faster or down in the current? Because there's always this question: if you're writing compilers, will these things help at all, or no, you need to handcraft it. And I'd love to ask someone who's actually seen it.
Yeah. Well, so I can't tell you all the things, but I can share my experiences. So we definitely encourage our team to use AI coding tools, and so use Claude Code and Cursor and all 57 other things, and so there's lots of different experiences. I've seen tons of benefit, and so for me, I still code, and so I use Cursor, for example, as my daily driver right now, and so I feel like it's, you know, a good 10% productivity, particularly for mechanical rewrites and stuff like this.
But you're a really good programmer, so that's pretty meaningful.
Yes, I represent a very sophisticated programmer, correct. Yes. But I really enjoy it because it frees me from a lot of the mechanical stuff, and so whether it makes me in aggregate more productive or not, it increases my enjoyment, which is good.
Now, important.
It is, actually. Now for other folks that are building prototypes, or PMs that are building wireframe models of things, it's transformative, because you can literally 10x somebody and you can make it so they could build something they otherwise wouldn't get around to doing. And so that is transformative. For production coding, I think that it's kind of hit and miss. And to me, honestly, I don't know if it's actually a net win or not, because I've seen many cases where people say, "Okay, just go let the agent try a thing." And then it grinds and grinds and grinds and a gajillion tokens later, it doesn't work. If they would have just done it, then it would have been done sooner, right? And so to me, again, there's lots of wins, but there's also some losses. And so I don't know how things will net out.
I do know it's moving very rapidly and I do want us to be using it. But also one of the things I'd say is I don't want people to turn their brains off. And one of the things I think is really important is that programmers, in our team but also in general, use it as a human assist, not a human replacement. And I think it's very important that we review the code. We understand the architecture. And for production workloads and production applications at least, vibe coding terrifies me. Not just because of what does it mean for jobs, but what does it mean six months from now when you want to change the architecture of something?
Yeah.
How are you gonna even understand how anything works, right? And so I do think that it's really important to keep humans in the loop and architects thinking about something and understanding things. Not just for security and performance and all the other reasons as well.
Well, Addy Osmani, I talked with Addy Osmani, who's been on the Chrome DevTools team for like 13 years now. I think a lot of the tooling he or his team built there. And he wrote this book called Beyond Vibe Coding, and he said that what he does is he enables a verbose kind of output for the agent and he actually reads through what the thing was. He makes sure that before he commits something or reviews that he just understands
what's there, so that he always stays in touch, so, you know, if this thing turned off, he could jump in. It's just like he now just chose not to do so.
Well, and I think that, again, the thing that I care about is the architecture of production. I mean, prototypes are a different deal, but for production you keep the architecture clean and right, and it doesn't need to be perfect but it needs to be curated, because I've seen AI coding tools go crazy with duplicating things in different places. Well, we know that's a problem, because then you want to go change something, and now you have bugs when you update two out of the three places, right? And so there's a lot of these things where the tools are amazing but they still need adult supervision.
So when it comes to hiring, you're building something really interesting and cool and honestly super ambitious. You already have a pretty big team, but you're growing. What is your hiring bar? And I think you previously mentioned that you're building a company of elite nerds. I'm interested, like,
what are things that you look for in folks, and, you know, for software engineers, experienced engineers especially, listening, what tactics or strategy would you suggest so that they get to the level that they could have a shot at a place like Modular, even if not necessarily Modular?
Yeah. So we hired two different kinds of folks. One is the super specialized, you know, compiler nerd or something like this, or the GPU programmer expert that knows how to do high-performance matrix multiplications and they've got 10 years of experience, and so you can go look at that track record. We also hire people fresh out of school.
Wow, this is refreshing.
And so I really do appreciate people that are very early career, because they haven't learned all the bad things yet. And so for me, I think maybe the more interesting thing we talk about is not the super expert that's already the senior engineer. It's about folks that are much earlier in their career. And so what I love to see is people that are really hungry. They have not given up and just said AI will do everything for me, but they have
intellectual curiosity. They're willing to work hard, of course, right? But then they're fearless, because everything, particularly in the AI space, is changing so rapidly. A lot of people will just freeze up and they won't do anything, versus if you say, "Okay, cool. Let's go figure it out. I want to learn how to do this." I mean, this has been my life's journey, is just, that sounds terrifying. People tell me it's impossible and it's doomed to failure, but how hard can it be? Let's just go figure it out and work through things and do that.
I personally really love when people have contributed to open source projects. I think that's the simplest way to prove that not only can you write code, but you could actually work with a team, which is a huge part of software engineering in reality. And so I love internships and things like this that people have done before. And when talking with people, you can tell whether they're excited about what they do or whether they're just, you know, kind of performatively going through the motions.
Now, one of the really challenging parts is that in interviews people get very nervous, and so one of the things I really believe in is giving people the native tools they're used to using. So let people code the way that they would be normally coding. And so today, for example, we want people to be using AI coding for the mechanical pieces, right? So saying, okay, well, you have to write on a whiteboard and you can't use your native tools, would be very strange. And so we try to make sure to meet people in something that actually approaches a real-world situation.
One question I've been meaning to ask, because you built several languages. This comes up, and I asked folks what are things they'd love to hear from you, and this was repeated quite a bit. Do you think there is any logic in having a new language that is designed to make it easy for LLMs to code with, and like attributes, for example, better pattern matching or other things that might be the strength of LLMs? And if so, what would these patterns be? This is really weird. I would have not thought we would talk about things like this, but now we should.
Absolutely. I'm happy to talk about it. And so I sometimes get asked, so given the AI is writing all the code, why are you building a programming language? Which is another way of asking the same question, maybe a little bit more aggro. And so to me, I don't think that optimizing for the LLM is the right thing to do at all.
Because go back to what we were just talking about. The important thing is reading the code. Always has been, by the way, right? Code is read more often than it's written. AI has made it way easier to crank out the code than it ever has been, and I don't think we're going back to where we were. So writing the code is actually not the key thing to me. It's about reading it.
And so what I care about with Mojo and with these new emerging systems we're building is really a combination of two things. One is expressivity. Can you express the full power of the hardware? Yes or no? If the answer is no, then you will never be able to get it. And so JavaScript will never be a good way to write a CUDA kernel or GPU kernel. It just will not, because it can't express the things you need to express, right? No slight against it. It's a good system, right? But if that's the goal, then you have to be able to do that. Now, assembly code can express the full power of the hardware, and that's not what I'm advocating for, right?
The other thing that goes with it is you need readability, right? And so you need this combination, this intersection between expressivity, so can you express the important thing, in our case performance, and then can you understand the code? Can you build abstractions? Can you build scalable systems? And that intersection, I think, is the key thing.
And so this is where you take Python syntax, for example. Python's super easy to read. It's actually very standard. A lot of people know it. Let's embrace that. Let's go all in on that. That's a good thing. And so Mojo embraces Python, the syntax of Python at least, and says, "Okay, well, now let's fix the problems with Python," which is basically the entire implementation. Let's replace the Python interpreter and all that kind of stuff. Let's keep this good part and make it so we can express that full power of the hardware. And with that, I think the LLMs will continue to get better and better at dealing with any weirdness. And what I've seen is that the LLMs are an amazing way to learn coding in a new language.
Yeah, this is interesting, because I feel we're starting to see, including in software, how to build good software, oftentimes what is good and efficient for teams is also what LLMs tend to do better with.
Yeah, I also get these funny questions. It's like, okay, well, a lot of people in agentic coding loops want to have error messages that feed back into the agents so that the agents know what to do, and so should we make better error messages for the agents?
Make it better for humans.
And my question is, what is better for an agent that's not also better for a human? Let's just have good error messages, right? The other key thing, and I think the most important thing for LLM-based coding and AI tools, is have a massive amount of open source. And so for us, if you go to the Modular repo, we have, I don't know, 700,000 lines of Mojo code that's open source.
And when you open source it, you open source the whole history as well, by the way.
Exactly. Exactly. And so that's super powerful, because now you can say, hey, go index this. And so now it knows a gigantic swath of everything you can possibly do. And so that, I think, is also really important.
Yeah. Now, to close with, you have mentioned that compilers are cool multiple times here.
I'm slightly biased, I admit.
Absolutely. But I'm getting convinced that compilers are cool. For software engineers who have not built or worked with compilers, what are some pointers you would give them of, hey, I'd like to get into understanding compilers, maybe building one? Should I try to come up with a language? What are some things that you point people to?
Yeah. So I'll tell you why I fell in love with compilers. If I go back to university, I was taking the standard computer science program of basic programming and then data structures and then an operating systems class and a GUI class and all this kind of stuff.
The thing I love about compilers is that you built project one. And so in this case, it was what's called the lexer, the thing that tokenizes the source code, and you had to use some data structures. So I'm like, oh, okay, cool. I'm using the things I learned, not just how to write code, but what a tree is and things like this. And then you get to the second part and you build on top of it. And so now you build a parser, the thing that decides the syntax of the language, using the tokenizer. And then you build the type checker, and you build on the first two. And then you build the next thing, and you build on those things.
And so what I loved about compilers as a class was that I got to build and learn and iterate, and then if I made a mistake, I had to go back and fix it, because you're building higher and higher and higher. And so I think that compilers, particularly in the university setting, reflect more of real software development in a way that a lot of other classes I experienced did not, because a lot of the other classes ended up being: build a thing, turn it in, throw it away; build another thing, turn it in, throw it away. And so that's why I fell in love with it.
Today, if you want to learn compilers, it's easier than ever. You can go to LLVM. There's a great tutorial called the Kaleidoscope tutorial that I wrote 15 years ago or something. It's been a long time. The Rust community has a lot of lovely compiler nerds in it, and so a lot of compiler-based technologies are built in Rust these days, which is also super cool. And so there's books and lessons and things like this.
I don't literally think everybody needs to become a compiler engineer, but I think that compilers don't get the credit they deserve, and it turns out there's a lot of really good jobs in compilers. And so if it happens to be that it's an interesting place and an interesting set of technologies, then I really do encourage people to do it, because we need more folks in this field.
Wonderful. Chris, this has been really, really interesting. Learned a lot. Thank you.
Yeah. Well, thank you for having me. I hope it was useful.
It was. It's rare to talk with a software engineer who is so passionate about programming as Chris is. And I hope you enjoyed this conversation as much as I did. I especially liked how Chris took us to the backstage of how things really happened at Apple, how it was Chris's own stubbornness and track record of getting things done that allowed him to push things through, like open sourcing the Swift programming language and creating the new language to start with.
For more deep dives on AI engineering and stories on developer tools at other companies, check out deep dives in The Pragmatic Engineer that I wrote, linked in the show notes below. If you enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks, and see you in the next
Article published
