Martin Kleppmann on Designing Data-Intensive Applications, the Cloud, and Research After Industry
The Pragmatic EngineerMartin Kleppmann wrote Designing Data-Intensive Applications (DDIA), a standard reference for engineers who build large back-end systems. Nine years after the first edition, a heavily updated second edition is out. In this conversation on The Pragmatic Engineer, Kleppmann covers the path from two startups to LinkedIn's data infrastructure team, then to writing and to academia. They explain what changed in the second edition, why cloud object stores changed how data systems are built, and why they think formal verification may matter more once AI writes much of our code. They also describe their current research on local-first software and on using cryptography to verify claims about physical supply chains. Their view throughout is that engineering is about making trade-offs visible so that people can choose deliberately, not about finding a single right answer.
A first startup that worked technically but found no business
Kleppmann studied computer science as an undergraduate and then was unsure what to do next. Starting a startup "seemed like an interesting thing to try," so they started one without knowing what it would build, and spent the first stretch searching for a promising idea.
That company was Go Test It, around 2008. Cross-browser compatibility was a serious pain at the time. Internet Explorer was still widely used, Chrome had just come out, and browsers were incompatible with each other. Go Test It was a hosted, automated cross-browser testing service built on Selenium, the open-source project that still exists. Customers wrote scripts that simulated a user clicking through a site and checked that the right behavior happened, and the service ran them so customers didn't have to maintain VMs with various operating systems themselves.
Kleppmann says the product worked technically, but adoption was the problem. Website builders agreed in principle that cross-browser testing mattered. In practice it was hard to get them to fit it into their workflow, get in the habit of using it, and invest in writing test scripts. Kleppmann notes that one or two companies from the same era did make a business of it, with Sauce Labs succeeding. Even for them, Kleppmann thinks, it was a slow-growing and difficult business.
The company was based in the UK and mostly bootstrapped. Kleppmann did consulting work to pay for hiring friends cheaply to help build the product, with only a very small amount of angel money. The venture didn't go far, but through it Kleppmann met the people who became co-founders of the second startup.
Rapportive, Y Combinator, and the move to San Francisco
The second company, Rapportive, went much better. It put social media inside Gmail. A browser extension modified the Gmail web interface so that when an email arrived from an unfamiliar sender, a sidebar showed a summary profile: picture, job title from LinkedIn, recent tweets, perhaps recent Facebook posts, and whatever else could be found. It started around 2010, became popular quickly, and on that basis the team raised money from Y Combinator, which Kleppmann describes as already well regarded but still fairly small.
The founders flew from the UK for the roughly three-month program, then obtained US work visas and settled in San Francisco. Kleppmann remembers this as exciting, like "going to the center of where it was all happening." They arrived knowing only one or two people in the Bay Area, who introduced them to others, who introduced them to still more. Kleppmann valued how open the area was to outsiders who could show up with an early-stage idea, raise money, and become established.
The technology was, in Kleppmann's words, "fairly unexciting": a Rails app on Postgres with some Redis. The one technically interesting part was that they had essentially built a graph database on top of Postgres.
An acquisition under pressure, and LinkedIn Intro
Rapportive was sold to LinkedIn in 2012, when the team was five people. Kleppmann calls it a success for everyone, though not a large amount of money, and is candid that it was not a triumphant exit. They had tried revenue-generating options without success, user growth was acceptable but not enough for a large round, and they were running out of money. Their visas also constrained them: they couldn't cut their own salaries, because doing so would have violated the visa conditions. Selling became, as Kleppmann puts it, "the least bad option." Given how little leverage they had, Kleppmann is pleased with the result, and says the process had the usual twists and moments when it seemed the deal would collapse.
Kleppmann speaks warmly of LinkedIn. The team was allowed to operate as an essentially independent unit and to keep building what it wanted. The Rapportive extension was put "on life support" while the team worked on a new product, eventually released as LinkedIn Intro. It got what Kleppmann calls "a slightly weird reception" and was shut down soon after launch. Kleppmann says there is a longer story there but doesn't tell it. After the shutdown the team was disbanded. Kleppmann still credits LinkedIn for giving them the freedom to build and launch the product even though it failed.
Kafka, Samza, and seeing how data systems fit together
After the team broke up, Kleppmann joined LinkedIn's stream processing team. Kafka had recently been developed there and had just been open-sourced, and Kleppmann worked on Samza, a stream processing framework built on Kafka.
Asked why LinkedIn built something that now looks so generic, Kleppmann points to Jay Kreps's blog post from that era, "The Log," which explains why Kafka was an append-only log and not a traditional message queue. In Kleppmann's account the motivation was data integration. Many databases and event-producing systems, such as user activity events, generated stream-shaped data. Many downstream systems needed that data, including the data warehouse and the Hadoop cluster used for machine learning. The practical problem was how to physically move data from one system to another. Kreps designed Kafka as that integration point: something close to a lowest common denominator, but still a general-purpose abstraction connecting many sources to many sinks.
This was Kleppmann's first hands-on experience with a really large system. The biggest company they had worked at before was Rapportive, with its single-instance database. At LinkedIn they got to hand-code MapReduce jobs in Java against a large Hadoop cluster, which they enjoyed. The real revelation came from the stream processing ideas and from Kreps's evangelism for Kafka. Kleppmann says they began to understand how the various data systems fit together, what they have in common, and what the underlying principles are. That understanding went directly into the book.
Leaving LinkedIn to write full-time
Kleppmann first moved back to the UK and kept working for LinkedIn remotely. Their girlfriend at the time, now their wife, was still in the UK, and Kleppmann didn't feel at home in the Bay Area, so they didn't push for her to move there. They love the Bay Area as a place to visit and still have many friends there, but wouldn't want to live there.
LinkedIn gave Kleppmann 50% of their time to work on the book alongside engineering duties. Kleppmann stresses that LinkedIn didn't have to do this and got little from it beyond a book it could use for internal training. Even so, the arrangement didn't work. Writing in parallel with an engineering job that included on-call duty meant too much context switching. Urgent on-call issues easily took over, leaving no room for the freedom needed to write something new. Kleppmann eventually left LinkedIn for an unpaid sabbatical ("i.e. unemployment") to write full-time. Only after that did they consider academia.
What the book was meant to be, and how it was researched
Kleppmann says the final book differs from the initial plan, but the goal stayed the same. It would be a broad conceptual overview that compares trade-offs across many kinds of tools, not a guide to any single system. It would be practitioner-focused, not a theoretical textbook: something people could use to build real systems. It was the book Kleppmann wished they'd had at Rapportive, where the team hit database performance problems and was "searching around in the dark" because it lacked the foundations to diagnose what was happening. With more background on how data systems work internally, Kleppmann believes, they would have had the intuition to debug those issues. Once they had learned it, they wanted to write it down so others wouldn't have to learn it the hard way.
Much of the learning came from curiosity and conversation. LinkedIn had senior data systems engineers who understood this material well but hadn't necessarily written it down. Kleppmann questioned them and built a mental model from those conversations. With those basics, research papers became approachable. Papers go much deeper into how and why systems are designed as they are, but they take a long time to read, so Kleppmann tried to extract the essential ideas. They also read many blog posts. The long reference lists at the end of each chapter are the material Kleppmann actually used, included as further reading for anyone who wants to go beyond the basics.
Kleppmann says the book's three-part structure (foundations of data systems, distributed data, derived data) was imposed mostly after the fact. The chapter topics, such as transactions, replication, sharding or partitioning, and consistency and consensus, were clear from the original proposal to the publisher. Kleppmann wrote one chapter at a time and started each with extensive background research. The internal structure of each chapter emerged then. For replication, for example, they concluded only after that research that the three main approaches were single-leader, multi-leader, and leaderless, and organized the chapter around them.
Estimates, deadlines, and the two editions
Asked whether estimating a book resembles estimating a software project, Kleppmann agrees that both take "vastly longer than expected." The first edition took about four years of elapsed time, or perhaps two and a half years of full-time-equivalent work, and missed the publisher's deadline by roughly two and a half years. O'Reilly was relaxed about it and let Kleppmann take the time to make it good. For the second edition O'Reilly was "a bit more aggressive and pushy" about deadlines. Kleppmann understands why, since by then the book was established and readers were waiting, but they missed the freedom they had the first time.
Reliable, scalable, maintainable, and scaling down
The subtitle of both editions refers to reliable, scalable, and maintainable systems. Kleppmann acknowledges these terms have no formal definitions. For them, reliability mainly means fault tolerance: the system should keep working overall when a network link fails or a node crashes. Much of the book covers techniques that support this, such as replication.
Kleppmann says scalability is "thrown around a lot" because it is fashionable and suggests success and millions of users. The book takes a more detached view: scalability is about the mechanisms for handling changes in load, such as adding computing capacity when load increases, using techniques like sharding. The book focuses on horizontal scaling because buying a bigger machine is less interesting to write about. What has become interesting about modern cloud and back-end services, in Kleppmann's view, is shared-nothing architecture: handling very high load using relatively cheap commodity machines.
Kleppmann says they have recently thought more about something they didn't consider much at first: scaling down. Ideally cost and computing capacity should be roughly proportional to load, which means a very lightly loaded service should be extremely cheap to run. That isn't guaranteed. With on-premises infrastructure a physical machine is the unit of deployment, and even if you split it into two dozen VMs, you still have to allocate resources to each. Kleppmann finds some serverless systems interesting because a service handling three requests a day is perfectly acceptable to them.
The second edition: bringing in Chris Riccomini and building on object stores
Kleppmann had known for several years that a second edition was needed because the first was becoming dated. By then they had an academic job centered on research and teaching, which they enjoy and didn't want to give up, so updating the book was a side project. The same context-switching problem returned, and progress was slow. Kleppmann also realized that, after moving into theory, they had lost touch with current industry practice, such as how people were using data lakes.
The solution was Chris Riccomini, a former LinkedIn colleague from the stream processing work and author of The Missing README. Kleppmann had read that book and considered Riccomini a strong writer. Riccomini also wrote a newsletter, Materialized View, on current trends in data systems, and had become a startup investor in that area. Kleppmann describes a good division of labor: Riccomini knew the current state of industry, and Kleppmann had strong views on teaching, meaning precise, carefully chosen wording that still reads easily.
The main change they planned from the start was cloud-native architecture, meaning data systems built on cloud services as the foundational abstraction. The first edition assumed machines with local disks: a database instance writes to its local disk, and the database software replicates data to another machine, which writes to its own disk. That was how computing worked for a long time. Now databases are being built on object stores, and replication happens at the object-store level instead of, or in addition to, the database level. Kleppmann distinguishes this from building on virtual block devices such as EBS. Those are cloud services too, but they still present the single-node abstraction of a block device with a file system on top. An object store is a new abstraction that looks and behaves differently from a file system. Some systems were starting to use it at the time of the first edition, and since then it has taken off. Kleppmann says the second edition weaves the idea through the whole narrative instead of confining it to one section.
Do engineers still need to understand the layer beneath?
The host asks whether managed services, which handle replication and come with uptime SLAs, remove the incentive for engineers to understand what lies underneath. Kleppmann sees this as the familiar history of computing: new, higher-level abstractions. Relying on a higher-level abstraction does mean not thinking about lower-level details. Using a garbage-collected language means not thinking about memory allocation. Kleppmann thinks that's a loss only in some contexts. People building low-level systems still need to care about memory. People writing higher-level business logic, Kleppmann thinks, are fine not caring. Data systems are similar: if you build higher-level systems that don't need to care about infrastructure, use the abstractions. Someone still has to build the cloud services from lower-level components, and those people will specialize more deeply in how to engineer, operate, and make them reliable, treating the higher-level builders as their customers.
Kleppmann adds that the book's philosophy is to give people a sense of how systems work internally, so that when something behaves strangely, such as unexpected performance, they have some intuition about why. The storage engine chapter explains B-trees and log-structured (LSM) storage engines. It isn't for people building their own databases, who need far more depth. It is for application developers who can use a storage engine better and diagnose problems if they know a little about how it works. The second edition applies the same philosophy to cloud services. Kleppmann's example is row-oriented versus column-oriented storage for analytics: a technical distinction that takes some background reading to understand but has large performance consequences. In cases like that, knowing the internals is "actually like a superpower."
Availability, cost, and a European case for multi-cloud
The discussion turns to how an engineer decides between multi-zone, multi-region, or multi-cloud setups. Kleppmann frames it as how much availability risk you'll accept versus the overheads: the computational overhead of the system, the human overhead of designing and operating it, and the cost. A more fault-tolerant system costs more to design and run. A simpler one may go down more often but is cheaper. Kleppmann says there's no right or wrong answer, and everyone has to decide where they sit on that spectrum. Multi-region pushes toward higher availability because it tolerates losing an entire region, but it affects which consistency models are possible across regions. The book tries to make these trade-offs explicit.
On multi-cloud, Kleppmann says a new concern has come up "just in the last month really": Europe's dependence on US cloud services. If geopolitics went badly and Europe were suddenly locked out of US cloud providers, the effect would be severe. Kleppmann hopes this won't happen and still considers it fairly unlikely, but "no longer unthinkable." From a European perspective they have been thinking about how to make systems resilient against that, which is a business risk as well as a kind of outage. A multi-cloud setup would let systems keep running with another provider if one company locked you out. It sits at the expensive, risk-reducing end of the spectrum, but Kleppmann thinks it is worth serious consideration for critical workloads where the geopolitical risk is significant.
Both agree that understanding and communicating risks and trade-offs will be a core part of engineering work. Kleppmann suggests that as AI writes more code, the job may depend less on expressing logic in a particular language and more on these high-level trade-offs.
How the cloud changed scaling, and why sharding matters less
Asked whether cloud primitives make scalability easier to reason about, Kleppmann separates two ends. At the high end, reaching very large scale is still hard. Object storage provides elastic capacity and removes disk capacity planning, but sharding affects application code and can't be made fully transparent. If one machine can't handle your workload, you still need substantial engineering thought, cloud or not. Where the cloud has clearly helped, Kleppmann says, is the low end: serverless systems that spin instances up and down quickly enable very lightweight services. Without cloud services this would be much harder, because memory and CPU would have to be allocated statically to a VM. The host mentions a small serverless site that costs about 13 cents a month, and Kleppmann calls this more efficient use of computing resources.
The host recalls that sharding was a major topic at Uber, including in interviews, and that it now seems less common for engineers to implement it themselves. Kleppmann attributes this less to the cloud than to more powerful hardware. A big machine can do a lot, so more workloads fit on one machine and still reach significant scale. Parallelism still matters, since you have to use hundreds of cores efficiently, and sharding is one way to get it. But sharding across machines is less pressing for many workloads. It hasn't disappeared, because some workloads still need it. Replication stays relevant even at small scale because it serves fault tolerance, not scalability.
"The Trouble with Distributed Systems"
The chapter with that title, Kleppmann explains, defends the assumptions of distributed systems theory by showing that they reflect reality. Theory usually assumes no upper bound on message delay. A message might arrive in 100 microseconds or in 10 years. Some theory does assume timing bounds, and Kleppmann calls that dangerous because network delays sometimes become much larger than usual. Theory also says nodes can crash, but in practice "crash" covers a software crash, a hardware failure, someone pulling the power cable, or a node that is still running but disconnected from the network. Clocks are another example: they're usually roughly correct but not precise enough to rely on. The lesson Kleppmann draws is that it's tempting to assume things behave well, and reliable systems require giving up those assumptions: "don't believe anyone who says oh failures are rare."
Kleppmann enjoyed writing the chapter. It is largely a collection of things that went wrong, built from postmortems published by tech companies and their root causes. A favorite example is sharks biting undersea cables. Kleppmann has heard that cable shielding has improved so sharks no longer cause damage, and that now cows on land step on cables and occasionally cause network outages.
The host notes that for a team like S3, the chapter describes daily life, since at that scale disk failures and even data center fires happen regularly. At a small company the same event may happen once a decade and be a big deal. Kleppmann agrees there is no universal answer. It's a trade-off between risk and cost, which makes it a business decision. The chapter aims to give people what they need to decide in an informed way, and Kleppmann doesn't want to make that decision for them.
What left the book and what came in
Some material could be cut. The first edition covered MapReduce in detail, but Kleppmann says bluntly: "MapReduce is dead. Nobody uses it anymore." Its successors, such as Spark and Flink, are in use. The second edition still mentions MapReduce, but as a teaching tool for understanding how partitioned batch processing systems work.
Coverage grew in areas that support AI. The book isn't about AI, but AI applications raise data systems concerns. Vector indexes are the clearest case. They were added to the storage engine chapter, which fit well because the chapter already compared indexing strategies and vector indexes are another one. Data frames were also added. Kleppmann notes they aren't exclusively an AI concept, but they are a good representation for training data and have become an important data model alongside relational, graph, and JSON documents. The first edition didn't cover them. Kleppmann describes these as expansions that reflect what people build without changing the book's direction.
Ethics becomes its own chapter
In the first edition, "Doing the Right Thing" was a subsection near the end. In the second edition it is the final chapter. The host quotes it: "We, the engineers building these systems, have a responsibility to carefully consider those consequences and consciously decide what kind of world we want to live in."
Kleppmann says the first version came from a sense that ethics had been largely ignored during their time in industry, especially in startups focused on building products customers would love. Consumer products were often designed around harvesting behavioral data because that data could be monetized through advertising, with little reflection on what was good or bad about it. Kleppmann didn't want to prescribe an approach, but wanted to point out that data protection laws now affect how data systems are designed, and that there is an ethical responsibility. If people go into tech to change the world, thinking about the effects of their technology is part of the job, and engineers are prone to neglecting it. Kleppmann says the section was also a way of working through these questions personally, since they hadn't thought much about ethics when they started building these systems.
Asked whether engineers are well placed to influence outcomes, perhaps more directly than regulators acting years later, Kleppmann agrees they have a strong voice. Engineers should present trade-offs so business leaders can decide well, and presenting trade-offs includes naming risks. Those risks go beyond technical ones like data corruption to societal harms, unintended consequences, and reputational damage to the company if a technology turns out to be harmful. Kleppmann wants these decisions made deliberately, "not just sweep it under the carpet."
Formal verification in an AI-assisted world
The host asks about a post Kleppmann wrote in December arguing that formal verification may become more important with AI. Kleppmann first describes the range of formal methods. At the introductory level, you write a high-level specification of a system's expected behavior in a language such as FizzBee or TLA+, and use a model checker, which Kleppmann describes as essentially a randomized test-case generator, to run through many scenarios and check that the desired properties hold. At the advanced level is formal proof: you write a mathematical specification and prove that an algorithm or implementation always satisfies it. Tests check a few example inputs. A proof can cover potentially infinite state spaces and show that a safety property holds in every possible case.
Kleppmann never used formal verification in industry because it took too much time. They took it up in academia, where they could spend months proving an algorithm correct. They find it valuable for subtle algorithms whose correctness is hard to judge from the code, especially high-stakes ones where a bug would corrupt data or open a security hole. They have written proofs with the Isabelle proof assistant and mention Rocq and Lean as alternatives. These proofs are very hard to write: learning the language takes a long time, and even afterwards writing the individual proof steps is laborious.
To show what proving involves, Kleppmann uses a simple example: proving that concatenating two lists yields a list whose length is the sum of the two lengths. You would use induction over one list. Concatenating a list of length i with an empty list gives length i. Appending a list of length one gives i + 1. Continuing by induction shows the result is i + j for every possible i and j. A unit test would check a few cases, such as lengths zero, one, and five. For something this simple you can convince yourself by reading the code, but for complex algorithms, Kleppmann says, our brains can't grasp them well enough to be sure without a proof.
For engineers who want to start, Kleppmann recommends model checking with TLA+ or FizzBee, which are much friendlier than Isabelle, Rocq, or Lean. The proof assistants require much more background knowledge, and Kleppmann admits the learning resources for formal proof aren't very good and they haven't found great books. They learned by pairing with lab colleagues who had years of experience, describing what they wanted to prove and being shown how to break it down step by step.
Kleppmann gives several reasons for thinking formal verification could become more important. LLMs are getting increasingly good at writing proofs, and if humans don't have to write them by hand, proofs become economical where they weren't before. LLMs also increase the need for proofs. With so much code being vibe-coded, manual human review would become the bottleneck, which would cancel much of AI's benefit, so automated ways to check correctness are needed. Extensive testing is a good start, but only proof covers every possible case. That matters most in security, where "it just takes one little bug" to compromise a whole system. For domains that need a complete absence of bugs, Kleppmann hopes LLMs will make formal verification accessible to people who previously found it too hard and too expensive.
Academia's long time horizons and local-first software
Kleppmann says academia covers a wide range, from purely theoretical work unconcerned with the real world to applied research aimed at real impact, and they place themselves at the applied end. The common difference from industry is time horizon. A startup has to ship within months and can't plan ten years ahead. A larger company working on infrastructure can think somewhat longer term because the requirements are better understood. Academia allows work that is long-term, not immediately commercial, or even contrary to commercial incentives.
Local-first software, which Kleppmann has worked on for years, is their example. The goal is to shift power from cloud operators back to end users, so users control their data and depend less on cloud services for their applications. Kleppmann argues this doesn't come naturally to companies. SaaS businesses can charge subscriptions because they can effectively "hold a gun to the customer's head" and say "pay us your subscription, otherwise we will delete all your data." Kleppmann says they understand the commercial reasons for this but consider it an unhealthy situation, and one that is hard to change from inside a business whose revenue depends on lock-in. In academia they can prioritize what they believe is right for users and treat the business model as secondary, because they don't depend on it.
In their vision, cloud services may still help sync data between, say, a phone and a laptop, because routing through a server is often the most convenient way to connect devices. The difference is that no single provider must be trusted. Data could sync through several providers at once, whichever responds first or all of them, and if one disappears, the others remain. That flexibility creates new research and engineering problems.
A hard problem: revoking access without a central server
Kleppmann's current example is access control. Granting and revoking a collaborator's access to a document is trivial with a central server, which checks roles. Across multiple providers or peer-to-peer, it gets hard. Suppose a user's edit permission is revoked while that user concurrently edits the document. Some devices see the edit first and the revocation second, and accept the edit. Others see the revocation first and reject the edit as unauthorized. The devices are now permanently inconsistent. A central server would simply decide which came first. Multiple servers might decide differently. A consensus protocol could resolve this, but Kleppmann calls consensus messy because it needs quorum votes and nodes to be online. The team is trying to solve it without consensus while keeping high availability, offline work, and server-free peer-to-peer sync. Kleppmann says they are close to solving it for Automerge, the CRDT library they work on, but it is much harder than the centralized version.
The host suggests synchronized clocks and timestamps would help. Kleppmann says clocks are useless here: a revoked user who wants to vandalize a document can backdate edits with earlier timestamps. Because actions come from end-user devices, the system has to handle potentially malicious actions.
The host observes that this is arguably a harder problem than most startups take on, since they would accept the constraint of a central server that makes business sense. Kleppmann agrees and calls that the right choice for startups. Research, with different incentives, can take the "idealistic, principled stance" that decentralization is worth the harder problem. If it's solved, others gain options: they can adopt decentralized technology without inventing it, while still weighing its trade-offs.
Teaching, and computer science education after AI
Kleppmann currently teaches an undergraduate course on concurrent and distributed systems, a master's course on cryptographic protocol engineering, a security seminar, and the undergraduate operating systems course, which they describe as a heavy teaching year. The distributed systems lectures are freely available on YouTube. They are more theoretical than the book, focusing on algorithms and on reasoning about their correctness when nodes crash, communication is unreliable, and clocks are wrong. The course is only eight lectures but goes deeper on algorithms than the book. One lecture covers the full Raft consensus algorithm, which Kleppmann calls complex but a good illustration of distributed systems' challenges and of how edge cases and failures can be handled. The message they want to convey is that consensus is subtle and easy to get wrong, but can be solved in a way that works well. Kleppmann describes their teaching style as unremarkable: slides annotated by hand on an iPad during lectures. They note that Cambridge favors theoretical, pen-and-paper courses over practical implementation, and they may add more practical exercises later. The cryptography course is already hands-on, with students implementing elliptic curves from scratch.
Before the AI boom, Kleppmann says, computer science teaching changed slowly, partly because Cambridge, being over 800 years old, thinks on long timescales and emphasizes fundamentals, many from the 1930s such as lambda calculus, instead of chasing fashions. AI has changed assessment entirely. Banning it can't be enforced and would be counterproductive, since students should learn to use new technology productively. The challenge is helping students use it in ways that support their learning instead of undermining it. Some students are mature enough to judge this themselves and many aren't, so guardrails are needed. Assessment also has to be fair and seen as fair: if students believe classmates earn high marks without effort, trust in the system erodes. Kleppmann says frankly that they don't have good answers yet. One step is a boot camp at the start of first year covering basic software engineering skills: version control, unit testing, and generative AI. How to handle assessment is still being worked out.
Kleppmann emphasizes that education's goals differ from industry's. In industry the desired outcome is usually a working product, so if AI gets you there faster with an equivalent result, use it. In education the essay itself isn't the point. "We don't ask the students to write essays because we love reading their amazing essays." The goal is the thinking process and the learning it produces. The host mentions a recent Anthropic study of junior engineers in which the group using AI learned little and the group without AI learned more. Kleppmann says one could "quibble" with the detailed methods, but the general principle seems right: learning sometimes requires struggle, though not too much. Using AI to get past a technicality and focus on the main learning goal is good. Where the goal is to wrestle with difficult ideas, students still need to do that themselves.
Bridging industry and academia
Kleppmann thinks the two worlds often view each other with disrespect. Industry people dismiss research as "theoretical" and irrelevant and miss useful insights. Academics dismiss industry work as "just engineering" without interesting thinking. Kleppmann sees one of their own goals as building respect in both directions: bringing research insights into practice, and letting real-world problems inform research.
Current research: cryptography for verifying claims about the physical world
Kleppmann has two main research areas. One is local-first software, pursued for about ten years through open source work, algorithm development, and formal verification. The aim is collaborative software like Google Docs or Figma that better protects users' data, depends less on a single provider that can lock you out, and gives users more agency and autonomy.
The second is a new area they are trying to establish: using cryptography to prove things about the physical world, with a focus on sustainability. One example is product carbon emissions. If buyers want to choose products with lower emissions, the numbers have to be accurate, and Kleppmann says they currently generally aren't, because incentives encourage lying, cheating, and creative accounting that amounts to greenwashing. Another example is new EU regulation on deforestation. Importers of goods such as coffee, cocoa, and palm oil must prove which plot of land a product came from and check satellite imagery to confirm it wasn't recently deforested. The difficulty is that companies won't disclose their suppliers or which ingredients they bought from whom, because that could reveal a secret recipe. Kleppmann hopes cryptography can show that accounting across a supply chain was done correctly without publicly revealing sensitive supplier or customer data.
Writing without AI, and moving between industry and academia
Kleppmann says they are not deeply involved with AI tools personally. They see them mostly through collaborators who use them well for software development, and they write very little code these days. They wrote their parts of the book entirely by hand and kept AI away from the text. They say this isn't a matter of principle and they aren't sure it's the right decision. For them, writing is how they figure things out, and figuring things out is the goal, so they have to do it themselves. They do see value in using AI to get feedback on ideas or test whether an idea holds up, in both industry and academia.
Asked for advice to students choosing between industry and academia, Kleppmann says the paths needn't be exclusive. Some of the best PhD students they've worked with spent a few years in industry after an undergraduate degree or master's, did real software engineering, then got bored and wanted more idealistic work or more freedom to choose research topics. Students who go straight from their degree into a PhD sometimes lack breadth of perspective. Movement in the other direction helps too. Kleppmann feels industry reasoning is often "short-circuit" reasoning, like adopting an idea because it appeared in a conference talk, whereas academia teaches nuanced, critical thinking about trade-offs and justifying why something is true. Kleppmann's conclusion is that people benefit from moving back and forth between industry and academia instead of treating them as separate career paths.
Should I consider multi-zone, multi-region, or even a multi-cloud setup? How much availability risk are you willing to take on versus the computational overheads, but also the human overheads actually designing and operating the system?
MapReduce is dead. Nobody uses it anymore. Other areas where we've increased the coverage are systems in support of AI, like vector indexes.
Is there any risk as software engineer that you're no longer incentivized to understand the underlying layer?
If you rely on a higher-level abstraction, you're no longer thinking about the lower-level details. If you're building higher-level business logic, actually, I think it's just fine.
LLMs increase the need for these formal proofs because we're live coding a bunch of stuff.
The reason I think that formal verification could become more important in the future. One is that
Designing Data-Intensive Applications has been the go-to book for anyone building large back-end systems. 9 years after publishing this book, the second edition is here. Martin Kleppmann is the author of this generational book. I sat down with him, and today we cover how working on Kafka at LinkedIn directly shaped the ideas that became the first edition of the book. What's new in the second edition, and why things like MapReduce got removed from this updated version. Formal methods, local-first software, decentralized access, and many more. If you care about how large systems work, where they're heading, and what the fundamentals are that don't change, this episode is for you.
This episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more.
This episode is brought to you by Sonar. Sonar, the makers of SonarQube, understands that code quality is about more than just avoiding syntax errors. It's about long-term maintainability by protecting the structural integrity of the system. As agents generate code at massive scale, they often ignore your system's structural integrity. This creates tangles, duplicated code, and other maintainability issues. These issues turn a modular design into a big ball of mud, making it increasingly difficult to extend. But, here's something that's really helpful, SonarQube's architecture management.
It moves architectural governance out of static wikis and into your automated workflow. It allows you to visualize your current architecture, define architectural boundaries, and manage architectural issues in real time. Whether it's a human or an AI agent at the keyboard, Sonar acts as a circuit breaker for structural decay. It ensures every commit respects the system's blueprint, protecting the long-term health of your most complex applications. Head to sonarsource.com/pragmatic to find out more.
So, Martin, welcome to the podcast.
Hi, Gregor. It's great to be here.
It's amazing to have you here. I don't think you need introduction to many software engineers, including myself. You're the author of this iconic book that I've had on my bookshelf for probably about 10 years, not much longer after it came out. Before we get into this book, which we're going to talk about, how did you get into the technology field?
Yes, well, I did an undergraduate computer science like many others, and then after that I wasn't quite sure what to do with my life, but I thought, well, starting a startup seems like an interesting thing to try. So, I started the startup having no clue what I was going to actually do, and then spent the first while searching around for things that might be interesting. The first startup didn't work out that well, but through that I met some others who then became my co-founders for the second startup, which worked better. And we sold that one to LinkedIn, and then after that I started being interested in teaching these distributed systems concepts. So, that's when I got into writing the book. And then during the writing of the book, I also switched over from industry back to academia.
Can we talk a little bit about your first and second startup?
Yeah, GoTest.it. This was like 2008 or something like that. It was the age where people were having really difficulties getting their JavaScript working cross-browser. Internet Explorer was still pretty big at the time. Chrome had just come out. All the browsers were incompatible with each other. And so, Go Test It was a cross-browser automated testing service for websites. It was based on Selenium, an open-source project that still exists. And the idea is you would write test scripts that automate a user clicking through the various interactions with a website and then just check that the right behavior happens. And so, yeah, it was based on Selenium but just provided as a hosted service so people wouldn't have to run various VMs with various operating systems themselves. It worked technically, but I found it really hard to actually get adoption for it. A lot of people building websites in theory said, "Oh, yeah, this is great. We need to test cross-browser." And in practice actually it was really difficult to get them to integrate it into their workflow and just get in the habit of using it and investing in writing the test scripts. So, that ended up not really going anywhere.
So, it's like there wasn't a business to be done or revenue to be generated in a meaningful sense.
Yeah, well, there's at least one, maybe two other companies from that same era that did manage to make a business. Sauce Labs is one that managed to actually succeed. But even for them it was a pretty slow-running business, I think. It was not an easy business to be in.
And for the startup, were you in the UK building it?
I was in the UK at the time, yes.
Was it bootstrapped? Did you raise some kind of funding? How big was the team? How can we imagine this?
It was mostly bootstrapped. So, I did a bunch of consulting in order to fund hiring some people and then hired some friends on the cheap to help contribute to actually building the product. And so, it was done all very cheaply. I had a very small amount of angel money in there, but mostly bootstrapped.
And then when you decided to not go forward with this, how did the next startup come? Rapportive, right?
Yeah, the second one was Rapportive. That went a lot better, so that was putting social media inside Gmail, basically. So, the idea was that if you get an email from someone you don't know, we had a little browser extension which manipulated the Gmail web interface, so that on the side, next to the email, we show you a summary social profile with a profile picture and job title pulled from LinkedIn and recent tweets pulled from Twitter and maybe recent Facebook posts or things like that. Just whatever we could find about that person, and put that as a social summary next to the email.
We started in 2010 or something like that. It then pretty quickly became quite popular, and so on the back of that we were then able to raise some money from Y Combinator, which was still fairly young at the time.
That was very young. You must have been one of the very early batches.
Yeah, I can't remember exactly when they started, but it was certainly in the early years. I think Y Combinator already built up quite a good reputation at the time, but it was still fairly small.
And then as part of Y Combinator, did you have to fly from the UK to San Francisco to attend that 10-week program, if I remember?
Exactly, yes. So, we initially came for the three months or whatever it was of the Y Combinator, but then we were able to get US work visas for ourselves and set up permanently in San Francisco.
How was that shift from the UK, where you spent going to university, your first startup, the first part of this, to coming to San Francisco?
It was very exciting because, you know, it felt like going to the center of where it was all happening, really. And we started it up not knowing anybody at all. We knew like one or two people in the entire Bay Area, but we contacted them and they introduced us to more people and they introduced us to more people. And so we were able to pretty quickly actually build up a network, and that's something that I really appreciated, that it was actually so open to outsiders like us who could just basically turn up with an idea and an early stage startup, and we managed to raise some money and managed to actually become somewhat established in the Bay Area.
And can you tell me how the company grew and at what point did the LinkedIn acquisition offer come in? And how can we imagine even, you were a founder of this company?
It was about in 2012 that we sold it. And we were five people at the time. So it's all still pretty small. Not vast amounts of money involved, but it was a success, I would say, for everybody involved. The acquisition process itself was fine. As always with these kinds of transactions, there were twists and turns and moments where we thought it would all fall apart, and then we were almost running out of money and hadn't really succeeded in raising another round. So we kind of had to sell or shut down. So we were under quite a bit of pressure. We couldn't reduce our own salaries because to do so would have violated the conditions of our visas. So we were in a slightly stuck situation. Given our lack of leverage in that situation, actually I'm pretty happy how it all turned out.
Yeah, it's nice that, you know, from 10 plus years we can talk about this honestly, because oftentimes you see an acquisition by LinkedIn and of course you might ask the founders and they would say this was either our dream or our goal or we will do so many things together, but something that you don't often hear is, well, that there was pressure on all of us as well. So did you go into this wanting to sell the company because you saw that either you needed to raise a new round or you sell to someone, and then you found LinkedIn to be the only or the best option to go into?
We tried a little bit to see what revenue generating options we had and hadn't really managed to make that work. So we were just burning money and our user growth was okay, but not really enough to go and raise a big round. So we were a little bit stuck there, and selling the company seemed like the least bad option there in a way. And I'm pretty happy how it turned out because, you know, LinkedIn was great actually. They were very good to us. They allowed us to operate as essentially an independent team within the company.
So your team stayed together?
Our team stayed together. We continued working on the product that we wanted to make.
Oh, so you got to keep working on Rapportive.
Yes. Well, actually, Rapportive, the Gmail browser extension, sort of got put on life support, but we were working on a new product at the time which did eventually get released under the name LinkedIn Intro. It kind of got a slightly weird reception at the time and it ended up getting shut down shortly after we released it. There's kind of a longer background story there, but I'm still really happy with LinkedIn, how they gave us the freedom to do this and allowed us to launch this product, and even though it didn't succeed, you know, they were very good to us throughout that process. And then after that got shut down, our team got disbanded. But we had a good run within LinkedIn building this product.
What tech stack did you work with at the time? What did you use?
Rapportive was fairly unexciting. It was a Rails app with a Postgres database basically, and some Redis and some similar things like that mixed in. So actually, you know, nothing particularly revolutionary. We essentially built a graph database on top of Postgres. So there was a little bit of technical interest in there, but you know, nothing particularly outrageous.
And then you spent time after LinkedIn Intro, you still worked inside LinkedIn. As I understand, you worked on data infrastructure, right?
Yes, data infrastructure. After our team got disbanded, I switched over to the stream processing team. So Kafka had just been developed at LinkedIn and
Yeah, they developed it, right? Oh, it was just being open sourced.
Yeah, I think it had just been open-sourced, and then I got to work on Samza, which was a stream processing framework on top of Kafka.
I always wanted to ask this question, so I think it comes here. Why did LinkedIn build Kafka, or develop Kafka? It's now such a foundational technology, but I was always curious why a company felt the necessity to build this thing that seems pretty generic, and it seems everyone would have needed it.
Yes, so I think Jay Kreps has a pretty good blog post from that era called The Log, where he explains his motivation behind Kafka and, you know, why make it an append-only log rather than a traditional message queue or something like that.
I think the motivation was really about data integration, because there were a whole bunch of databases and event generating systems, you know, like activity events from users, for example. They were all generating data in a sort of stream shape. And then a bunch of downstream systems wanted to consume this, like wanted to get it into the data warehouse and wanted to be able to get it into the Hadoop cluster at the time in order to run machine learning and things over it. And there was just this data integration problem of actually how do you physically get the data out of one system and into another? And Jay designed Kafka as this integration point, essentially almost kind of the lowest common denominator, but still a general purpose abstraction for integrating various sources and downstream data sinks.
Working at LinkedIn, you know, with Kafka and at LinkedIn scale, what did you learn or what surprised you about working at this type of scale? As I understand, this was the first time that you hands-on worked on a really large system, right?
That's right, yes, because previously the biggest company I had worked in was Rapportive with five people. We had a sizable database, but it was still a single-instance database and not really that big in the grand scheme of things. And then suddenly I was at LinkedIn, and oh, we got to use their big Hadoop cluster. That was fun. Hand coding MapReduce jobs in Java at the time. And so I learned a huge amount there. Especially when the stream processing ideas came up, and Jay was evangelizing the use of Kafka and the things you could do with it. That was kind of a revelation for me really, where suddenly it felt like, ah, this kind of makes sense. I've started to understand how these various data systems fit together, what they have in common, what the fundamental principles are. And so that experience then fed directly into the writing of the book.
At what point did you decide to leave LinkedIn? I'm looking through your career: start out in the UK, do a startup, do a second startup, Y Combinator, move to San Francisco, get acquired by LinkedIn. The arc that most people would draw would be, okay, do something more in Silicon Valley, or maybe start another startup, etc. And instead you decided to leave LinkedIn.
Yeah, so first I decided to move back to the UK actually, and I continued working for LinkedIn remotely.
Okay.
That was mostly because my girlfriend at the time, now wife, was still in the UK, and a long distance relationship is not a lot of fun. And I didn't feel that at home in the Bay Area. So I wasn't really encouraging her to move to the Bay Area either. I thought it was better for me to go back to Europe, and I'm very happy with that decision. I still have a lot of great friends in the Bay Area. I
love it as a place to visit, but I wouldn't want to live here, honestly. And then I was still remotely working for LinkedIn, and that worked all right for a while when I then started writing the book. LinkedIn even gave me 50% of my time free to work on my book alongside my software engineering duties, which is really great.
Amazing.
Yeah.
That's so nice of them.
Absolutely. And they don't have to do that, and LinkedIn didn't directly get anything out of it in response other than a book that they could use for internal training purposes.
Well, shout out to LinkedIn for this.
Yeah, absolutely. So, then I did find that actually trying to write a book in parallel with doing a software engineering job and being on call, et cetera, I just wasn't able to do it. It's just too much context switching, and it's very easy for the urgent things from the on call to dominate, and then not to have the freedom that you need in order to write something new. And so then after a while I decided, okay, it's probably better if I focus full-time on the book. So, I then left LinkedIn and just took a sabbatical, unpaid sabbatical, i.e. unemployment, to just focus full-time on the book for a while. And then it's only after that that I actually even considered getting into academia.
So, how did the idea of a book come? What was the point where you decided you would write, and in your mind what were you designing to write? Was it already, you know, this book with this layout, or you had an early idea back then?
I had an idea that, of course, the final product ended up looking somewhat different, but the overall goal, I think, stayed the same. So, I knew I wanted to write something that was a broad conceptual overview. So, not about how you use any one specific system or tool, but comparing the trade-offs between many different types of tools. And I knew that I wanted it to be practitioner focused, not a theoretical textbook, but something that people could use to build real systems. That was basically the goal with which I approached this, and this was exactly the book that I wish I had had when I was starting out and working at Rapportive, for example, because we were all searching around in the dark where we're having performance problems with our database and we had no idea what to do, basically, because we're totally lacking the foundations to actually understand what was going on and how to diagnose the issues. And so I felt that, well, if I had had a bit more background on how these data systems actually work internally, then I could have had an intuition about how to debug these kinds of performance issues. And then after a while, after I'd learned more about how data systems work, I thought, well, okay, it's time to write this down so that others don't have to learn it the hard way, but can hopefully just get a better idea of how these systems work and thus be better at managing their own data systems.
To start with, how did you learn about, for example, how databases work? Because again, from your story at Rapportive, you built systems, you've had some performance issues at a smaller scale, to be fair, compared to LinkedIn, then you worked at LinkedIn and you saw a little bit of how the sausage was made, but I know a lot of software engineers who have been in this path and they still don't really know how the fundamental systems work. They just know, "Okay, we have a platform team inside our company and they build it. I could read the RFCs, but it's a lot of work." Or the planning docs, "I could look at the source code." It feels to me that even at that point, you just went down and tried to dig in. What resources did you use? How did you find out those basics which you later put into the book?
A lot of it was just kind of being curious and talking to people, actually, and just asking them lots of questions. At LinkedIn, there were a bunch of senior data systems engineers who understood this stuff very well, but hadn't maybe necessarily written it down.
Mhm.
And so I just talked to a bunch of them and quizzed them, and that way started building an image in my own mind of how this stuff works. And then once I sort of got the basics from these conversations, then I was able to go and read research papers, for example. They go into much more detail of exactly how and why things are designed in such a way, but, you know, it is time-consuming to read those things. So then what I tried to do was pull out what are really the essential ideas. I just read a ton of blog posts as well. And so, the reason why you see so many references at the end of each chapter in the book is, well, that is actually the material that I myself used in order to understand what was going on. And then I thought, well, okay, if I found these things useful, then I'll also cite them in the book as a way for any reader who wants to go beyond the basics covered in the book. Here are some good sources for further reading.
Yeah, the structure of the book, this first book at least, is foundations of data systems, distributed data, and derived data, if I understood it as three big parts. Did you already have the structure in mind when you started writing the book, or did it shape as you went?
This three-part structure, it's not that critical in the design of the book, really. That's sort of more after the fact. I thought, oh well, it seems like we can group the chapters into roughly this sort of structure. But the topics of the chapters were more or less what I had envisaged. So, I knew that I wanted to talk about what a transaction actually is. I knew that I wanted to talk about replication. I knew that I wanted to talk about sharding or partitioning. I knew that I wanted to talk about consistency and consensus. So, those sort of high-level topics, I think, were clear from my initial book proposal to the publisher. The details within each chapter, that is something that I often figured out once I got to that chapter. So, I wrote one chapter at a time and started each chapter's work with just a lot of background research to actually get up to speed on the topic myself. And it's often only then that, say, for replication, I decided, okay, well, it seems like the three major ways of doing this are single leader, multi-leader, or leaderless. Okay?
Uh-huh.
I would decide on that structure essentially when I started writing each chapter, and then tried to fit the various points I wanted to make into this narrative structure.
As a fellow author who also wrote a book, one thing I've noticed is there's a bit of a parallel between estimating a book and estimating a software project, in that you come in with an estimate, and if you've never done it before, you tend to be wildly off. How was this in your journey? And in addition, you also had a publisher, and publishers are a little bit like project managers. They, you know, like to have a schedule. They like to try to keep you on track. They like to ask, when is it done? How did you manage that part as well? And in the end, how long did you estimate it would take when you started, and how long did it actually take?
As always, it takes vastly longer than expected. It's the same for software projects as it is for writing, I think. So, I think it took me about four years to write the first edition. And that was not four years of full-time, maybe two and a half years of full-time equivalent or something like that, but written over the course of about four years. So, it definitely took a long time. The publisher deadline I missed by a ludicrous margin. I think I missed it by about two and a half years or something like that. But fortunately, O'Reilly were pretty laid-back with the first edition and were happy for me to just take my time and make it good.
When it came to the second edition, then actually O'Reilly got a bit more aggressive and pushy about sticking to deadlines. I guess by that point, the book had been established and people were waiting eagerly for the second edition. So, I kind of understand the desire to want to accelerate it, but at the same time I really appreciated the freedom that I had for the first edition to work on my own schedule. And I had a bit less of that with the second.
The tagline for the first edition, which I believe is the same as the second edition: the big ideas behind reliable, scalable, and maintainable systems. Reliable, scalable, and maintainable, what do these adjectives mean to you?
Yeah, so they're all slightly vaguely defined, right? So, there's not a formal definition of those things, but for me, reliability means fault tolerance, primarily. So, meaning that a system should, on the whole, continue working even if a network link is interrupted or a node crashes or something like that. So, a lot of the book is about techniques that support fault tolerance, like replication, for example. So, that's reliability. Scalability is one of those terms that get thrown around a lot, and it's sort of
So much.
And it's like fashionable and cool to make things scalable, you know, because it suggests success and millions of users. And so, of course, everyone wants things to be scalable because everyone wants success. For this book, I tried to take a bit more of a dispassionate kind of approach and said, "Scalability is just what mechanisms we have for dealing with changes in load. If load increases, how can we add computing capacity to a system, for example, so that the system still continues working?" And then the techniques that you use to achieve scalability, well, they are like sharding, for example.
But in this case, with your definition of scalability, do I understand that you're mostly referring to horizontal scalability, so that you can add compute up or down, pretty much?
Yeah, I guess because that's the more interesting one. Like, yes, you can always buy a bigger machine, but what's interesting about that? Exactly, there's just not that much to be said about it. I mean, there are details of how you scale even on a single machine, but I think part of what has become interesting about modern cloud services and just back-end services in general is how they've introduced this idea of horizontal scalability and shared-nothing systems. So, we can build systems that, you know, are able to cope with very high load even if the individual components are just fairly cheap commodity machines. But maybe part of the scalability story which I wasn't thinking about as much at the time, but started thinking about more recently, is not just scaling up, but scaling down as well. So, actually, how do you run a service in such a way that if it has a very small amount of load, it's really cheap to run it? That's, in a way, the same question as how do you continue running a service if it has very high load. Generally, you just want the cost and the computing capacity to be roughly proportional to the load that you have. And at the low end, that means actually being able to scale down to something that is extremely cheap to run. And that's not necessarily a given. That's something that is hard with on-premises software, for example, because if you've got a physical machine, that's a unit of deployment. And yes, you could carve it up into two dozen virtual machines and make those small virtual machines, but it still requires some sort of resource allocation. So, part of what's interesting about some serverless systems, for example, is actually their ability to scale down and say, "Okay, if you're going to handle just three requests per day, that's just fine as well."
Can you tell me about the second edition? When did the idea come about?
Yeah, it had been clear for a couple of years that a second edition was needed, just because the first edition was getting a bit dated. There were changes in technology that just hadn't been reflected in the first edition. So, I wanted to update it, but, you know, I now have an academic job. I'm actually doing research and teaching as my main thing, and updating the book is just a sort of sideline business in some sense. So, it actually took quite a while to make progress with that, because I was always doing it alongside other projects, and essentially back to that context switching problem that I had while writing the first edition, but just now with an academic job that I didn't want to just drop, because I actually quite enjoy it. Initially, then, I made very slow progress with the second edition. And also I kind of realized that I had slightly lost touch with current industry practices because, you know, I had switched over to the academic side, I'd gone much deeper on the theory, but I was no longer up to speed on what people were doing with, say, data lakes or things like that. So then at some point I remembered Chris Riccomini, an old colleague from LinkedIn. I had worked with him on the stream processing stuff some time ago.
With him? He's the author of The Missing README.
Exactly.
Wow, what a small world.
Yeah, and I had read Chris's book The Missing README and thought, oh, he's a great writer. And I had worked with him as a software engineer and found him a great colleague. And also he had been writing this newsletter called Materialized View on the latest trends in data systems, essentially, and become a startup investor in that space. And so at some point I thought, well, actually, I have to get in touch with Chris and ask him whether he wants to help out with the second edition. And he was keen to do that. And that turned into such a good collaboration, because he was up to date on what the cutting edge was in terms of technology in industry. I had strong opinions on how to teach, essentially. So, how to explain things in the book, make sure that we were explaining everything in a way that was very precise, very carefully chosen words, but at the same time very accessible, so that it's hopefully easy to read. And so we took essentially my writing style plus Chris's knowledge of the latest industry trends to bring the book up to date. And that was a great collaboration.
What are the big things that you added, and which of these did you know would be missing, and which ones did you realize during the writing process that, okay, this needs to be in here now?
Yeah, so the thing we knew from the start that we wanted to reflect was cloud-native systems architecture. It's a bit of a vague term, but what I mean with that is essentially building data systems on top of cloud services as the foundational abstraction. In the first edition, the assumption was basically that you have some machines, each machine has some local disks. You can run a database instance on a machine. It will write its data to the local disk. If you want to replicate it to another machine, then, well, the database software will replicate it at the database level to another machine, which will also write the data to its local disks. For a long time, that was exactly the way computers worked.
And now suddenly people are building databases on top of object stores, for example. And now the replication happens at the object store level, no longer at the database level. Or maybe there's still some replication at the database level, but it really changes the nature of things if you're building on top of an object store. And this is different from, say, building on top of a virtual block device like EBS, because these block devices, although they are cloud services, still offer the abstraction that is a sort of single-node operating system abstraction of a block device on top of which you run a file system. Whereas an object store is just a brand new abstraction. It looks
different from a file system. It behaves differently. And so then building on top of that as a foundational abstraction is something that people were starting to do at the time of the first edition, but since the first edition, that has really taken off. A whole lot of systems have been built in that style now. And so that's an idea that we really wanted to incorporate and we've woven that in throughout the book. So it's not just one section here, but it's an idea that we've integrated throughout the entire narrative.
There's now a lot of managed services as well. The primitives that we use, but there's also so many managed services that all the cloud providers use, and a lot of engineers often just use the managed services as is because they take care of replication. They have SLAs for uptime and so on. But when you build on top of these things, then you kind of use those as primitives as well. Is there any risk as a software engineer that you're no longer incentivized to understand the underlying layer, or are we building better systems because of that? How do you think about this? It feels there's a move of abstraction because of cloud, right?
Yeah, it's definitely a shift to different and higher-level abstractions. But, you know, that's been the story of the entire computing industry since the start. It's building new abstractions. So, it is true that if you rely on a higher-level abstraction, you're no longer thinking about the lower-level details. And so, if you're using a programming language with a garbage collector, you're no longer thinking about memory allocation. And so, is that a loss? Well, maybe. If you're building low-level systems, you should still have to care about memory allocation. If you're building higher-level business logic, actually I think it's just fine for people not to care about memory management.
So, I think there's an analogous thing here with data systems: if you're building the higher-level systems that don't need to particularly care about the underlying infrastructure, then that's fine. Just use the higher-level abstractions, nothing wrong with that. But somebody still has to build those lower-level abstractions from lower-level components. Somebody's got to implement the cloud services.
Martin talked about trade-offs that come with using cloud services. And this is a good time to talk about our season sponsor, WorkOS. If you've read Designing Data-Intensive Applications, you know that building systems that scale is all about trade-offs, but one thing isn't a trade-off, that's enterprise features. The moment you land bigger customers, you need SSO, directory sync, RBAC, audit logs, all the things they expect out of the box. Building that yourself can take months. WorkOS gives you APIs to ship it in days, so you can stay focused on your core product. That's why companies like OpenAI and Anthropic run on WorkOS. Visit workos.com to learn more.
I'd also like to mention our presenting sponsor, Statsig. Statsig built a unified platform that enables both experimentation and continuous shipping. Built-in experimentation means that every rollout automatically becomes a learning opportunity with proper statistical analysis showing you exactly how features impact your metrics. Feature flags let you ship continuously with confidence. And because it's all in one platform with the same product data, teams across your organization can collaborate and make data-driven decisions. To learn more, head to statsig.com/pragmatic. With this, let's get back to Martin and the trade-offs that come with using cloud services.
And so, those people will have to then specialize even more in actually the details of how you engineer those cloud services, how you make them reliable, how you operate them, and so on. The skills are still there, it's just a bit of specialization happening, where some people can worry about the higher-level things without having to concern themselves with the lower-level things. Some people focus on the lower-level things and treat the higher-level aspects as their customers.
Interesting. So, it sounds to me that if you're an engineer who is utilizing a lot of these services, you might not need to know how they exactly work.
Yes, and I would say the underlying philosophy of the entire book is to give people insights into just the sort of essence of how the systems work internally, so that if, for example, they start having weird performance behavior, you can have a bit of intuition for why it's doing that and how you might solve it. So, for example, say the storage engine chapter tells you about how B-trees work and how log-structured LSM-tree storage engines work. And the book is not intended for people who are going to actually build their own databases and implement their own storage engines. If you want to do that, you have to go into much greater depth than this book covers. But the idea is that as an app developer, if you know just a little bit about how the storage engine works internally, you'll be in a much better place to use it in a way that gives you good performance, for example, and to diagnose any issues. That philosophy we've kept also in the context of cloud services, where yes, a cloud service hides some of the operational details that app developers don't need to think about anymore, but they should still know a bit about how they work internally just so that they can use them effectively.
I guess our argument is that trade-off: deciding on which service to use, which characteristics to look out for for your use case, right?
Exactly. And you know there are huge differences of, say, if you're doing analytics, whether you're using row-oriented storage or column-oriented storage. That's a bit of a technical distinction and takes a little bit of background reading to even understand what that means, but it has a massive performance implication in terms of the final behavior of the system. And so those are those places where I feel like knowing a bit about the internals is actually like a superpower.
Yeah, and I guess for engineers, the one thing that we always need to argue about, or should need to argue about, is at the very least cost versus performance. And by performance I mean latency to the user. And of course resilience: if something happens, you know, a machine goes down, a zone goes down, a region goes down, how our product is affected and what's acceptable.
The basic idea there seems to be how much availability risk are you willing to take on versus the overheads, both in terms of the system itself, like the computational overheads, but also the human overheads of actually designing and operating the system. And the cost overhead. Yeah, exactly. And so yes, you can have a system that is more able to tolerate various types of faults, but which is more expensive to design and operate, versus a simpler system that, you know, might go down a bit more often, but which is cheaper. And there's no right and wrong with that really. Everyone needs to figure out where they sit on that trade-off space themselves. And I would say that multi-region is pushing in the direction of higher availability because it means you could tolerate the outage of an entire region. But then it has implications on the consistency model that you can get across different regions, for example. So, that's a trade-off that the book tries to make very explicit to help people reason through what is the right choice for them.
In terms of multi-cloud, for example, one thing that I've been concerned about just in the last month really is European dependence on US cloud services. Yes. So, what if geopolitics was to go horribly wrong and tensions escalate and Europe finds itself suddenly locked out of US cloud services? I hope that doesn't happen. I still think it's fairly unlikely, but it's no longer unthinkable. And as a result, I, coming from this European perspective, have been thinking a fair bit about how we can engineer systems to be resilient against that sort of thing. And that's, you know, not just a regional outage, it's a business risk essentially, and a multi-cloud setup could help mitigate against that sort of risk, so that at least, for example, if one company locks you out, then you could still have systems on another company. Again, that's very much towards the expensive but high-availability, risk-reduction end of the spectrum. But for the people who have, you know, really critical workloads where they think the geopolitical risk is a significant enough risk, I think it's seriously worth considering that kind of setup.
I'm thinking that as engineers, we do have the responsibility, because who else will do this?
Yes, totally. And I totally agree with you as well that understanding what the risks are and communicating what the trade-offs are, I think is going to be a core part of our roles as engineers moving forward as well. Maybe as AI writes more and more of our code, it's less about the details of how you express logic in a particular programming language and much more about those kinds of high-level trade-offs.
How has the definition of scale changed in this book? Because as we talk about cloud, before cloud, building a scalable system sounded pretty involved, because building a horizontally scalable system is complicated. All the pieces you need to put in, in the first book you detail a lot of this. With cloud, a lot of the services actually do define how they allow horizontal scaling and what the trade-offs are. Do you feel that it's made it a lot easier to reason about scale, scalability, when you're using these primitives?
So I think achieving really high scale is still challenging, because even though we have cloud services like object storage, for example, which provide you this very elastic storage model, at least you don't have to worry about capacity planning on your disks anymore and running out of disk space, because those kinds of operational things are taken care of. But if you need sharding, for example, that's something that actually does reflect on the application code as well. You can't really make that entirely transparent. And so if you are at a sufficiently large scale that sharding is required, because a single machine is not powerful enough to process your workload, then I think even with cloud systems you still have to do quite a bit of engineering thinking of how to realize that.
Where I think the cloud has helped quite a bit is actually at the lower end, of scaling down. If you want to have a very lightweight service that processes only a small number of requests, what we've got with serverless systems, being able to very quickly spin up and spin down a very lightweight instance, that's quite a good innovation that has enabled those very low-scale services. And that's something that would be much harder to do without cloud services, because you would have to statically allocate a certain amount of memory and certain CPU resources to a particular virtual machine.
I love serverless. I have a small website that runs on serverless and my bill is like 13 cents per month because it has very little load.
Absolutely. It's just making more efficient use of computational resources.
Let's talk about sharding. In the first book, and when you wrote the first book, when I was working at Uber, we talked a lot about sharding and there were a lot of internal implementations. Our interviews involved asking about sharding because we were designing systems that were sharded. I did sense that over time, as cloud systems started to become available that give you turnkey solutions, that act more like platforms where you send the data and it takes care of these things, fewer engineers have to actually implement sharding. With cloud-native systems, in your research, what have you seen? What are the cases where putting sharding in place is still important, and where are the places where it might have just disappeared as a concern? I mean, it's still nice to know, but you might not have to implement it.
I think it's probably less of an effect of cloud and more of just hardware getting more powerful. Actually, a big machine nowadays can do a lot. And that means that more and more workloads you can just run on a single machine, and that is sufficient actually to achieve quite significant scale already. There are still concerns of how you actually efficiently make use of the hundreds of CPU cores that you have on a single machine. So, parallelism is still a required thing to think about there, and sharding is one way of achieving parallelism. But at least this sort of sharding across multiple machines has maybe become less of a pressing issue, just because more and more workloads can just run on a single machine. Some people still have very large-scale workloads that do have to be sharded across multiple machines. So, it's not going away entirely. And replication is still relevant even at smaller scales, because that's for fault tolerance, that's not for scalability.
You have a chapter called The Trouble with Distributed Systems, which goes through a lot of things that can go wrong. Without going through the whole chapter, can you recall some of the things that are memorable to you, or some of the things that you feel are important to remember?
Yeah, the whole idea of this chapter is that in distributed systems theory, there are certain things that we tend to assume. For example, we just assume that there's no upper bound on how long it might take for a message to go over the network. So, you send a message, it might arrive within 100 microseconds, or it might take 10 years. And distributed systems theory just doesn't make any assumptions about that sort of timing if we can avoid it. Or rather, some theory does make those assumptions, but it's a dangerous assumption to make, because occasionally the network delay does become much higher than what is typical.
Another thing is about crashes. For example, distributed systems theory just says nodes can crash. But what does that actually mean? What in practice does it mean for a node to become unavailable? Because it might be a software crash, but it might be a hardware failure, it might be somebody unplugging the power cable. It might be that the node is actually still running, but it's just become disconnected from the network. The point of this book chapter really is to defend and justify those theoretical models that we use for analyzing distributed systems, and just giving a lot of stories and case studies that show that, you know, actually tons of stuff does go wrong, and don't believe anyone who says, oh, failures are rare, don't worry about it, it's fine. The moral of this chapter is really that actually, you know, if you want to make things reliable, you really do have to worry about a whole bunch of weird, unusual, but certainly possible edge cases. Timing is another one of those things. You know, it's very easy to assume that your clocks are correct, and most of the time the clocks are pretty correct, but we just can't rely on it because actually they're just not precise enough on the whole. And so a lot of it is about this: it's very tempting to make certain assumptions that things are well behaved, and in distributed systems we just have to try to get away from those assumptions if we want the systems to work reliably even in the face of things going wrong.
It was a really fun chapter to write because, you know, it's essentially a big collection of stuff that has gone wrong, and so I went through a bunch of postmortems published by various tech companies, for example, in order to see, okay, what was the root cause of how things went wrong, and what kind of lessons can we draw from this that apply to the book in general, and there you
You know, there's some fun stuff like the sharks biting undersea cables and damaging them, and that just makes for a great story. And then I hear that in recent years the shielding of undersea cables has got better and therefore the sharks are not biting them anymore, but instead the cows on land are stepping on cables and occasionally causing network interruptions that way. And that sort of thing just makes it a bit more fun.
That chapter is so interesting also because, depending on what kind of teams you work on or what kind of people you talk with, when I talk with the S3 team, for them that whole chapter is just their day-to-day. It's not a weird thing when a hard drive goes up, or, okay, maybe it might be a weird thing to have a fire in a data center, but they're prepared for all of those things. They're at the scale where these things just happen on a regular cadence because they're one of the largest scales. Whereas at a smaller company, even if you read this chapter, you will treat this as, well, this could happen. But when it actually happens, it will be a once in 10 years and it will be a big deal.
Yeah, but I think there's no right answer. So, it's a trade-off between risk and cost, broadly speaking. And that means a business decision has to be made in terms of where the business wants to lie on that trade-off. And so, the goal of this chapter is really just to give people the information in order to make an educated decision. But I don't want to make that decision for people. That's for businesses themselves to decide.
That's very clear. Have you come across some concepts or services mentioned in the book in the first edition and now in the second edition that are becoming either more popular or less popular over time, more or less referenced by your readers, thinking about things like streaming systems, batch processing, or anything else?
Yeah, so there are some things that we've been able to take out of the book compared to the first edition. In particular, for example, coverage of MapReduce was quite detailed in the first edition. But basically, MapReduce is dead. Nobody uses it anymore. Its successors, like in the form of Spark and Flink, for example, they are used. And so, we still reference MapReduce in the second edition, but more as a learning tool in order to understand how these kinds of partitioned, sharded batch processing systems work. So, that's one thing where we've been able to reduce the coverage.
But other areas where we've increased the coverage are, for example, systems in support of AI. And so, even though this is not an AI book, there are still data systems concerns that arise when needing to support AI applications. A classic one is vector indexes, for example. And so, we've added some coverage of vector indexes to the storage engine chapter. It fit in really well there because it already covers various different indexing strategies anyway. And so, vector indexes, you know, it's just another indexing strategy.
We also added some coverage of data frames, for example. That's not an exclusively AI thing, but data frames are quite a good data representation for training data, for example. And that was not one of the data models that we discussed in the first edition, but we decided to add it to the second edition because it has actually become a very important data model that people are using alongside all of the classic data models like relational and graph and JSON documents and so on. And so, these are places where we've just expanded the coverage a bit to reflect the kinds of systems people are building, for example, to support AI, without it changing the direction of the book entirely.
The final subsection in the first edition, the first few, I guess, subparts were titled "Doing the Right Thing." And in the second edition, this has its own chapter. The final chapter is "Doing the Right Thing." And I quote a little bit from it: "We, the engineers building these systems, have a responsibility to carefully consider those consequences and consciously decide what kind of world we want to live in." Can we talk a little bit about this section and the importance of it?
Absolutely. Yeah, so the motivation for putting in an ethics section there in the first edition was that I just felt it had been quite ignored as a concern during my time in industry. Especially in startups, people were very focused on building a product that their customers would love and really deprioritizing these sorts of ethical questions in the process. And so, for example, with consumer-facing products, it might be that the products are very much geared towards essentially data harvesting, collecting behavioral data, because that's what can be monetized in the form of advertising. And there seemed to be just very little reflection on what was good and bad about these sorts of things. So, I really just wanted to encourage a bit of thinking there.
Not really wanting to prescribe a particular approach too much there, but at least to point out that there is this thing such as data protection legislation now, which we do have to think about in the architecture of our data systems. And there is an ethical responsibility. You know, people say that you get into tech in order to change the world. If you want to change the world, then thinking about the impact that your technologies have on the world is part of your job. It's a really essential part, really, and something that engineers are often prone to ignoring as we focus just on the technology and less on the effects that that technology will have out in the real world.
And so, this chapter is really just an attempt to get people thinking about it a bit. And it's sort of a reflection of my own process as well, because as I started working on these systems, I didn't really think about ethical things particularly either. So, I felt like I had to put that section in there for myself as well as for the readers, because it was my own way of grappling with these questions a bit.
Is it fair to say that as engineers building these systems that will have an impact on wider things, potentially society-wide impact, we are just in such a good position to directly influence and maybe even change course? So, do I understand that this section is a bit of a reminder that by building it, we have a huge opportunity to shape these? We probably have a lot stronger voices, maybe as strong voices as the regulator might have years down the road, right?
Exactly. I think engineers have a very strong voice there. And like we talked about earlier, engineers need to articulate trade-offs in such a way that business leaders can then make educated decisions about how to address those trade-offs. And part of those trade-offs is pointing out risks. And risks include not just technical risks, like the data might get corrupted, but they include societal risks as well. For example, what negative effects, what harms might arise from this technology? What sort of unintended consequences, possibly? Or what's the risk of reputational damage if it turns out that the technology has some harmful effects? You know, that can reflect badly on the company that made it. And that has to be part of the trade-off discussion. And I just want people to make intentional and deliberate decisions about these kinds of things and not just sweep it under the carpet.
One of the hot topics these days is of course AI. And you've written a very interesting post about this just in December about formal verification and your conviction that formal verification might be more important with AI. For those of us engineers who have heard of formal verification, can we talk about what this is and how you envision this becoming more important?
Yeah, so there's a whole range of formal methods. One approach is to, for example, use a specification language like FizzBee or TLA+ or something like that to describe the expected behavior of a system at a high level and then use a model checker, which is essentially like a randomized test case generator, to just play through a lot of scenarios and see whether the system has those desired behaviors in all the different scenarios. That's the sort of intro level formal verification, I would say.
The more advanced level is to use actual formal proof. And in that case, you can write a specification of some system in a formal language, usually using mathematical notation, and then make a mathematical proof that a certain algorithm or certain implementation always satisfies that specification. And the distinction to testing there is that, well, in testing you just try a couple of examples, give the algorithm some example inputs and check whether you get the expected output in those particular examples. But a proof can reason about potentially infinite state spaces. So, it can tell you things about every possible thing that could possibly happen in the entire universe, and show that, for example, a certain safety property is always given in those.
Formal verification is a lot of work. I never used it in my time in industry because it's just too time-consuming, basically. I only got into formal verification when I was in academia and I could afford to take the time to spend a few months proving an algorithm correct. But there I started finding this very useful, especially if I was working on very subtle algorithms where it's very hard to tell just from reading the implementation whether this actually is always correct under all possible cases. But if it's an important algorithm where, for example, it will corrupt data if there's a mistake in it, or it will have a security vulnerability if there's a mistake in it, then when it's high-stakes things like that, I feel it's worthwhile to have a formal verification and to really make sure that the code really is correct.
And so, I've done some formal proofs using the Isabelle proof assistant, for example. There are a couple of others as well, like Rocq and Lean and so on. These proofs are really hard to write. It takes a long time to learn the language of writing those proofs. And then even once you know the language, it's just really laborious to actually write the individual proof steps.
And when you say it's hard to write, just as someone who knows how to code in so many different languages, can you just explain what it means to be hard to write? Does it feel like a strict programming language with all sorts of rules, or lots of math formulas? What makes it hard for you to learn it and get good at it?
Yeah, so you're trying to make a proof that a certain piece of code always satisfies a certain property. In some cases that property might be quite easy to specify. Let's say, as a really simple example, you have two lists and you want to concatenate them, and then you want to prove that the length of the concatenated list equals the sum of the lengths of the two individual lists. A very, very simple property. How would you prove something like this? Well, you would have a function that concatenates two lists, and then you would probably do a proof by induction over one of the lists that shows that, okay, well, if you have one list of length I and another list of length zero, then the sum of the two is I. If you have a list of length I appended with a list of length one, well, then it's I plus one, and so on. And then by using a proof by induction, you can then show that the length of the concatenated list is I plus J, where I and J are the lengths of the two input lists, for every possible value of I and J.
And this is something that, in a test, you would maybe test for the cases of J equals zero, J equals one and J equals five, and then you're done.
And J equals INT_MAX. Yes. But the edge cases, that's what we do. That's how I write my unit tests.
Exactly. And so this is a trivial example. With list concatenation, you can easily just read the code and convince yourself that it's correct. But if it's a much more complex algorithm, then our brains just can't grok the algorithm well enough to really convince ourselves that it's correct if we don't prove it. And that's where these proofs then become handy.
If I'm an engineer and I would be interested in getting into formal verification, for example because I have the notion that it will be more important with AI, and of course it will be easier to write these things, where would you point engineers to get started, or how did you get started in this field?
I would suggest starting with model checking. So, something like TLA+ or FizzBee is much friendlier to get started with compared to proof assistants like Isabelle, Rocq, and Lean. These proof assistants just require a whole lot of additional knowledge. And the resources for learning about writing these formal proofs aren't, to be honest, particularly good. I haven't really found really great books on it either. The way I learned it was by working with some colleagues in my lab who had learned it through years of prior experience. And I just sat down with them and paired with them at a desk, where I described the thing I was trying to prove and they showed me how to prove it step by step, how to break it down.
I'm interested to see if your thinking will be correct, which is that this thing will go more mainstream. And hopefully we'll have better books and resources for it as well.
Yes, I do hope so too. The reason I believe that formal verification could become more important in the future is there are several aspects to it. One is that LLMs are getting increasingly good at writing these proofs. And if we don't have to write the proofs by hand as humans, it just becomes feasible to do them in situations where previously it would not have been economical. But also LLMs increase the need for these formal proofs because, you know, we're vibe coding a bunch of stuff. If we have to manually review all of that code, then that will become the bottleneck. So, we can't really have humans reviewing all of the generated code either if we really want to get the benefits of AI. So, we need some automated way of checking whether the code is correct.
And writing lots of tests is a very good starting point. But the thing that proof can do that tests can't is to consider absolutely every possible thing that could happen. And that's really important in a security context, for example, where it just takes one little bug to create a vulnerability that destroys the security of the whole system. And so, I feel for those domains where we really want to ensure there's a complete absence of bugs, that's the kind of place where formal verification can really shine. And I'm hoping that LLMs will actually make that a lot more accessible to people who would have previously not considered using formal verification because it was just too hard and too expensive.
You've worked in the industry and then you went into academia. Can you tell us what the difference is? Myself and most people watching work in what you would call industry, in the tech industry, where we work at different companies. You know, we're bootstrapping our own or we're just building our things. How does academia contrast to this? What do you and your colleagues do inside of academia?
Yeah, within academia there are lots of different styles really. There's not one thing. Some people go full-on theoretical, mathematical, don't care about the real world at all, just want to work on things that are intellectually interesting. And that's fine. And some people are very much at the applied end of wanting to do research that is likely to have a real-world impact. I'm more on the applied end, and that's fine, too. But a common
distinction there is that academia can just think much longer term. So, you know, if you're doing a startup, you have to ship something within a few months. You can't afford to think 10 years into the future. No. Maybe you'll have sort of a long-term vision that you're gradually getting towards, but you do have to really ship things on a fairly short time scale. At a bigger company, maybe if you're working on infrastructure, you can think on a bit of a longer time scale because the requirements of what are needed is perhaps better understood. And in that case, you know, making sure that the system is scalable, operationally robust, and so on, it's then fairly clear what the requirements are and it's still a matter of implementing it, but in that case, you can think a bit longer term. But in academia, what I really appreciate is the freedom to work on things that are long-term and which are not immediately commercially viable or which are not aligned with the incentives of commercial companies.
So, one research area that I've been on for several years now is what we call local-first software, which is this idea that we want to take away a bit of the power from cloud operators and give it back to end users. So, end users should be more in control of their own data and less dependent on cloud services for providing the applications and the data that the users need.
And that's something that doesn't naturally come to companies, right? Because software as a service businesses, for example, the whole reason why they can charge a subscription is because they are able to essentially hold a gun to the customer's head and say, "Pay us your subscription, otherwise we will delete all your data." And I totally understand the commercial imperatives that lead to that, but it also leads to this situation where people have a gun against their head all of the time. That isn't really a healthy situation to be in, in my opinion. But changing that in such a way to take away that gun from customers' heads is difficult if you're in a business whose revenue depends on perpetuating that kind of lock-in situation. And there I feel like in academia, I have the freedom to work on things that go against this commercial incentive of companies and say, "Actually, no, I'm going to do what I think is right for the users." And then I'm going to say the commercial model of the companies making the software is second priority. And I can afford to do that because I'm not dependent on this commercial model.
To add to this, it's very interesting engineering problems, right?
Yes, and it's wonderful to get to work on interesting engineering and computer science problems while at the same time trying to pursue this higher level vision.
For local-first software, what are some of these really interesting engineering challenges that we will need to solve, or we need to solve, to get to more viable local-first software? May that be, say, note taking, it's a very popular one, right?
Yeah, so with our vision of local-first software, we're trying to get away from this dependency on centralized cloud services. There may still be cloud services involved in syncing data between your phone and your laptop, say, because often going via a cloud service is just the most convenient way of establishing that kind of communication. But we just don't want to have to trust a cloud service providing a particular function. And if you can get away from assuming this one cloud service, you could for example have multiple cloud services on multiple cloud providers side by side, and you just sync by whichever happens to respond first or sync with all of them, and then if one of them disappears, no problem because you've got the other one. And so it gives us a huge amount of freedom and flexibility if we get away from this assumption of centralized cloud services. But that introduces a whole bunch of interesting research and engineering challenges.
Because one thing that we've been working on lately, say, is access control. You know, simple problem: you have a document, you want to be able to grant collaborators access, and you want to be able to revoke that access again. Totally obvious, should be totally straightforward. In a centralized cloud service model, it is totally straightforward.
Yeah, you have the roles, you confirm those sort of things, and you check for the right roles, and that's it.
Yeah. But if you want to run your system over multiple providers, or even in a peer-to-peer setting, then, well, what could happen is that a user gets their edit permissions revoked, and concurrently, that user makes an edit to the document whose permissions have just changed. And now, some devices may see the edit to the document first, and the revocation second, and so they would accept the edit to the document. And another device may see it the other way around. They may see the revocation first, and then the edit to the document second, and they'll drop the edit to the document because they think it's not authorized. And now, those devices have become inconsistent with each other, permanently inconsistent. So, that means if we actually want to ensure consistency, even for this fairly basic setup, we now have to somehow figure out how to resolve the situation of an edit that is concurrent with the revocation of the user who made that edit. Solving that problem then means, in a decentralized setting, where we don't have just a single server that can make that decision. In a centralized setting, you know, you just have one server, it decides did the edit to the document come first, or did the revocation come first? And that one server makes that decision. But if you have multiple servers, they might make different decisions. So, then you could have a consensus protocol, but then consensus is messy because it requires some quorum votes, and requires nodes to be online, and so we've been trying to do the whole thing without doing consensus, while preserving high availability, while preserving the ability for users to work offline, preserving the ability to synchronize peer-to-peer without any servers, for example.
But that just makes the engineering challenge a lot harder. And it's solvable, and we are close to solving it for Automerge, which is the CRDT library that I work on. But it's just much less straightforward than it is in the centralized case. But that's a nice example of where interesting engineering challenges arise from this desire to get away from centralized services.
And then we were just talking about clocks earlier, but then the obvious thing that came to mind is, well, if all of them had the same clock exactly to the microsecond, you could just use a clock. You could use this timestamp. But as you said, in distributed systems, we cannot always trust that the clocks are always synchronized. So I assume a lot of the things that you have been researching and writing about are just coming back to
Absolutely. And in this particular setting of a user getting their edit permissions revoked, if a revoked user still wants to, say, vandalize a document, they can just backdate their edits, give it an earlier timestamp. So, relying on clocks is absolutely useless here because people can forge the timestamps from those clocks and thereby potentially undermine the access control mechanism. So, in this kind of system, we have to worry about potentially maliciously generated actions as well when the actions come from end user devices.
This is fascinating because it feels to me that you're solving a hard or maybe even harder engineering challenge than some startups would, because startups would go the easy route. They would take on a constraint, in this case a centralized server, which makes business sense, makes revenue sense. But because you are not doing this, you now need to look for a solution for a harder problem. And if you solve this harder problem, you can give a building block that can move the industry forward. Just give an option for either a business or an individual or an institution to, you know, have an option to not just use centralized, but use this decentralized local-first approach. And then of course they can weigh the trade-off and decide whichever makes sense.
Exactly. And that's what I mean with this long-term thinking. This is an example of it where, because it's research, we can afford to take this idealistic, principled stance and say, yes, we're going to solve this harder engineering problem because we think decentralization is a valuable feature. And we know perfectly well that most startups are not going to solve this problem because they will just do the easy pragmatic thing, which is the right thing for startups to do.
But we have a different set of incentives and we can afford to put in the time to try and solve those hard problems. And as you said, if we can solve them, then it creates more optionality for anyone, any users of this technology. They can, if they want to, choose to use this decentralized tech. And there's still trade-offs around it, but at least if they're not having to invent it from scratch, it'll be a lot easier to adopt this kind of decentralized tech for those who want to use it.
So, in academia, you're also teaching?
Mhm.
What courses do you teach?
At the moment, I have a concurrent and distributed systems course for the undergraduates and a cryptographic protocol engineering course for the master's students. And then additionally this year, I have a seminar course on security and I'm also teaching the undergraduate operating systems course. I've got quite a lot of teaching this year.
The distributed systems course, it's available on YouTube. Can you summarize what people who would go through this course, which again is freely available, thank you to you and the university for making it available, what would they learn throughout those courses?
Yes, so that distributed systems course, it's a bit more theoretical than what is in the book. So, it's more focused on algorithms and how we convince ourselves that the algorithms behave correctly under the assumptions of distributed systems that we talked about, of nodes may crash, communication might be unreliable, clocks might be wrong, et cetera. So, that's really it. It's not a very long course. It's just eight lectures' worth of material, but it goes into substantially more detail on the algorithms than the book. So, for example, one of the lectures goes through the entire Raft consensus algorithm, which is pretty complex.
But I really wanted to show the students exactly how it works because it's just such a nice illustration of the challenges of distributed systems and the various measures we need to take in order to handle the various types of edge cases and failures that can happen. And showing that the problem can be overcome. It's not easy and the algorithms are very subtle and it's very easy to have bugs in them, but it is possible to solve consensus in a way that works pretty well. And so that's really the message I'm trying to get across with this course.
And you mentioned that when you're writing the book together with Chris, you brought a lot of industry insight and being up-to-date, and you brought your experience of teaching and what works.
I don't think I have a particularly unique teaching style. In lectures I'll go through slides. I like to annotate the slides by hand during the lectures. I just draw on an iPad to make it a little bit more interactive, but other than that it is fairly theoretical. That's partly the way the Cambridge system works. It kind of favors theoretical and pen-and-paper courses over, say, implementation-focused practical courses. I think it would certainly be possible to do a practical course on this and I may incorporate a bit more practical exercise in the future, but right now it's mostly a theoretical pen-and-paper course and that is fine. The cryptography course that I do is much more hands-on. So, that's about actually getting the students to implement some elliptic curves from scratch, for example.
And how have you seen it in your time in academia, which is now a longer time period? How have you seen computer science education changing? How do you think it might change further in the future, especially as we're seeing AI be part of industry and probably the world as well?
Yeah, I mean prior to the AI explosion happening, actually the rate of change was very slow in computer science teaching. Partly that might be Cambridge, you know, Cambridge is over 800 years old. Everyone thinks on longer time scales. People don't tend to rush into the latest fads and instead try to focus on the fundamentals, and a lot of the fundamentals of computer science were developed in the 1930s already and are still true today, you know, lambda calculus and those types of things, for example. And so we have quite a bit of a focus on those sort of fundamentals rather than chasing the latest fashionable thing.
That said, AI has totally changed the way we can assess coursework, for example, because of course now we can try banning AI, but it's impossible to actually enforce such a ban and also it's kind of counterproductive because we do want students to engage with new technologies and figure out how to use them productively for themselves. But we want to somehow do that in a way that supports their own learning and doesn't undermine it. So how do we get the students to use AI in a responsible way, in a way that's mature? And we can't necessarily rely on the students being mature enough to know for themselves what is a helpful use of AI and what is a form of use of AI that undermines their own learning, because some of them are quite mature and able to decide that for themselves, but many are not, and so we need to provide some guardrails for them.
And we do need to make sure that when we have assessed work, for example, it's fair and it's perceived as fair by the students. And if the students feel that some of their co-students are getting really good marks without doing any work, that undermines the trust in the entire system. And so, we have to be very careful with how we approach this.
And to be honest, we don't really have good answers yet. So, we do now, for example, have a boot camp right at the start of the first year for the new students to expose them to basic software engineering skills, which is: this is version control, this is unit testing, this is generative AI. And the sort of basics that really everyone should be familiar with. And then the hope is that they will use that throughout their degree in order to just improve the work that they do. But how exactly we handle things for assessment, for example, we're still in the process of figuring out.
Though it sounds like the pace of change is going to be fast in the industry and also academia will probably adopt it, and we'll see, you know, what comes after.
Yes, there's a difference though, which is in the desired outcomes. I think with industry, generally, the desired outcome is a working product, for example. In academia, the actual artifacts that the students produce, like an essay that the students write, that's not really the point. We don't ask the students to write essays because we love reading their amazing essays. We ask them to write essays because we want them to go through a thought process, which helps them learn something. And it's that thought process and that learning which is really the desired outcome here. And so, that means that we do have to approach it a little differently, because generally in industry, you know, if you can use AI to get a job done faster and you get to an equivalent result, do it, absolutely. Because yes, that is the desired
outcome. Whereas in education, we do have to think about how we ensure that the learning outcomes and the thought processes are still preserved, such that the students benefit intellectually.
It's very relevant, especially Anthropic had a recent study where they looked at junior engineers. One group used AI, the other one did not. And they found, unsurprisingly, that — what would you also explain — the group who used AI, they had little to no learning, whereas the group that did not, they actually learned it.
Yes, I saw that study as well. I think the detailed methods of that study we might be able to quibble with a bit, but I think the general principle seems true that, yes, sometimes in order to learn something you just have to struggle with it a bit. Not struggle too much, so if people are stuck on some technicality and they can use AI to get unblocked and then be able to focus really on the main learning outcome, then I think it's good to use these types of tools, but if the point is to actually grapple with some difficult ideas and think them through in their own minds, then we need to still find ways to make sure the students are doing that.
You work both in industry and academia. What do you think industry could learn from academia and academia could learn from industry?
The two really could be closer together because often they regard each other with sort of disrespect really. The industry people will say, "Ah, that's theoretical. That's academic. It's got nothing to do with the real world." And they're really missing a trick there because actually there are a lot of interesting insights from research that are very relevant to the real world, but they're not necessarily making their way across that chasm. In the other direction, the academics will say, "Oh, this industry stuff, you know, that's just engineering. They're not actually doing any interesting thinking. It's just writing routine stuff." I see it as one of my goals to try and build better respect in both directions by bringing interesting insights from research into industrial practice, but also by informing our research by the problems that arise in the real world. And so that way, joining those two things up a bit better.
What are your current research topics that you're working on? Ones that you're excited about?
I have two main areas I'm working on at the moment. One is local-first software, so that's this idea that we want collaborative software like Google Docs, like Figma, etc., but in a way that gives better protection to users' data. That's less dependent on a single cloud provider who can lock you out of your files and that's therefore more resilient, gives users greater agency and greater autonomy over their own data. So that's an area that I've been working on for the last 10 years or so through a mixture of open source work and algorithm development and formal verification and so on.
I'm now also trying to set up a brand new research area in a totally different topic, which is on using cryptography to prove things about the physical world. So I'm interested there in especially sustainability-related things. So for example, if you want to verify that the carbon emissions involved in manufacturing a particular product were X and you want to be sure that that number is correct because maybe you want to include emissions as part of your purchasing decision and choose the product with the lower emissions. For that to be meaningful, then the emissions number has to be correct. And unfortunately, at the moment, the numbers are generally not correct because the incentives are to lie and cheat and to use creative accounting techniques, all as a way of greenwashing basically. Or a related thing is happening in the EU, for example, which is bringing in new regulations on preventing deforestation of tropical rainforests. So that, for example, coffee, cocoa, palm oil, etc. imported into the EU, the importer needs to prove exactly which plot of land it actually came from and then check against satellite imagery that that was not recently deforested.
And so I've been looking into using cryptography as a tool of proving things about the supply chains of these physical products, but without revealing commercially sensitive information. For example, a company will not want to reveal who its suppliers were and which ingredient to its process it purchased from which supplier, because that might reveal something about its secret recipe that it uses. And so, the hope here is that cryptography can allow us to prove that, for example, the accounting has been done correctly across supply chains, but without having to reveal publicly any of this sensitive data about suppliers or other customers.
What is your view from your vantage point on the impact that AI is having on academia, not just for students studying, beyond that, and also industry with your industry contacts?
Yeah, I mean, I'm not that deeply into the AI things, really. I'm seeing it more through my collaborators who are making very good use of AI tools for software development, especially. I personally write very little code these days, and so I haven't had that much need or occasion to actually use AI agents myself personally.
When writing prose, like working on the book, for example, I prefer to still do that the old-fashioned way of just write every word by hand. So, I haven't let AI anywhere near the text of the book, for example. And I don't know if that's the right decision. It's not really a principle thing that I think it would be wrong to do so. It's more that for myself, the process of writing is the way how I figure things out. And figuring things out is really my goal here. So, I'm trying to figure it out in my own head, and for that I just have to write it myself. There doesn't seem to be any way around it. But using AI is a way of getting feedback on ideas or exploring whether an idea really holds up to scrutiny or things like that. That seems like a very productive use of the technology, and that applies for both industry and academia, I would say.
So, as closing, for a student or a young professional who is still studying and considering their route into either industry or academia. What have you seen, who thrives in one or the other?
Yeah, my feeling is they're not really that mutually exclusive, or rather some of the best PhD students I've worked with, for example, actually have a few years of industry experience. So they might have done an undergraduate, maybe done a master's, then spent a few years in industry actually doing real software engineering, learning about the real world, and then maybe at some point got bored and thought, oh actually, you know, I want to work on maybe more idealistic things or have more freedom to choose their own research topics, and then start getting interested in doing a PhD. And that I find is quite a healthy route. You do get people who go, you know, straight from their undergraduate degree and master's into doing a PhD. But sometimes those people can just lack a bit of the breadth of perspective, and so I think having seen a bit of just real-world engineering is actually really helpful for people even if they then want to stay in research.
But in the opposite direction I think it can work very well too, because in research and academia we just get to think things through a lot more carefully than people often do in industry. Often people in industry, I feel like, sort of have short-circuit reasoning. Maybe don't quite reason something through from first principles, but just like, oh, I heard this from a conference talk, I'm just going to go with that. But what academia can teach is the sort of nuanced and critical thinking to really reason through trade-offs, for example, and to really justify why something is true. And so I think it's really good actually if people can weave in and out of industry and academia a bit and not regard it as two totally mutually exclusive career paths, but actually have a bit of switching between the two.
Well, Martin, thank you very much. I expected us to talk a lot more about your book, which we did, but I have a newfound curiosity and respect for all the important and interesting academic work that you and everyone else is doing. So, thank you so much for this.
Thank you for the great interview. This was really interesting.
I hope you enjoyed this rare conversation with Martin Kleppmann. I found it interesting to learn that the first edition of the book assumed that you have machines with local disks, but actually today this is not how most engineers build systems anymore. Cloud-native primitives like S3 change how you build systems, and this is why this book just needed a refresh.
I also appreciated Martin's take on whether engineers still need to understand how systems are journals when they're using managed services. If you're building business logic on top of these services, you probably don't need to know every detail, but it can become useful to be able to look deeper, especially when you need to debug your system.
By the end of our conversation, I gained a lot of appreciation for the academic research that Martin is doing. The local-first software work, the access control problem in decentralized systems, using cryptography to verify supply chain emissions. A lot of these are hard engineering problems that few startups will take on. It was nice to understand how academia is in a good position to do work that has a long-term focus.
Do check out the show notes below for related Pragmatic Engineer deep dives. If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks, and see you in the next one.
Article published
