Martin Kleppmann on Designing Data-Intensive Applications, the Cloud, and Research After Industry

Open on YouTube ↗
Overview

Martin Kleppmann wrote Designing Data-Intensive Applications (DDIA), a standard reference for engineers who build large back-end systems. Nine years after the first edition, a heavily updated second edition is out. In this conversation on The Pragmatic Engineer, Kleppmann covers the path from two startups to LinkedIn's data infrastructure team, then to writing and to academia. They explain what changed in the second edition, why cloud object stores changed how data systems are built, and why they think formal verification may matter more once AI writes much of our code. They also describe their current research on local-first software and on using cryptography to verify claims about physical supply chains. Their view throughout is that engineering is about making trade-offs visible so that people can choose deliberately, not about finding a single right answer.

38 min read

A first startup that worked technically but found no business

Kleppmann studied computer science as an undergraduate and then was unsure what to do next. Starting a startup "seemed like an interesting thing to try," so they started one without knowing what it would build, and spent the first stretch searching for a promising idea.

That company was Go Test It, around 2008. Cross-browser compatibility was a serious pain at the time. Internet Explorer was still widely used, Chrome had just come out, and browsers were incompatible with each other. Go Test It was a hosted, automated cross-browser testing service built on Selenium, the open-source project that still exists. Customers wrote scripts that simulated a user clicking through a site and checked that the right behavior happened, and the service ran them so customers didn't have to maintain VMs with various operating systems themselves.

Kleppmann says the product worked technically, but adoption was the problem. Website builders agreed in principle that cross-browser testing mattered. In practice it was hard to get them to fit it into their workflow, get in the habit of using it, and invest in writing test scripts. Kleppmann notes that one or two companies from the same era did make a business of it, with Sauce Labs succeeding. Even for them, Kleppmann thinks, it was a slow-growing and difficult business.

The company was based in the UK and mostly bootstrapped. Kleppmann did consulting work to pay for hiring friends cheaply to help build the product, with only a very small amount of angel money. The venture didn't go far, but through it Kleppmann met the people who became co-founders of the second startup.

Rapportive, Y Combinator, and the move to San Francisco

The second company, Rapportive, went much better. It put social media inside Gmail. A browser extension modified the Gmail web interface so that when an email arrived from an unfamiliar sender, a sidebar showed a summary profile: picture, job title from LinkedIn, recent tweets, perhaps recent Facebook posts, and whatever else could be found. It started around 2010, became popular quickly, and on that basis the team raised money from Y Combinator, which Kleppmann describes as already well regarded but still fairly small.

The founders flew from the UK for the roughly three-month program, then obtained US work visas and settled in San Francisco. Kleppmann remembers this as exciting, like "going to the center of where it was all happening." They arrived knowing only one or two people in the Bay Area, who introduced them to others, who introduced them to still more. Kleppmann valued how open the area was to outsiders who could show up with an early-stage idea, raise money, and become established.

The technology was, in Kleppmann's words, "fairly unexciting": a Rails app on Postgres with some Redis. The one technically interesting part was that they had essentially built a graph database on top of Postgres.

An acquisition under pressure, and LinkedIn Intro

Rapportive was sold to LinkedIn in 2012, when the team was five people. Kleppmann calls it a success for everyone, though not a large amount of money, and is candid that it was not a triumphant exit. They had tried revenue-generating options without success, user growth was acceptable but not enough for a large round, and they were running out of money. Their visas also constrained them: they couldn't cut their own salaries, because doing so would have violated the visa conditions. Selling became, as Kleppmann puts it, "the least bad option." Given how little leverage they had, Kleppmann is pleased with the result, and says the process had the usual twists and moments when it seemed the deal would collapse.

Kleppmann speaks warmly of LinkedIn. The team was allowed to operate as an essentially independent unit and to keep building what it wanted. The Rapportive extension was put "on life support" while the team worked on a new product, eventually released as LinkedIn Intro. It got what Kleppmann calls "a slightly weird reception" and was shut down soon after launch. Kleppmann says there is a longer story there but doesn't tell it. After the shutdown the team was disbanded. Kleppmann still credits LinkedIn for giving them the freedom to build and launch the product even though it failed.

Kafka, Samza, and seeing how data systems fit together

After the team broke up, Kleppmann joined LinkedIn's stream processing team. Kafka had recently been developed there and had just been open-sourced, and Kleppmann worked on Samza, a stream processing framework built on Kafka.

Asked why LinkedIn built something that now looks so generic, Kleppmann points to Jay Kreps's blog post from that era, "The Log," which explains why Kafka was an append-only log and not a traditional message queue. In Kleppmann's account the motivation was data integration. Many databases and event-producing systems, such as user activity events, generated stream-shaped data. Many downstream systems needed that data, including the data warehouse and the Hadoop cluster used for machine learning. The practical problem was how to physically move data from one system to another. Kreps designed Kafka as that integration point: something close to a lowest common denominator, but still a general-purpose abstraction connecting many sources to many sinks.

This was Kleppmann's first hands-on experience with a really large system. The biggest company they had worked at before was Rapportive, with its single-instance database. At LinkedIn they got to hand-code MapReduce jobs in Java against a large Hadoop cluster, which they enjoyed. The real revelation came from the stream processing ideas and from Kreps's evangelism for Kafka. Kleppmann says they began to understand how the various data systems fit together, what they have in common, and what the underlying principles are. That understanding went directly into the book.

Leaving LinkedIn to write full-time

Kleppmann first moved back to the UK and kept working for LinkedIn remotely. Their girlfriend at the time, now their wife, was still in the UK, and Kleppmann didn't feel at home in the Bay Area, so they didn't push for her to move there. They love the Bay Area as a place to visit and still have many friends there, but wouldn't want to live there.

LinkedIn gave Kleppmann 50% of their time to work on the book alongside engineering duties. Kleppmann stresses that LinkedIn didn't have to do this and got little from it beyond a book it could use for internal training. Even so, the arrangement didn't work. Writing in parallel with an engineering job that included on-call duty meant too much context switching. Urgent on-call issues easily took over, leaving no room for the freedom needed to write something new. Kleppmann eventually left LinkedIn for an unpaid sabbatical ("i.e. unemployment") to write full-time. Only after that did they consider academia.

What the book was meant to be, and how it was researched

Kleppmann says the final book differs from the initial plan, but the goal stayed the same. It would be a broad conceptual overview that compares trade-offs across many kinds of tools, not a guide to any single system. It would be practitioner-focused, not a theoretical textbook: something people could use to build real systems. It was the book Kleppmann wished they'd had at Rapportive, where the team hit database performance problems and was "searching around in the dark" because it lacked the foundations to diagnose what was happening. With more background on how data systems work internally, Kleppmann believes, they would have had the intuition to debug those issues. Once they had learned it, they wanted to write it down so others wouldn't have to learn it the hard way.

Much of the learning came from curiosity and conversation. LinkedIn had senior data systems engineers who understood this material well but hadn't necessarily written it down. Kleppmann questioned them and built a mental model from those conversations. With those basics, research papers became approachable. Papers go much deeper into how and why systems are designed as they are, but they take a long time to read, so Kleppmann tried to extract the essential ideas. They also read many blog posts. The long reference lists at the end of each chapter are the material Kleppmann actually used, included as further reading for anyone who wants to go beyond the basics.

Kleppmann says the book's three-part structure (foundations of data systems, distributed data, derived data) was imposed mostly after the fact. The chapter topics, such as transactions, replication, sharding or partitioning, and consistency and consensus, were clear from the original proposal to the publisher. Kleppmann wrote one chapter at a time and started each with extensive background research. The internal structure of each chapter emerged then. For replication, for example, they concluded only after that research that the three main approaches were single-leader, multi-leader, and leaderless, and organized the chapter around them.

Estimates, deadlines, and the two editions

Asked whether estimating a book resembles estimating a software project, Kleppmann agrees that both take "vastly longer than expected." The first edition took about four years of elapsed time, or perhaps two and a half years of full-time-equivalent work, and missed the publisher's deadline by roughly two and a half years. O'Reilly was relaxed about it and let Kleppmann take the time to make it good. For the second edition O'Reilly was "a bit more aggressive and pushy" about deadlines. Kleppmann understands why, since by then the book was established and readers were waiting, but they missed the freedom they had the first time.

Reliable, scalable, maintainable, and scaling down

The subtitle of both editions refers to reliable, scalable, and maintainable systems. Kleppmann acknowledges these terms have no formal definitions. For them, reliability mainly means fault tolerance: the system should keep working overall when a network link fails or a node crashes. Much of the book covers techniques that support this, such as replication.

Kleppmann says scalability is "thrown around a lot" because it is fashionable and suggests success and millions of users. The book takes a more detached view: scalability is about the mechanisms for handling changes in load, such as adding computing capacity when load increases, using techniques like sharding. The book focuses on horizontal scaling because buying a bigger machine is less interesting to write about. What has become interesting about modern cloud and back-end services, in Kleppmann's view, is shared-nothing architecture: handling very high load using relatively cheap commodity machines.

Kleppmann says they have recently thought more about something they didn't consider much at first: scaling down. Ideally cost and computing capacity should be roughly proportional to load, which means a very lightly loaded service should be extremely cheap to run. That isn't guaranteed. With on-premises infrastructure a physical machine is the unit of deployment, and even if you split it into two dozen VMs, you still have to allocate resources to each. Kleppmann finds some serverless systems interesting because a service handling three requests a day is perfectly acceptable to them.

The second edition: bringing in Chris Riccomini and building on object stores

Kleppmann had known for several years that a second edition was needed because the first was becoming dated. By then they had an academic job centered on research and teaching, which they enjoy and didn't want to give up, so updating the book was a side project. The same context-switching problem returned, and progress was slow. Kleppmann also realized that, after moving into theory, they had lost touch with current industry practice, such as how people were using data lakes.

The solution was Chris Riccomini, a former LinkedIn colleague from the stream processing work and author of The Missing README. Kleppmann had read that book and considered Riccomini a strong writer. Riccomini also wrote a newsletter, Materialized View, on current trends in data systems, and had become a startup investor in that area. Kleppmann describes a good division of labor: Riccomini knew the current state of industry, and Kleppmann had strong views on teaching, meaning precise, carefully chosen wording that still reads easily.

The main change they planned from the start was cloud-native architecture, meaning data systems built on cloud services as the foundational abstraction. The first edition assumed machines with local disks: a database instance writes to its local disk, and the database software replicates data to another machine, which writes to its own disk. That was how computing worked for a long time. Now databases are being built on object stores, and replication happens at the object-store level instead of, or in addition to, the database level. Kleppmann distinguishes this from building on virtual block devices such as EBS. Those are cloud services too, but they still present the single-node abstraction of a block device with a file system on top. An object store is a new abstraction that looks and behaves differently from a file system. Some systems were starting to use it at the time of the first edition, and since then it has taken off. Kleppmann says the second edition weaves the idea through the whole narrative instead of confining it to one section.

Do engineers still need to understand the layer beneath?

The host asks whether managed services, which handle replication and come with uptime SLAs, remove the incentive for engineers to understand what lies underneath. Kleppmann sees this as the familiar history of computing: new, higher-level abstractions. Relying on a higher-level abstraction does mean not thinking about lower-level details. Using a garbage-collected language means not thinking about memory allocation. Kleppmann thinks that's a loss only in some contexts. People building low-level systems still need to care about memory. People writing higher-level business logic, Kleppmann thinks, are fine not caring. Data systems are similar: if you build higher-level systems that don't need to care about infrastructure, use the abstractions. Someone still has to build the cloud services from lower-level components, and those people will specialize more deeply in how to engineer, operate, and make them reliable, treating the higher-level builders as their customers.

Kleppmann adds that the book's philosophy is to give people a sense of how systems work internally, so that when something behaves strangely, such as unexpected performance, they have some intuition about why. The storage engine chapter explains B-trees and log-structured (LSM) storage engines. It isn't for people building their own databases, who need far more depth. It is for application developers who can use a storage engine better and diagnose problems if they know a little about how it works. The second edition applies the same philosophy to cloud services. Kleppmann's example is row-oriented versus column-oriented storage for analytics: a technical distinction that takes some background reading to understand but has large performance consequences. In cases like that, knowing the internals is "actually like a superpower."

Availability, cost, and a European case for multi-cloud

The discussion turns to how an engineer decides between multi-zone, multi-region, or multi-cloud setups. Kleppmann frames it as how much availability risk you'll accept versus the overheads: the computational overhead of the system, the human overhead of designing and operating it, and the cost. A more fault-tolerant system costs more to design and run. A simpler one may go down more often but is cheaper. Kleppmann says there's no right or wrong answer, and everyone has to decide where they sit on that spectrum. Multi-region pushes toward higher availability because it tolerates losing an entire region, but it affects which consistency models are possible across regions. The book tries to make these trade-offs explicit.

On multi-cloud, Kleppmann says a new concern has come up "just in the last month really": Europe's dependence on US cloud services. If geopolitics went badly and Europe were suddenly locked out of US cloud providers, the effect would be severe. Kleppmann hopes this won't happen and still considers it fairly unlikely, but "no longer unthinkable." From a European perspective they have been thinking about how to make systems resilient against that, which is a business risk as well as a kind of outage. A multi-cloud setup would let systems keep running with another provider if one company locked you out. It sits at the expensive, risk-reducing end of the spectrum, but Kleppmann thinks it is worth serious consideration for critical workloads where the geopolitical risk is significant.

Both agree that understanding and communicating risks and trade-offs will be a core part of engineering work. Kleppmann suggests that as AI writes more code, the job may depend less on expressing logic in a particular language and more on these high-level trade-offs.

How the cloud changed scaling, and why sharding matters less

Asked whether cloud primitives make scalability easier to reason about, Kleppmann separates two ends. At the high end, reaching very large scale is still hard. Object storage provides elastic capacity and removes disk capacity planning, but sharding affects application code and can't be made fully transparent. If one machine can't handle your workload, you still need substantial engineering thought, cloud or not. Where the cloud has clearly helped, Kleppmann says, is the low end: serverless systems that spin instances up and down quickly enable very lightweight services. Without cloud services this would be much harder, because memory and CPU would have to be allocated statically to a VM. The host mentions a small serverless site that costs about 13 cents a month, and Kleppmann calls this more efficient use of computing resources.

The host recalls that sharding was a major topic at Uber, including in interviews, and that it now seems less common for engineers to implement it themselves. Kleppmann attributes this less to the cloud than to more powerful hardware. A big machine can do a lot, so more workloads fit on one machine and still reach significant scale. Parallelism still matters, since you have to use hundreds of cores efficiently, and sharding is one way to get it. But sharding across machines is less pressing for many workloads. It hasn't disappeared, because some workloads still need it. Replication stays relevant even at small scale because it serves fault tolerance, not scalability.

"The Trouble with Distributed Systems"

The chapter with that title, Kleppmann explains, defends the assumptions of distributed systems theory by showing that they reflect reality. Theory usually assumes no upper bound on message delay. A message might arrive in 100 microseconds or in 10 years. Some theory does assume timing bounds, and Kleppmann calls that dangerous because network delays sometimes become much larger than usual. Theory also says nodes can crash, but in practice "crash" covers a software crash, a hardware failure, someone pulling the power cable, or a node that is still running but disconnected from the network. Clocks are another example: they're usually roughly correct but not precise enough to rely on. The lesson Kleppmann draws is that it's tempting to assume things behave well, and reliable systems require giving up those assumptions: "don't believe anyone who says oh failures are rare."

Kleppmann enjoyed writing the chapter. It is largely a collection of things that went wrong, built from postmortems published by tech companies and their root causes. A favorite example is sharks biting undersea cables. Kleppmann has heard that cable shielding has improved so sharks no longer cause damage, and that now cows on land step on cables and occasionally cause network outages.

The host notes that for a team like S3, the chapter describes daily life, since at that scale disk failures and even data center fires happen regularly. At a small company the same event may happen once a decade and be a big deal. Kleppmann agrees there is no universal answer. It's a trade-off between risk and cost, which makes it a business decision. The chapter aims to give people what they need to decide in an informed way, and Kleppmann doesn't want to make that decision for them.

What left the book and what came in

Some material could be cut. The first edition covered MapReduce in detail, but Kleppmann says bluntly: "MapReduce is dead. Nobody uses it anymore." Its successors, such as Spark and Flink, are in use. The second edition still mentions MapReduce, but as a teaching tool for understanding how partitioned batch processing systems work.

Coverage grew in areas that support AI. The book isn't about AI, but AI applications raise data systems concerns. Vector indexes are the clearest case. They were added to the storage engine chapter, which fit well because the chapter already compared indexing strategies and vector indexes are another one. Data frames were also added. Kleppmann notes they aren't exclusively an AI concept, but they are a good representation for training data and have become an important data model alongside relational, graph, and JSON documents. The first edition didn't cover them. Kleppmann describes these as expansions that reflect what people build without changing the book's direction.

Ethics becomes its own chapter

In the first edition, "Doing the Right Thing" was a subsection near the end. In the second edition it is the final chapter. The host quotes it: "We, the engineers building these systems, have a responsibility to carefully consider those consequences and consciously decide what kind of world we want to live in."

Kleppmann says the first version came from a sense that ethics had been largely ignored during their time in industry, especially in startups focused on building products customers would love. Consumer products were often designed around harvesting behavioral data because that data could be monetized through advertising, with little reflection on what was good or bad about it. Kleppmann didn't want to prescribe an approach, but wanted to point out that data protection laws now affect how data systems are designed, and that there is an ethical responsibility. If people go into tech to change the world, thinking about the effects of their technology is part of the job, and engineers are prone to neglecting it. Kleppmann says the section was also a way of working through these questions personally, since they hadn't thought much about ethics when they started building these systems.

Asked whether engineers are well placed to influence outcomes, perhaps more directly than regulators acting years later, Kleppmann agrees they have a strong voice. Engineers should present trade-offs so business leaders can decide well, and presenting trade-offs includes naming risks. Those risks go beyond technical ones like data corruption to societal harms, unintended consequences, and reputational damage to the company if a technology turns out to be harmful. Kleppmann wants these decisions made deliberately, "not just sweep it under the carpet."

Formal verification in an AI-assisted world

The host asks about a post Kleppmann wrote in December arguing that formal verification may become more important with AI. Kleppmann first describes the range of formal methods. At the introductory level, you write a high-level specification of a system's expected behavior in a language such as FizzBee or TLA+, and use a model checker, which Kleppmann describes as essentially a randomized test-case generator, to run through many scenarios and check that the desired properties hold. At the advanced level is formal proof: you write a mathematical specification and prove that an algorithm or implementation always satisfies it. Tests check a few example inputs. A proof can cover potentially infinite state spaces and show that a safety property holds in every possible case.

Kleppmann never used formal verification in industry because it took too much time. They took it up in academia, where they could spend months proving an algorithm correct. They find it valuable for subtle algorithms whose correctness is hard to judge from the code, especially high-stakes ones where a bug would corrupt data or open a security hole. They have written proofs with the Isabelle proof assistant and mention Rocq and Lean as alternatives. These proofs are very hard to write: learning the language takes a long time, and even afterwards writing the individual proof steps is laborious.

To show what proving involves, Kleppmann uses a simple example: proving that concatenating two lists yields a list whose length is the sum of the two lengths. You would use induction over one list. Concatenating a list of length i with an empty list gives length i. Appending a list of length one gives i + 1. Continuing by induction shows the result is i + j for every possible i and j. A unit test would check a few cases, such as lengths zero, one, and five. For something this simple you can convince yourself by reading the code, but for complex algorithms, Kleppmann says, our brains can't grasp them well enough to be sure without a proof.

For engineers who want to start, Kleppmann recommends model checking with TLA+ or FizzBee, which are much friendlier than Isabelle, Rocq, or Lean. The proof assistants require much more background knowledge, and Kleppmann admits the learning resources for formal proof aren't very good and they haven't found great books. They learned by pairing with lab colleagues who had years of experience, describing what they wanted to prove and being shown how to break it down step by step.

Kleppmann gives several reasons for thinking formal verification could become more important. LLMs are getting increasingly good at writing proofs, and if humans don't have to write them by hand, proofs become economical where they weren't before. LLMs also increase the need for proofs. With so much code being vibe-coded, manual human review would become the bottleneck, which would cancel much of AI's benefit, so automated ways to check correctness are needed. Extensive testing is a good start, but only proof covers every possible case. That matters most in security, where "it just takes one little bug" to compromise a whole system. For domains that need a complete absence of bugs, Kleppmann hopes LLMs will make formal verification accessible to people who previously found it too hard and too expensive.

Academia's long time horizons and local-first software

Kleppmann says academia covers a wide range, from purely theoretical work unconcerned with the real world to applied research aimed at real impact, and they place themselves at the applied end. The common difference from industry is time horizon. A startup has to ship within months and can't plan ten years ahead. A larger company working on infrastructure can think somewhat longer term because the requirements are better understood. Academia allows work that is long-term, not immediately commercial, or even contrary to commercial incentives.

Local-first software, which Kleppmann has worked on for years, is their example. The goal is to shift power from cloud operators back to end users, so users control their data and depend less on cloud services for their applications. Kleppmann argues this doesn't come naturally to companies. SaaS businesses can charge subscriptions because they can effectively "hold a gun to the customer's head" and say "pay us your subscription, otherwise we will delete all your data." Kleppmann says they understand the commercial reasons for this but consider it an unhealthy situation, and one that is hard to change from inside a business whose revenue depends on lock-in. In academia they can prioritize what they believe is right for users and treat the business model as secondary, because they don't depend on it.

In their vision, cloud services may still help sync data between, say, a phone and a laptop, because routing through a server is often the most convenient way to connect devices. The difference is that no single provider must be trusted. Data could sync through several providers at once, whichever responds first or all of them, and if one disappears, the others remain. That flexibility creates new research and engineering problems.

A hard problem: revoking access without a central server

Kleppmann's current example is access control. Granting and revoking a collaborator's access to a document is trivial with a central server, which checks roles. Across multiple providers or peer-to-peer, it gets hard. Suppose a user's edit permission is revoked while that user concurrently edits the document. Some devices see the edit first and the revocation second, and accept the edit. Others see the revocation first and reject the edit as unauthorized. The devices are now permanently inconsistent. A central server would simply decide which came first. Multiple servers might decide differently. A consensus protocol could resolve this, but Kleppmann calls consensus messy because it needs quorum votes and nodes to be online. The team is trying to solve it without consensus while keeping high availability, offline work, and server-free peer-to-peer sync. Kleppmann says they are close to solving it for Automerge, the CRDT library they work on, but it is much harder than the centralized version.

The host suggests synchronized clocks and timestamps would help. Kleppmann says clocks are useless here: a revoked user who wants to vandalize a document can backdate edits with earlier timestamps. Because actions come from end-user devices, the system has to handle potentially malicious actions.

The host observes that this is arguably a harder problem than most startups take on, since they would accept the constraint of a central server that makes business sense. Kleppmann agrees and calls that the right choice for startups. Research, with different incentives, can take the "idealistic, principled stance" that decentralization is worth the harder problem. If it's solved, others gain options: they can adopt decentralized technology without inventing it, while still weighing its trade-offs.

Teaching, and computer science education after AI

Kleppmann currently teaches an undergraduate course on concurrent and distributed systems, a master's course on cryptographic protocol engineering, a security seminar, and the undergraduate operating systems course, which they describe as a heavy teaching year. The distributed systems lectures are freely available on YouTube. They are more theoretical than the book, focusing on algorithms and on reasoning about their correctness when nodes crash, communication is unreliable, and clocks are wrong. The course is only eight lectures but goes deeper on algorithms than the book. One lecture covers the full Raft consensus algorithm, which Kleppmann calls complex but a good illustration of distributed systems' challenges and of how edge cases and failures can be handled. The message they want to convey is that consensus is subtle and easy to get wrong, but can be solved in a way that works well. Kleppmann describes their teaching style as unremarkable: slides annotated by hand on an iPad during lectures. They note that Cambridge favors theoretical, pen-and-paper courses over practical implementation, and they may add more practical exercises later. The cryptography course is already hands-on, with students implementing elliptic curves from scratch.

Before the AI boom, Kleppmann says, computer science teaching changed slowly, partly because Cambridge, being over 800 years old, thinks on long timescales and emphasizes fundamentals, many from the 1930s such as lambda calculus, instead of chasing fashions. AI has changed assessment entirely. Banning it can't be enforced and would be counterproductive, since students should learn to use new technology productively. The challenge is helping students use it in ways that support their learning instead of undermining it. Some students are mature enough to judge this themselves and many aren't, so guardrails are needed. Assessment also has to be fair and seen as fair: if students believe classmates earn high marks without effort, trust in the system erodes. Kleppmann says frankly that they don't have good answers yet. One step is a boot camp at the start of first year covering basic software engineering skills: version control, unit testing, and generative AI. How to handle assessment is still being worked out.

Kleppmann emphasizes that education's goals differ from industry's. In industry the desired outcome is usually a working product, so if AI gets you there faster with an equivalent result, use it. In education the essay itself isn't the point. "We don't ask the students to write essays because we love reading their amazing essays." The goal is the thinking process and the learning it produces. The host mentions a recent Anthropic study of junior engineers in which the group using AI learned little and the group without AI learned more. Kleppmann says one could "quibble" with the detailed methods, but the general principle seems right: learning sometimes requires struggle, though not too much. Using AI to get past a technicality and focus on the main learning goal is good. Where the goal is to wrestle with difficult ideas, students still need to do that themselves.

Bridging industry and academia

Kleppmann thinks the two worlds often view each other with disrespect. Industry people dismiss research as "theoretical" and irrelevant and miss useful insights. Academics dismiss industry work as "just engineering" without interesting thinking. Kleppmann sees one of their own goals as building respect in both directions: bringing research insights into practice, and letting real-world problems inform research.

Current research: cryptography for verifying claims about the physical world

Kleppmann has two main research areas. One is local-first software, pursued for about ten years through open source work, algorithm development, and formal verification. The aim is collaborative software like Google Docs or Figma that better protects users' data, depends less on a single provider that can lock you out, and gives users more agency and autonomy.

The second is a new area they are trying to establish: using cryptography to prove things about the physical world, with a focus on sustainability. One example is product carbon emissions. If buyers want to choose products with lower emissions, the numbers have to be accurate, and Kleppmann says they currently generally aren't, because incentives encourage lying, cheating, and creative accounting that amounts to greenwashing. Another example is new EU regulation on deforestation. Importers of goods such as coffee, cocoa, and palm oil must prove which plot of land a product came from and check satellite imagery to confirm it wasn't recently deforested. The difficulty is that companies won't disclose their suppliers or which ingredients they bought from whom, because that could reveal a secret recipe. Kleppmann hopes cryptography can show that accounting across a supply chain was done correctly without publicly revealing sensitive supplier or customer data.

Writing without AI, and moving between industry and academia

Kleppmann says they are not deeply involved with AI tools personally. They see them mostly through collaborators who use them well for software development, and they write very little code these days. They wrote their parts of the book entirely by hand and kept AI away from the text. They say this isn't a matter of principle and they aren't sure it's the right decision. For them, writing is how they figure things out, and figuring things out is the goal, so they have to do it themselves. They do see value in using AI to get feedback on ideas or test whether an idea holds up, in both industry and academia.

Asked for advice to students choosing between industry and academia, Kleppmann says the paths needn't be exclusive. Some of the best PhD students they've worked with spent a few years in industry after an undergraduate degree or master's, did real software engineering, then got bored and wanted more idealistic work or more freedom to choose research topics. Students who go straight from their degree into a PhD sometimes lack breadth of perspective. Movement in the other direction helps too. Kleppmann feels industry reasoning is often "short-circuit" reasoning, like adopting an idea because it appeared in a conference talk, whereas academia teaches nuanced, critical thinking about trade-offs and justifying why something is true. Kleppmann's conclusion is that people benefit from moving back and forth between industry and academia instead of treating them as separate career paths.