Martin Fowler on AI and Software Engineering: From Determinism to Non-Determinism

Open on YouTube ↗
Overview

Martin Fowler, Chief Scientist at Thoughtworks, co-author of the Agile Manifesto, and author of Refactoring and Patterns of Enterprise Application Architecture, sat down in person with the host of The Pragmatic Engineer to discuss how large language models are changing software development. His central claim is that AI is the biggest change of his career, comparable to the move from assembly language to the first high-level languages. What makes it different, he argues, is less the rise in abstraction than the move from deterministic to non-deterministic tools. The conversation also covers his career, the Thoughtworks Technology Radar, vibe coding, testing, refactoring, design patterns, the Agile Manifesto, and the state of the industry.

31 min read

An Accidental Route into Software

Fowler describes his entry into software in the late 1970s and early 1980s as largely accidental. At school he got poor marks for writing but did well in mathematics and physics. He also says he is "hopeless with my hands," unable to do things like remove rusted nuts from his car. That pushed him toward electronics, which needed little more than a soldering iron, and then to computers, which didn't need even that. Before university he spent a year at the UK Atomic Energy Authority writing Fortran IV. After a degree mixing electronic engineering and computer science, he chose computing over traditional engineering jobs, which he saw as poorly paid and low in status.

His first job was at the consulting firm Coopers & Lybrand. He wasn't hired as a strategist. He looked after Unix workstations because he was one of the few people who knew Unix. The small group he worked with had adopted ideas from a consultant who packaged them under the label "object orientation." Looking back, Fowler says there was "a lot of snake oil involved." He adds that this is a little unfair, because some of the ideas were very good, and the work got him into object-oriented thinking in the mid-1980s.

There he met Jim Odell, an American independent consultant and teacher who became his biggest early mentor. Fowler then spent a couple of years at a small company whose UK office, with four people, was its largest. After that he went independent, helped by work that Odell passed to him. He remembers thinking he would never work for a company again. He wrote his first books as an independent consultant, moved to the United States in 1993, and worked with Kent Beck on Chrysler's C3 project, which he calls the birth project of Extreme Programming.

Joining Thoughtworks and the "Chief Scientist" Title

Thoughtworks began as a client. It was running a project of about 100 people that Fowler says was "clearly going to crash and burn." He helped them see what was happening and recover, and they invited him to join. His reason for accepting was that other clients would say good ideas were hard to implement and stop there, while Thoughtworks would say they were hard to implement and try anyway, and usually pulled it off. He expected to stay a couple of years. That was 25 years ago.

He jokes that as Chief Scientist he is "chief of nobody" and does no science. The title was popular at the time for public-facing thinkers; he recalls Grady Booch holding it at Rational. The irony, he says, is that Thoughtworks employees could then choose their own titles, but he was not allowed to. Thoughtworks rejected his preferred alternatives, such as "flagpole," "battering ram," or his favorite, "loudmouth."

How the Technology Radar Is Made

The host lists items from the newest Radar, which came out on the day of recording. In "Adopt" are pre-commit hooks, ClickHouse for analytics, and vLLM for running LLMs efficiently. In "Trial" are Claude Code and FastMCP. Many AI-related items sit in "Assess." The host asks how Thoughtworks stays so close to the industry.

Fowler traces the Radar to a technology advisory board set up just over ten years ago by then-CTO Rebecca Parsons, who wanted practitioners to keep her informed about projects. Fowler was on the board mainly because he was a public face of the company. At one meeting, her technical assistant suggested building a picture of which technologies projects were using and how useful they were, because the company struggled to spread good ideas internally. That struggle existed when it had a few thousand people and is harder now that it has about 10,000. The assistant came up with the radar metaphor and its rings. Thoughtworks habitually publishes what it builds for internal use, so the Radar went public.

The process has changed as the company grew by an order of magnitude. People now nominate "blips," which are entries on the radar. They brief someone connected to them by geography, line of business, or technology, who in turn briefs what is now called the Doppler group ("we can be a bit loose with our metaphors"). Blip-gathering sessions happen a month or two before the meeting, and in the meeting the group goes through candidates one by one. Fowler calls it a bottom-up exercise. He admits he is now so far from day-to-day work that he doesn't recognize most of the technologies, though he sometimes picks up themes. Microservices, about ten years ago, was one such theme; it came through the Radar and led to his writing on the subject with James Lewis.

Both men note that Thoughtworks encourages clients to build their own radars. Fowler says client radars differ: they can be more prescriptive about what teams should do, and they can simply decline to engage with a technology. Thoughtworks can't do that, because if its clients use something, it has to learn it.

The Biggest Shift Since High-Level Languages

Asked what earlier change compares to AI, Fowler says it is the biggest of his career. Across the whole history of software, he thinks the closest parallel is the move from assembly to early high-level languages like COBOL and Fortran, which came before his time. He did a little assembly at university, which was useful mainly because it made him never want to do it again.

He describes what that transition meant. Assembly instructions differed from chip to chip, and even simple tasks required convoluted sequences of moves between memory and registers. Even Fortran IV, which he calls relatively poor, let you write conditionals and loops. It had no "else," allowed only one statement per "if," and needed gotos for blocks, but it moved programmers away from the hardware toward something more abstract. It also decoupled code, to some degree, from whether it ran on a mainframe or a minicomputer.

He thinks LLMs bring a similar degree of mind shift, but the character is different. There is some increase in abstraction. The larger part, he says, is "the shift from determinism to non-determinism," where you suddenly work in an environment that is non-deterministic, "which completely changes" how you have to think.

Abstraction, Rigorous Language, and Engineering Tolerances

The host asks whether English is simply the next abstraction layer above high-level languages. Fowler agrees there is some jump in abstraction but says it is smaller than the determinism jump. He emphasizes a feature of high-level languages he didn't mention earlier: the ability to create your own abstractions within them. Fortran allowed some of this through subroutines, and object-oriented and functional languages like Lisp allowed much more. He cites an old Lisp adage that you should create your own language in Lisp and then solve your problem in it. Balancing problem-solving with language-building, he says, is what leads to maintainable, flexible code.

AI helps a little here, he suggests, because abstractions can be built more fluidly. The catch is that their implementations are now non-deterministic, and "we've got to learn a whole new set of balancing tricks." He points to his colleague Unmesh Joshi, who uses the LLM to co-build an abstraction and then uses that abstraction to talk to the LLM more effectively.

He recalls reading, in a book he couldn't name during the interview, that if you describe chess games to an LLM in plain English it doesn't learn to play, but if you describe the same games in chess notation, it can. Fowler thinks this is partly because notation shrinks the token count, but mainly because it is a more rigorous notation. That suggests we may need rigorous ways of speaking to LLMs. He connects this to domain-driven design's ubiquitous language and to his work about a decade ago on domain-specific languages and language workbenches.

The host asks whether this is the first time non-determinism has become so widespread in software, since earlier neural nets were niche. Fowler agrees it is a whole new way of thinking, but sees parallels in other engineering fields. His wife is a structural engineer who always thinks in terms of tolerances: how much margin to add beyond what the math says, because material properties vary and you design for the worst case. Software, he argues, needs similar thinking about the tolerances of non-determinism, and must avoid skating too close to the edge. He expects "some noticeable crashes," particularly in security, because people are already skating too close.

Where LLMs Already Help: Prototypes and Legacy Code

Fowler names two areas of clear success at opposite ends of the scale. The first is rapid prototyping and exploration, the territory of vibe coding. Someone unsure of an idea can spend a couple of days exploring it far faster than before. That covers throwaway explorations and small disposable tools, including ones built by people who don't consider themselves developers. There is good reason to be suspicious of taking this too far, he says, but kept within its bounds it is very valuable.

The second is understanding legacy systems. Colleagues at Thoughtworks did semantic analysis of code, loaded the results into a graph database, and queried it in a RAG-like style. They could ask what happens to a piece of data, or which parts of the code touch it as it flows through the program. Fowler calls this "incredibly effective." Using generative AI to understand legacy code is one of only four items in the Radar's Adopt ring. He explains that it came from people already doing legacy work who tried the approach. Legacy modernization is a constant concern for Thoughtworks, since every company older than a few years has the problem "in spades."

What's Still Unresolved

Beyond those two areas, Fowler says practices are still being worked out. For building decent-quality software one-on-one with an LLM, he sees signs that you need very thin, rapid slices. Every slice should be treated as a PR from "a rather dodgy collaborator who's very productive in the lines of code sense," which means reviewing everything carefully. He mentions Kent Beck's name for the tool, "the genie," and Birgitta Böckeler's anthropomorphic donkey, Dusty. Used well, he says, the speed-up is real but not as large as advocates claim. It is "non-trivial" and worth learning. He names Birgitta, Kent, and Steve Yegge as people pushing this work forward.

Most of this experience, he notes, comes from greenfield work. Whether LLMs can safely modify legacy code is still open. He tells a story from James Lewis, who asked Cursor to rename a class in a modest program. It took an hour and a half and used about 10% of his monthly token allocation. The host recalls that JetBrains' ReSharper made class renaming a paid right-click feature around 20 years ago, and that early Swift in Xcode lacked such refactorings, which drew complaints. Fowler says Lewis did it to see what would happen, since the deterministic tooling has existed for a long time. He takes the point to be that modifying existing systems is still "really up in the air."

Team dynamics are the other open question. Software is built by teams, and Fowler expects that to continue. Even if AI made developers an order of magnitude more productive, which he doesn't believe it will, a team of 10 would still be needed where 100 were before, and he sees no sign of demand for software dropping. How to work with AI in a team setting is something "we're still trying to figure out."

Vibe Coding and the Learning Loop

Fowler uses the original definition of vibe coding: you don't look at the output code at all, and you may not know programming. He considers it good for explorations and disposable work, but not for anything with long-term life.

His example involves Unmesh Joshi, who asked an LLM to produce an illustrative "pseudo-graph" of capability over time for an article and committed it. Fowler wanted to move the labels closer to their lines. He opened the SVG and found it "astonishingly" convoluted and "gobsmackingly weird," although he had written the previous version himself in about a dozen lines. The lesson he draws is that vibe-coded output can't be tweaked. You throw it away and hope regeneration gets you what you want.

The deeper problem, which is the core of an article by Unmesh that had just been published, is that vibe coding removes the learning loop. Programming involves constant back-and-forth between ideas and what the computer does. Fowler agrees with Unmesh that this process can't be shortcut. If you don't look at the output, you don't learn. If you don't learn, you can't tweak, modify, evolve, or grow what you produced. "All you can do is nuke it from orbit and start again."

The host admits to getting tired while reviewing large volumes of AI output and letting things slip, and says many engineers worry about rigorous review as code volume grows. Asked about approaches that help people keep learning, Fowler says he hasn't seen a huge amount. He finds Unmesh's direction most promising: working with the LLM to build a language that communicates more precisely what you want. He adds another use Unmesh has highlighted, which is getting oriented in unfamiliar environments. James Lewis, for example, was using the Godot game engine on a Mac with a language he didn't know well. Fowler himself asks LLMs how to do things in R that he has done twenty times but can't remember. Generating a starting skeleton project is another such use.

Stack Overflow, Testing, and Tools That Lie

The host compares the moment to Stack Overflow's arrival, when junior developers pasted snippets without understanding them. He cites a top-voted but flawed email-validation answer that spread widely. Fowler agrees it is similar but "on steroids." He raises a further question: who will write Stack Overflow answers in the future? He has no problem with pasting LLM output to see if it works. After that, you should understand why it works, consider whether it is structured the way you'd like, not be afraid to refactor it, and write a test for anything that works.

He cites Simon Willison as someone who constantly stresses testing, and notes that Birgitta, coming from Thoughtworks' Extreme Programming culture, says the same. The host describes LLMs claiming all tests passed when running them showed five failures. Fowler replies that they "lie to you all the time." If they really were junior developers, as people sometimes describe them, "I would be having some words with HR." The host adds a personal case: asked to add today's date to a config comment, the model copied the previous entry's date and, when corrected, put in yesterday's. They agree on "don't trust, but do verify."

Summarizing where Thoughtworks developers are finding success, Fowler repeats prototyping, legacy understanding, and exploring new technologies or even domains, "as long as you trust it significantly less than you would trust Wikipedia 10 years ago."

Spec-Driven Development and Domain Languages

On spec-driven development, which the host says Birgitta is exploring, Fowler sees a risk of repeating waterfall if people write a large spec up front and pay little attention to code. What matters to him is doing "the smallest amount of spec you can possibly get to make some forward progress," then building, testing, and ideally deploying, in thin slices with a human verifying each time. He says Birgitta agrees on that point.

What interests him is whether specs can become a more rigorous, domain-language-like way of expressing abstractions, more fluid than code allows but still parallel to the codebase in the spirit of ubiquitous language. He notes that some programmers have already written business logic in, say, Ruby that domain experts could read. The experts couldn't write it, but they could spot what was wrong and suggest changes that a programmer could then make syntactically correct. That takes deliberate effort in how the language is designed. The open question is whether LLMs will let this happen more widely.

The Enterprise World

Fowler says the corporate enterprise is the world he knows best. There, developers are a small minority, business processes are complex, and legacy problems are usually worse. When the host offers banks as an example, Fowler replies that banks tend to be more technologically advanced than most corporations; retailers, airlines, and government agencies are often further behind.

He describes speaking at an agile conference for the Federal Reserve in Boston. Its staff are not currently allowed to touch LLMs because the consequences of error are so serious. On a tour of the cash-handling area, he saw the care and controls around sorting and counting notes. That made him think of an adage that to understand an organization's software development, you look at its core business. An airline's concern with safety, he says, shapes its thinking in the same way, "or ought to."

Asked how cautious enterprises differ from nimble companies, Fowler stresses that big enterprises are not monolithic. Some parts are adventurous, as his own small group at Coopers & Lybrand was, and "the variation within an enterprise often is bigger than the variation between enterprises."

Refactoring: Origins, Revision, and Renewed Relevance

Fowler first encountered refactoring when Kent Beck showed him, in a Detroit hotel room early in the C3 project, how he refactored Smalltalk code. Fowler had always cared about making code comprehensible, but he was struck by how small Beck's steps were. Because they were small, they didn't go wrong, and they composed into large changes. Beck was busy with the first Extreme Programming book, so Fowler decided to write the refactoring book himself. He took careful notes each time he refactored, which became the "mechanics" sections, and added an example for each. He used Java rather than Smalltalk because Smalltalk was "dying sadly" and Java was, in late-1990s thinking, "the only programming language we'd ever need."

He stresses that Beck didn't invent refactoring. Ralph Johnson's group at the University of Illinois at Urbana-Champaign built the first Refactoring Browser in Smalltalk, with John Brant and Don Roberts. IBM's VisualAge team, originally Smalltalk people, was aware of it. JetBrains, with early IntelliJ IDEA and later ReSharper, made automated refactoring something developers could rely on. Knowing how to refactor by hand still matters, he says, because some languages lack these tools and some refactorings aren't automated.

The word became common and was misused to mean any change. Fowler insists refactoring means very small behavior-preserving changes, "each step is so small that it's not worth doing," strung together to achieve a lot. The host recalls colleagues announcing "I'm doing a refactoring" at stand-up for days on end.

The 2019 second edition switched to JavaScript, both to reach a broader audience and to be less object-centered ("extract function" instead of "extract method"). Fowler also wanted to refresh the examples and give the book "another 20 years of life" to "keep me going until I croak."

On whether refactoring has caught on, Fowler says it is hard for him to judge because he mostly talks with Thoughtworks people. He reads plenty online that makes him shake his head at how refactoring is described, let alone practiced. He maintains that the disciplined approach is actually faster. Progress has been real, especially in tooling, "maybe not as much as I'd have hoped for."

Looking ahead, he says he isn't seeing it yet, but expects refactoring to become increasingly important. If AI produces a lot of code of questionable quality that works, refactoring is how to improve it while keeping it working. He says current LLMs "definitely cannot refactor on their own." He points to Adam Tornhill's work combining LLMs with other tools. The host notes that describing a small refactoring to a command-line agent often takes longer than doing it in an IDE.

Fowler sees promise in LLMs as a front end to deterministic tools, much as people use LLMs to draft SQL they can then check and adjust. He recalls a large company's API migration across a big codebase, about a year earlier, that was reported as an LLM achievement. By his account it was roughly 10% LLM and 90% another tool, whose name he couldn't recall. The interesting interplay, he says, is "using the LLM as a starting point to drive a deterministic tool," so that you can see what the deterministic tool is doing.

Why Patterns Went Out of Fashion

The host notes that design patterns were common interview topics around 2002, when Patterns of Enterprise Application Architecture appeared with more than 40 patterns such as Lazy Load and Identity Map, but faded from mainstream discussion in the 2010s. Fowler explains that patterns aim to create a shared vocabulary, much as medicine uses Greek and Latin jargon. They are useful for describing alternatives and deciding when to apply something, not for cramming as many into a system as possible. He suspects they may have lost favor partly because people used them "like pinning medals on a chest." He recently worked with Unmesh Joshi on Patterns of Distributed Systems and still finds the form valuable. Why patterns became unfashionable, he says, is hard for him to judge, and they may come back.

The host suggests startups treated patterns and strict UML as legacy baggage, preferring whiteboard boxes, and that each generation reacts against the last. He also relays Grady Booch's view that well-architected cloud building blocks such as managed databases reduced the need for architectural thinking. Fowler suspects there are still patterns for using those services but hasn't explored them, partly because colleagues haven't brought him draft articles on the topic.

Both agree that every organization develops its own jargon, including internal system names. Fowler says books like his try to offer shared words across environments, and the mismatches only show when you move between them. He tells of someone who joined an established bank from a startup and said that after three years they finally understood the problem. Such systems are not logical, Fowler says, because they are built by humans: vendors favored by one manager, reorganizations, and personal history accumulate into "a complicated mess." Uber, he says, is lucky to be young, but if it survives 50 years it will look like American Express, whose re-architecture planning the host once learned had taken three years.

The Agile Manifesto

Fowler traces the manifesto to a gathering Kent Beck held about a year earlier in rural Oregon. The group debated whether Extreme Programming should stay narrow or broaden. Beck chose narrow, which left open what to do with the broader set of ideas that overlapped with Scrum and others. That led to the 2001 meeting in Utah, chosen over Dave Thomas's preferred Anguilla. Fowler says he wasn't heavily involved in organizing it, though Bob Martin insists they discussed it over lunch in Chicago.

He remembers little of the meeting itself and regrets not keeping a journal. He would love to know how the "this over that" structure of the values came about. He has a fairly clear memory, which he cautions may be unreliable, of Bob Martin pushing for a manifesto, and of himself thinking the document would be ignored but the exercise of writing it would help the group understand each other. Its impact came as a shock. He quotes Alistair Cockburn: a brilliant idea will be either ignored or misinterpreted, and you don't get to choose which. He points out that the text says "we are uncovering," making it a snapshot of 2001.

For Fowler, agile's success is practical. In 2000, clients resisted what Thoughtworks wanted: tests, automated builds, small increments. The standard approach was a five-year plan with two years of design before implementation and then testing. Thoughtworks wanted to run that entire process for a subset of requirements in a month, "and of course we really wanted to do it in a week." Now clients let them work much closer to how they want. He wanted "the world to be safe" for people who wanted to work that way. Despite bad side effects, he thinks the industry is better off. Measured against the original vision, though, progress is "a pale shadow," and he suspects most of the surviving 17 authors would agree.

Cycle Time in the AI Era

Asked whether AI changes agile, since customers may want quality up front, Fowler says it is too early to know. He still bets on small slices reviewed by humans. He hopes AI lets teams do slices faster, and he would rather have smaller, more frequent slices than more work per slice. He says the biggest gains have come from faster cycles: look at queues in your flow, and if an idea takes two weeks to reach running code, work out how to make it one.

The host describes how Boris, one of Claude Code's creators at Anthropic, built 20 interactive prototypes of a progress-display feature in two days, recording each prompt and output. Fowler ties this back to feedback loops: tightening them so we learn faster about what we are trying to do.

How Fowler Learns, and Whom He Trusts

Fowler says his main way of learning now is editing articles by practitioners for his website. He doesn't do day-to-day production work; the only production code he writes runs his site. Experimentation comes second, when he has time. He reads Birgitta and Simon Willison, whose output rate he envies, and follows Kent Beck, joking that much of his career has been "leeching off Kent's ideas." He occasionally watches videos, "although I really hate watching videos."

He is thinking about how to judge sources in what he calls an epistemological crisis, and plans to write about it. One signal he trusts is a lack of certainty. In the late 1990s he looked for architecture writing from the Microsoft world and found Jimmy Nilsson, whose book was openly tentative about how he currently saw things. Fowler found that trustworthy. He is also drawn to writers who lay out trade-offs. Anyone who says always, or never, use microservices "can be completely discounted." Clients who ask for cookbook answers will get into trouble, he says, because anyone offering one either doesn't understand the problem or is hiding the nuance.

Advice for Junior Engineers

Juniors should use and explore AI tools, Fowler says, but they lack the judgment to tell whether output is good. His advice is the same as ever: find a good senior mentor, who "is worth their weight in gold," and consider prioritizing that above much else in a career. Meeting Jim Odell was "blind luck" and the best thing that happened to him, and he still thinks of Kent Beck as a kind of mentor. With AI, remember it is "gullible and it's likely to lie to you." Ask it why it gives that advice and what its sources are, as you would ask a person about their context, because it is regurgitating what it saw online, good or bad.

Rapid Fire

Fowler's current favorite language is Ruby, from long familiarity, but his love is Smalltalk, which he calls unmatched for fun in the 1990s. He mentions the Pharo project and wishes he had weeks to explore Smalltalk again. For books, he recommends Daniel Kahneman's Thinking, Fast and Slow for building intuition about probability and statistics. He thinks the world would be better if more people understood statistics, and wishes school maths had emphasized it over calculus. Tabletop gaming has sharpened his probabilistic reasoning. He also recommends The Power Broker, about Robert Moses, who was never elected yet controlled more money than New York's mayor or governor for about 40 years. He values it for showing how power works in a democracy, often out of plain sight, and for writing so good he would stop to appreciate passages. Reading excellent writing, he says, makes you a better writer. At 1,200 pages it's long, and he is working through the same author's multi-volume, still-unfinished biography of Lyndon B. Johnson. His board game pick is Concordia, fairly abstract, easy to learn, and rich in decisions.

The State of the Industry

Fowler is positive in the long term: technology can still do huge things, and demand exceeds what we can imagine. In the short term, he describes something like a depression in the developed world's software industry, citing figures he has heard of a quarter to half a million jobs lost. Thoughtworks grew about 20% a year until around 2021 and then "hit a wall," as clients stopped spending. AI is "clearly bubbly," he says, but bubbles are unpredictable in size, timing, and aftermath. Unlike blockchain and crypto, he believes there is real value in AI, though how it plays out is unknown. He sees it as a repeat of the late 1990s and 2000s at perhaps an order of magnitude more scale.

The most important thing to hit the industry, in his view, is not AI but the end of zero interest rates. Job losses began before AI because of it. Macroeconomic uncertainty, including what he calls "Looney driving the bus" in the United States, keeps businesses from investing. The result is a strange mix of near-depression in software alongside an AI bubble.

He still thinks software is a good profession to enter, though the timing isn't as good as 2005. He does not believe AI will wipe out software development. It will change it as dramatically as the move from assembly to high-level languages, but the core skills remain. Those are less about writing code than about knowing what to write, which means communication, especially with users. From his healthcare modeling work, he notes that he learned a great deal about healthcare processes, but you would never want him treating you, because he is not a doctor. That is why domain experts must stay involved, and why he sees effective collaboration as what distinguishes the very best developers.