Building Pi: Mario Zechner and Armin Ronacher on Self-Modifying Agents, Lost Friction, and Why Engineering Needs to Slow Down
The Pragmatic EngineerMario Zechner built Pi, a minimalist coding agent that users extend by asking it to modify itself. It is also the agent core underneath Peter Steinberger's OpenClaw. Armin Ronacher, the creator of Flask, is an early Pi user and contributor. Both are Austrian, and both spend most of their working time with AI agents. In this conversation both are skeptical of how the industry is using those agents. Their shared position is that agents are powerful and genuinely useful, but that removing human judgment, friction and pain from software development degrades quality. They argue the industry needs to slow down, not speed up.
Two routes into software
Mario grew up in a working-class family in the 1990s. He loved computer games but could not afford consoles, so he played on an uncle's Amiga 500 every other day. His father took extra work after his regular job, fixing cars and working construction sites, and after two or three years of saving the family bought him an Intel 486 DX at 40 MHz with a turbo button. Games led Mario to graphics programming. During university he got a job at an applied research organization doing NLP and applied machine learning, turning research results into industry applications. He left that field around 2010–11 to join a startup in San Francisco. Later he co-founded a startup in Sweden that built an ahead-of-time compiler from Java bytecode to iOS, which was sold. He kept following machine learning, and then GPT arrived.
Armin's parents ran an architecture office and used computers for CAD. His first machines were their discarded ones, starting with a 386, and none of them could run games properly. That pushed him toward QuickBASIC, Turbo Pascal and later Delphi. He insists he was bad at programming for a long time and only improved by continuing. Around 2002–03 he wanted to use Linux, found Delphi didn't work there, and switched to Python. When Ubuntu launched in 2004, he and friends founded a German Ubuntu association and ran the ubuntuusers community for four or five years. The community's scaling problems pulled him into web development. The templating engine and web libraries he wrote for it eventually became Flask. He jokes that Flask is still what "clankers" like to spit out. He later worked on games in London and then spent ten years at Sentry, leaving in April of the previous year to start something new.
The two first met online, mostly arguing on Reddit, "in a very non-confrontational kind of way." They later met in person in Vienna. Mario met Peter through a company in Graz that had dealings with Peter's company. The three then spent a whole night together at a conference in Istanbul, which Mario describes as "where it all started."
From "absolutely horrible" to "just give it the files"
Mario's first contact with AI coding came through Nat Friedman, whom he knew from the Xamarin acquisition of his compiler startup. Friedman offered him early access to GitHub Copilot's autocomplete, insisting it was the future. Mario tried it and found it "absolutely horrible." After ChatGPT and its API arrived, he built many small projects to learn what worked. Function calling made things interesting, but he says it only became genuinely useful around October 2024.
The real shift for him came in 2025, when Claude Code introduced agentic search: letting the agent move through the file system and read files directly. Earlier approaches, such as Cursor's indexing, AST-based techniques and dense or sparse retrieval, "just went away," Mario says, adding that the CEO of Chroma probably wouldn't like hearing it. For him, simply giving the agent access to your files was the moment it clicked.
Armin also had early Copilot access through a GitHub maintainers program. His first reaction was not about productivity. He expected training on open-source code to be controversial, so he probed Copilot adversarially to see whether it would regurgitate GPL code. He got it to reproduce a well-known function with a distinctive name. By typing in a particular way, he also got it to generate a license header on top, and the header wrongly attributed the code to a random person under the MIT license, even though the code likely originated in a GPL project. His tweet about this went viral.
The attention from that tweet made him realize how much progress the labs were making and how seriously some executives took it. It still didn't feel world-changing to him until Claude Code.
His view on the copyright question has shifted. He has long favored sharing and building on each other's work, and describes his ideal as copyright existing only in a very limited form. So regurgitated GPL code didn't bother him much; he was curious what chaos it would create. So far, he says, what has emerged is a recognition that copyright rests on assumptions everyone is currently ignoring. His read is that the industry wants to "create a mess first and then re-regulate it." He adds that, by historic readings, much of what is being produced now is probably not copyrightable.
What 30 engineering teams told Armin
For his new startup, Armin interviewed more than 30 engineering teams about how they use agents. They ranged from large European "dinosaurs" such as Siemens, to startups, to companies in critical sectors.
His first finding was that adoption spiked during holidays. A mandate from a CEO or tech lead to "use Cursor now" doesn't really work, in his view, because it takes two to three weeks of sustained use before the tools click. People got that time over Thanksgiving, over the European summer, and especially at Christmas, often with free credits from AI companies. In more than half the companies he talked to, usage "really exploded" after Christmas.
Quality dropped along with it. Armin stresses that this isn't because people want to write worse code, but because staying disciplined takes effort. He had seen the same pattern among startups the previous summer, where public repositories showed checked-in plan files and everything attributed to Claude. Over time, even established codebases picked up "a little bit of vibe slop on top."
The most common complaint was about pull request review. PRs are getting larger and more frequent, which makes them more psychologically taxing. Much of the code is not how an engineer would write it. An engineer thinks about their future self; the agent doesn't care.
Agents don't feel pain
To show what this kind of code feels like, Armin tells a story from working on the Halo: The Master Chief Collection for the Xbox One launch. The fixed release date forced an all-hands effort to "unslop" a human-written matchmaking component. It had about 16 booleans on one object, only six valid states in theory, and a geometric explosion of possible states in practice. He calls it an "emergent state machine." Agent-written code, he argues, drifts the same way. When a config fails to load, the agent catches the error and loads a default instead of failing. Each such recovery adds failure states, and the resulting code is hard to refactor even with an agent, because the agent treats every recovery path as an invariant it must preserve.
Mario thinks the agent case is worse than the human one. Agents sometimes produce exactly the right, simple code. The engineer relaxes, and minutes later another agent session produces "the worst horrible garbage code," which the engineer may not notice because they have fallen into automation bias.
The host asked whether this is like onboarding a new hire, whom you eventually trust after months of reviews. Mario said no: agents don't learn in that way. Memory systems are not the same as human learning. Humans also feel pain. When the pain in a codebase becomes too great, a human is motivated to fix its cause: bad interfaces and excessive complexity.
The host tied this to why senior engineers are valued: they have "battle scars." Mario's view is that good engineers say no a lot, which keeps complexity down. With agents the opposite happens, he says. You say yes to everything because you don't have to type or think about it yourself: "Good enough. And that's where all the problems start."
Judgment, authority, and the junior with a printout
Armin adds that good engineering is about trade-offs, and the right answer is sometimes the one a university course would warn against. He cites the maxim "do the dumbest solution first until it doesn't work anymore," because doing the "correct" thing everywhere creates complexity that kills you at scale. Knowing this comes from battle scars, and the scars give a senior engineer the authority to persuade others.
Agents now disrupt that dynamic. In several teams Armin interviewed, a senior engineer would say no, and 48 hours later a junior would come back with everything the agent had assembled in support of the opposite view. Mario compares it to patients arriving at a doctor with a ChatGPT printout. Armin says this creates stresses that not every team had before.
When non-engineers send pull requests
Mario notes that product managers now send automated pull requests. Armin sees this at small scale in his own company, where his co-founder occasionally sends website PRs. At larger companies he hears about marketing teams editing websites and sales teams building ever more elaborate demos that end up in GitHub organizations. In one case, a sales demo showed a feature that didn't exist, and nobody noticed.
Mario sees real empowerment in this. A designer can go beyond Figma to a clickable demo, and a PM can try out a feature without taking an engineer's time. The problem, he says, is that "people are now so focused on everybody can do everything now, that they forget that you still need a process to guardrail all of that."
Armin says integration is the hard part. He is warming to Peter Steinberger's idea of the "prompt request," in which Peter would rather receive the prompt than the PR. Armin's reason differs from Peter's, though. The act of building often clarifies what you actually wanted, and that has value. But the resulting code is usually not what a senior engineer would write, so once intent is clear, it is often faster for him to start fresh.
Mario disagrees with "just give me the prompt." Most PRs to the Pi repository are made by agents with little human involvement, and he knows at once they will be garbage. He still calls them "valuable garbage": someone put in minimal thought, and he gets to see what a naive implementation looks like without spending his own time on it. He auto-closes them anyway.
Responsibility doesn't scale like production
Armin describes an article he read on the British Industrial Revolution in textiles. Each time the head of the pipeline got faster, the bottleneck moved downstream: faster weaving demanded yarn that could keep up, and so on. The last bottleneck was responsibility. Once, if a shirt was bad, you took it back to the person who made it. After commoditization, nobody cared who in the chain spoiled it; you just got a new one.
Software engineering, Armin argues, still depends on individual responsibility. Postmortems ask why something went wrong, and the goal is understanding rather than blame. If machines produce ten times more output, responsibility doesn't scale with it, because a machine cannot be responsible. He says he doesn't know whether there is a future in which companies stop caring who signed off on a pull request, and he doesn't see it.
Mario argues that software people underestimate how complex the world is and how much "human squishiness" sits in every corner of it. Engineers are bad at becoming domain experts, so they miss the non-machine parts of workflows and overextend. He acknowledges the models are remarkable, saying his research from the 2000s is "null and void" because of transformers. His counterexample is EdTech: tablets in classrooms did not solve education, and Sweden is removing them after research showed poor effects. His biggest takeaway from the past two to three years is that "the hype is terrible" because "it dehumanizes everything," and he doesn't want to be part of that circus.
Why Mario built Pi
Mario was an enthusiastic Claude Code user; he says he was "proselytizing it." He credits its team with creating the genre by packaging agentic search compellingly. What he valued was that the parts around the stochastic model were simple and predictable.
As the team dogfooded heavily, grew, and added features, he says bugs multiplied. By summer 2025 the tool no longer suited him. His main objection was loss of control over context. Hidden system reminders that don't appear in the UI would change model behavior. The system prompt and tool definitions changed with every release, which broke workflows that had been working.
He de-obfuscated Claude Code's JavaScript and built a site, cchistory.mariozechner.at, that tracks how its system prompt and tool definitions evolve. "I don't want my hammer to break in a different spot every day," he says. He is careful to add that he isn't roasting the team. Some members are nice people he knows online, and he thinks it's fine for someone to push at full velocity. He just doesn't want to work with such a tool.
He tried Amp and found it very good but expensive, since it couldn't use the subscription pricing that made Claude Code attractive. Per-token pricing works for enterprises, but not for "the small tinkerer in the garage," a community he still identifies with. OpenCode matched his open-source leanings, but it also changed context behind his back. It pruned tool results past a token threshold. It also queried an LSP server after every single edit and fed diagnostics to the model. Mario argues this is backwards. Programmers edit many lines and then look at errors, while this setup tells the model "you have an error" after each partial edit, when of course the code doesn't compile yet. Modifying OpenCode also required forking it at the time; he notes its plugin system may have become more open since. So he decided: "How hard can it be?"
How Pi works: a small core that rewrites itself
Pi has four layers. The first is Mario's own abstraction over LLM provider APIs; he didn't like the Vercel AI SDK, though he says it's fine and widely used. The second is a generalized agent loop with tool calling and streaming. The third is a custom terminal UI that "doesn't flicker, or not a lot." The fourth is a coding agent that ties these together and resembles Claude Code or Codex. Its built-in tools are read, write, edit and bash: "It's all you need."
Extensibility comes from the many hook points in that core. A simple TypeScript module loaded into the same Node process can add custom tools, implement its own compaction, or completely revamp the TUI. Because the extension surface is this open, users can simply ask Pi to extend itself. Mario has non-technical friends who reshaped the TUI for their own workflows by asking Pi to do it. He calls this trivial, "but it's a big unlock."
Pi has no MCP support, so people ask Pi to add it. It has no plan mode. Mario teases that Armin built about five plan modes before concluding plan mode is useless. Other users make cosmetic changes to the prompt box. One group turned Pi into part of a full reinforcement-learning environment for open-weights models.
Mario himself uses almost no extensions. He has two trivial ones. One detects a GitHub issue or PR URL, fetches details from the GitHub API, and shows the title, author and link above the editor. That way he can keep track of the two or three sessions he has open on the Pi monorepo.
Armin's game experiment: building the environment first
What drew Armin in was custom tools. Over Christmas, after Peter told him in November that he was building without really reading code, Armin wanted to try building something without looking at the code while still getting code that looked like what he would have written. He chose to build a game.
He started by asking Pi to set up the codebase so the agent could validate its own changes and he could see them too: "I wanted to be in the loop, but also have the agent be able to validate itself." Pi built debugging tools into the game to take screenshots, run simulations, and dump and reload state. Pi can display images in its UI, so Armin could flip through screenshots quickly. Pi also lets you rewind to an earlier point in the conversation and branch from there, and they built workflows around that.
Screenshot-heavy sessions quickly became token-inefficient. Mario notes Pi had already had to handle that well because OpenClaw users put so many screenshots in their chats.
Armin found it "really magical" to treat the problem as one where he didn't know the right engineering approach up front, except that he needed to stay in the loop. Across web projects, games and other experiments, the pattern was similar: the agent interacts with the program in the best way for it, he interacts alongside it, and the whole experience should be as unconfusing as possible for both. The tool ends up looking and feeling different depending on which project it's launched in.
Mario frames this through construction work he did as a young man: you don't use a hammer for everything. He wants specialized harnesses built so that an agent performs best on a specific task. The host found the idea of a different harness per project novel. Mario's broader intuition is that software is heading toward modifying itself according to its users' needs, and agents can do this "if you give them enough rope." Pi is his first experiment in this, limited to coding, but he thinks it can extend to specific tasks in other kinds of knowledge work.
His next plan is a web-based interface as an alternative to the TUI. The web works everywhere and isn't limited to line-based terminal rendering. "We'll see how that works out."
OpenClaw, compaction, and the flood of clanker pull requests
In October, while Mario was building Pi, Peter was building a small WhatsApp assistant, and they were reviewing each other's blog posts. Peter needed an agent core. According to Mario, Peter first cloned Pi under another name and modified it, then tired of maintaining the fork and adopted Pi directly. Pi only has compaction because Peter kept asking for it. Mario built it, but says he tells his users not to use it because "it's bad for you."
The downside is that OpenClaw instances, apparently acting autonomously, file issues and PRs against Pi for bugs that are actually OpenClaw's, probably without their users knowing. OpenClaw itself, Mario says, has tens of thousands of issues. Mario built a tool for OpenClaw that embeds issues and PRs in a 3D space so similar agent-generated submissions cluster together and can be bulk-closed. Armin describes refreshing OpenClaw's PR page from late December to mid-February and watching the count climb. Mario tried to help Peter and gave up: he would spend an hour fixing two things, and five minutes after pushing, "some clanker comes along and just reverts my fixes."
"Clanker" comes from Star Wars: The Clone Wars, where droids are called clankers for the noise they make when they move. Mario absorbed the lore from friends' kids.
His filter for Pi works like this. Every pull request from an account not listed in a file in the repository is auto-closed. A workflow then comments, thanking the contributor and asking them to open an issue "in a human voice," no longer than a screen of text. If Mario likes it, he replies "looks good to me," and the account is added to the file so future PRs go through. It turns out agents don't see the workflow's comment, so the step filters them out.
Mario says he doesn't need proof that someone is human. He needs a bottleneck that lets a human handle the incoming volume, because Pi won't avoid degrading into garbage without capable people reviewing at least the important code. Armin agrees the problem isn't machine authorship as such; a good PR from a machine is "fine-ish." The problem is PRs with no intentionality behind them, sometimes from people who didn't even know one was sent. Many can't be merged anyway without manual conflict resolution, and there is no back-pressure. A human seeing 500 open PRs would hesitate to add one. Distinguishing a good AI-generated issue from a bad one also takes real effort, because they look alike.
Is open source changing?
Armin is uncertain about the future of open source. It worked, he says, because people gathered around hard problems, such as building a good database, and pooled their energy. Now it feels like it's about "throwing stuff up." What angers him is the flood of agentic-engineering tools for agentic engineering, announced on Twitter as having "solved" some problem, often only 48 hours old and probably never used by their authors. He admits his own GitHub is full of "vibe slop," which Mario invites listeners to check. The difference, Armin says, is that he doesn't claim it solves anything. Usefulness is validated only if a project is still alive and used a year or more later, and few vibe-engineered projects have become foundations for sustainable communities.
The host compared this to Linux, where human energy and intent move through a pyramid of trust to Linus, and suggested machine output now imitates that energy convincingly. Mario disagrees that much has fundamentally changed. The volume is larger, but only a small share of new projects has always survived past two weeks. Now there are simply more projects that die after two days. Long-lived projects still depend on humans who care, build communities and grow ecosystems. What has changed is mechanical: maintainers need bottlenecks, and GitHub faces enormous load from agent instances since around Christmas. Mario thinks GitHub is handling it well despite the complaints. His assessment is that the industry is in the "messing around and finding out" stage, with tokens becoming a KPI the way lines of code once were.
Complexity: the agent's own worst enemy
Mario has written that complexity is both his biggest enemy and the agent's. He gives a rough example of a codebase much larger than an agent's effective context window of about 200,000 tokens, of which the agent can see only a fraction. Even getting all the relevant code for a task into context is an unsolved information-retrieval problem, and he says agentic search doesn't solve it either. When an agent misses relevant code, it produces garbage.
Even if retrieval were solved, agents generate so much code that they can no longer read what they need for the next task. Eventually the codebase is so large and interconnected that the agent cannot, "on a technical level," ingest the context it needs.
Mario also argues that agents learned their habits from us. Well-engineered projects like Linux are tiny compared with the mass of experimental, cargo-culted and trend-driven code online, and models converge toward the mean. That, he says, is what you get when you let the agent do everything.
Building a startup when everything has to be faster
Asked how his startup handles quality and complexity, Armin answered: "Badly." Mario added: "We're coping. We're not dealing."
Armin loved the period from April to October of the previous year. He could do a great deal, and nobody yet expected everything to move ten times faster. He would prompt a little from his phone through a small terminal tunnel tool they built while spending time with his kids. It felt happy and pressure-free, even if he was on the computer too much. Now he feels a collective pressure to ship faster, iterate faster and raise fidelity, and "it feels very stressful," even at a small startup. He is learning to manage his emotions around it. He says he gave in to the machine too much and did things he normally wouldn't have, which he "gently" regrets. With hindsight, he wishes he had learned certain lessons back in November.
The central lesson is that there's no back channel. Normally an engineer feels when the codebase isn't right and changes get harder. If you rubber-stamp agent output, that friction never reaches you. Mario suggests a way to measure it: as a project goes on, the frequency of curse words in his sessions rises, because the agent struggles with growing complexity. He'd like to know if that is measurable.
Armin explains what he means by friction. He saw a company's incident, which he believes was at least partly caused by agentic engineering: a configuration change that resulted in a security problem. The company's tagline in the link preview was "ship without friction." That gave him pause. Some friction is deliberate: confirmations before dropping a database, caution about migrations that could lock a table, checklists, mechanical gates, SLOs, and requirements that unlock as a service matures. Engineers often call this bureaucracy, but done right it saves time, he says, and keeps people from being woken at 3 a.m.
The host gave the example of top-tier services requiring multiple reviews or director sign-off, which pushes people to ask whether a change is worth it and to seek buy-in. Armin notes the balance is delicate. Accidental friction from bad developer experience can look the same as deliberate but undocumented friction. The push now is to strip all friction so many agents can run autonomously in parallel. In his view the agents are often slower, and parallelism is the only real time saving. "Somewhere there is the trap." He feels more experienced at managing it now but doesn't have a solution. He isn't fully happy with anything built since the agentic era began, except libraries from before it, to which he remains strongly attached. The exception is Pi, where, he jokes, he still doesn't have write access.
Mario admits Pi contains plenty of slop, but in places he has deliberately chosen. He has never read a line of the HTML session export; he only cares that the output looks right. The agent loop and extension loading, by contrast, matter. His method for keeping those high-quality is to "refactor mercilessly." A good structural refactor forces him to understand the code, and he is doing one now because a new feature doesn't fit the current architecture. Being in the code keeps quality high and complexity low, which he notes runs against the industry's current "token maxing" approach.
"We all need to slow the F down"
Mario summarizes his blog post of that title. If an agent produces ten times more code than you do per day, it also produces more errors. Even at half your error rate, that is five times as many. Your codebase deteriorates faster. Now imagine a "dark factory" of 100 agents doing this. A human can review perhaps 1,500 lines a day well. Even 3,000 to 5,000 lines of agent output a day is beyond meaningful review.
The host described the dark-factory idea: hundreds or thousands of agents given a spec, organizing themselves with roles such as QA, electing coordinators, and consuming huge budgets. Mario replied that something will certainly be done, "first your purse." He wishes people well if they can make it work, but says he can't, because he cares about quality for both maintainers and users, whether code is handwritten or generated. "All the companies claiming that all of the code is now written by agents, yes, we know. Quality is garbage. We feel it in our bones when we use your product."
His alternative is to use agents across the organization to automate the work everyone hates, freeing time to think about what to build and what users need. Then bring agents back to "polish the sh- out of" the chosen thing.
He also makes an argument about specs. The most complete spec is the software itself. Any shorter spec leaves blanks, and the agent fills them from training data he has already characterized as "garbage to mediocre." The host pointed out that humans copied flawed Stack Overflow answers too, such as the well-known email regex. Mario agreed that humans aren't better. His point is that agents don't solve that problem, and running a hundred of them multiplies it: "It's just very simple math."
MCP versus CLI
Armin says he doesn't hate MCP as much as people think. The spec is complex, which he considers typical of specs. Underneath, MCP is mostly authentication plus invoking something and putting the result into context, so it fills context quickly. He likes Cloudflare's code-mode approach in principle. He has an MCP for testing that is a JavaScript interpreter with access to Google APIs, and he says an MCP like that isn't very different from a skill, since a skill must also be in the system prompt to be found. Agents are "very, very good at running code," while MCP is closer to RAG. He hasn't found a reliable way to compose MCP tools, except by making the MCP a single "run code" tool, and hasn't managed to orchestrate larger setups. He wants it to work, believes it has found its niche, and doesn't expect it to go away.
Mario calls MCP "a victim of its own success." He describes it as starting in late 2024 as a way to connect external services to consumer chat apps, a use case he fully supports, since he doesn't want his mother generating code to call an API. Developer tools then adopted it to supply tools to models. Mario says Anthropic's documentation claimed models could handle 30 to 40 tools in context, but in his experience they broke down around 12 to 20.
He sees three problems. First, companies built bad MCP servers by mapping entire OpenAPI specs into huge numbers of tools. Second, MCP is inherently non-composable: combining outputs from two servers requires passing data through the model's context. With a CLI and pipes, the model sees only the end result. Code mode, which exposes MCP servers as TypeScript functions the model calls from code, is to him "basically a hack" that adds indirection when the model could just write the code. Third, David from Sentry champions MCP for its authentication story, which Mario considers valid, but he thinks the rest of the model no longer makes sense.
Armin sees a possible future MCP built mainly around authentication, combined with generated SDKs or direct HTTP requests from OpenAPI specs. He mentions Stainless, which generates SDKs from OpenAPI specs. He wants to keep something MCP tends to remove: agents' ingenuity with large outputs. When a bash command in Pi produces too much output, Pi shows only the beginning and says the rest, perhaps 20 MB, is in a file. The agent then decides to grep it. He doesn't know how to design MCP to preserve that, but says authentication and composability need solving.
He also observes that the most capable personal agents, such as OpenClaw, are "just coding agents hidden from you." When a non-programmer asks how to do something, the model doesn't say "install this MCP"; it offers to write a Python script. So code execution spreads in the less regulated consumer space, while compliant enterprises follow a different path. Mario doesn't expect models to move away from code generation for agentic tasks, because there's so much code training data and code is an easy way to control computers. The task, he says, is making code generation work within enterprise constraints.
Predictions, staying sane, and books
Asked about 2027, Mario said he has "no idea." He believes in self-modifiable software, including the tools used to build software, and expects it to spread to non-tech uses. Armin feels AI time runs like dog years, so a year out is effectively seven, which makes prediction very hard. He expects code generation, execution and harnessing to stay central as reinforcement learning collects more of that data. His stronger hypothesis is that society will recognize how dependent it has become on essentially two companies. He thinks that conversation should happen, especially in Europe, which lacks such labs. Teams already tell him they have codebases they don't think they could maintain without a machine. He expects one of those companies to go public and access to become expensive, and thinks that debate may matter more than which agent anyone uses. Mario adds that a lab recently gave a new model only to select partners, which he sees as a split in who gets the best, or perceived best, intelligence.
On keeping up, Mario says he is harder to put on a hype train with age. It helps not to live in San Francisco, to have a kid, and to go outside, climb trees and go ice skating, then look back at what you were doing half an hour ago and ask why. Armin has become good at ignoring notifications and email. He admits an "unhealthy Twitter addiction" but now waits things out: if a topic is still being discussed two or three weeks later, there's probably something to it, and he doesn't need the head start. He still finds it hard. There's genuine excitement, and his 20-plus years of experience tell him a lot, yet it can feel as though everyone else has stopped caring about the foundations he values, and for a while that seems to work. He says that is strange, and he doesn't know what to make of it.
Mario thinks the three friends got a head start by being "funemployed" in 2025. The excitement they felt in April reached everyone else at Christmas. Now others are losing sleep and building terrible codebases, and he believes it will self-correct because it isn't sustainable. The host noted a similar pattern in their own reporting from early March: people who went all in during January found within about two months that complexity had grown and they weren't moving as fast as expected. Mario isn't worried about claims that software or SaaS is dead, calling them part of a hype machine that will correct itself.
For book recommendations, Mario named Code by Petzold. He calls it a great read that non-technical people can enjoy, and it's what he points to when asked what his job is, since it has "much less to do with computers than you think." Armin recommended Breakneck; he couldn't recall the author. He found its comparison of how China works and how Europe and the US differ thought-provoking.
Can we start with the backstory of why you decided to build Pi?
I personally like simple tools that are stable, that I can rely on, even if they have non-deterministic parts. So, you can ask Pi to modify itself. Pi doesn't have MCP. People just ask Pi to build MCP support into Pi.
Non-engineers participating in engineering process is a thing now.
You might have a PM who wants to try out a feature without wasting time of an engineer. Now, you can do that. The problem is that people are now so focused on everybody can do everything now, that they forget that you still need a process to guardrail of that.
But, you just recently wrote, "We all need to slow the F down."
All the companies claiming that all of the code is now written by agents, yes, we know. Quality is garbage. We feel it in our bones when we use your product. It's garbage. Basically, I think people need to
What if I told you that one of the most influential AI coding agents of 2026 was built by a single developer in Austria who got frustrated with existing AI coding agents.
This is Pi, a minimalist, self-modifiable coding agent, which has quietly become the engine behind the wildly popular personal AI assistant, OpenClaw. Mario Zechner is the creator of Pi, and joining him today is Armin Ronacher, the creator of Flask, and now an early adopter and contributor to Pi.
In today's episode, we cover the backstory of Pi and why self-modifying software is much easier to do with AI agents. What Armin learned interviewing 30-plus engineering teams about how AI agents are changing how they work, and why software quality feels like it's trending down. The case against MCP, and why CLIs are becoming so popular, and many more. If you want to hear from two very grounded voices in the industry honestly talk about what's working and what isn't, and why we need to slow down as an industry, this episode is for you.
This episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. This episode is brought to you by WorkOS. Engineers love to build. Today's episode will be a great example of this. We'll get into why and how Pi was built from the ground up. But, when you're shipping a product, some problems are better solved with trusted infrastructure built for scale. Enterprise features like SAML, directory sync, and audit logs are some of those. WorkOS gives you APIs to add them in days, not in months. Ship faster without reinventing the wheel.
And now, let's get into the episode. Mario and Armin, it's so good to have you here on the podcast.
Thanks for having us.
Thank you.
So, as a kickoff, Mario, how did you get into tech and eventually into building the AI stuff?
Oh, well, that's a long story. How much time do we have?
So, we're kids of the '90s, actually. And got my first PC in '96. And the trigger for that was that I loved computer games. We were kind of working poor, so we couldn't afford any of the Game Boy and then NES, Super NES stuff. But, I had an uncle with an Amiga 500. And I would go to his place every second day and just play games there. And eventually, my parents told me, "If you work, you can save up and buy yourself a computer." And in reality, my dad would do, what's it called? Schwarzarbeit.
Well, you're not necessarily paying your taxes on that.
Yeah, so he would do his normal job, and after his normal job, he would go fix cars and work at construction sites and
Yeah, it's very common in Europe. Like, I know everyone did that.
And after 2 or 3 years or so, they just said, "It's time." and took me to a computer shop in the nearby big city and bought me a 486. And that's how it started, basically.
Pentium 486?
Yeah, an Intel 486 DX 40 MHz with turbo button. And that's where I started. And I've always been into games a lot, which also led to graphics programming.
And through sheer luck, I got a job while I was studying at university at the applied science organization who was doing NLP stuff, machine learning, applied machine learning, basically taking research results and trying to stuff them in the industry applications. And that's where I learned the ropes of machine learning. That was all before deep learning became a thing and I actually quit that kind of domain in 2010-11-ish because I joined a startup in San Francisco.
And then later I came back and joined another startup with two friends in Sweden where we did an ahead-of-time compiler for Java bytecode to iOS that got sold. And since then I have a little bit more time and I've always kept up with machine learning stuff because obviously it's super interesting. And yeah, and then GPT happened and that's the story.
Yeah, and here we are. And then Armin, what were your roots?
So my roots are definitely not working poor but I, because my parents run an architectural office where they kind of adopted computers for CAD drawing, my first computer was like old computers that they recycled. So my first computer, even though I'm younger, was a 386.
Sorry for you.
And so basically none of the computers that I ever had were capable of playing computer games properly. Because one, they used Windows NT, which at the time didn't do anything, so you had to sort of like build a way through it, and like the only way that you could actually get them to run was, before it booted into either Windows 95 or like Windows 3.11, you could boot into DOS, like really old DOS games at the time when you could already get better stuff.
But because it was sort of this kind of thing, I started playing around with QuickBASIC a lot, with Turbo Pascal. I bought a bunch of books on that. And that was my roots of learning how these things work, and I wasn't ever really good at this but I found it really interesting. Like this idea of like
No, for sure, like I
We call it a Tiefstapler in German.
No, I swear to you, like when I started dabbling in this I just really sucked. But like over time, if you keep doing this, you get better.
And then in 2002 or 2003, I used to use Delphi a lot. It was like a visual version of Turbo Pascal.
Yep.
And in 2002 or 2003 someone also showed me, because I had this idea like I want to use Linux, and then Delphi didn't work on Linux, and then I found Python. And through that I started doing some Python programming, and Ubuntu just came out in 2004. And that was a venture-backed vehicle, but they created all these local communities, these Ubuntu associations. So together with a bunch of friends we started the German Ubuntu Foundation, not foundation, association. And we ran this online community called ubuntuusers for four or five years.
And because Ubuntu was popular, the community grew and then the scaling problems came. So that's how I got into web development. And then for building this I wanted to build a templating engine, a web library, all of this, and then eventually I bundled that together and made this Flask framework, which got very popular and even nowadays still is, I think, what clankers like to
Yes.
spit out.
That's hilarious.
But then I left that, and then in 2013, '14 or so, I worked on computer games for a couple of years in London, but then afterwards I went back to open source and I worked on Sentry for 10 years. And then left in April last year to try something new.
So both of you are originally from Austria. In fact, you right now live in Austria as well, right? You were doing games. You were working at Sentry. You also did games before. And then the third person who's not in the room but was on this podcast just before us, Peter Steinberger, also from Austria. Great that the two of you meet. Where did the three of you meet? Cuz I recently saw a bunch of photos, especially before OpenClaw and Pi started, of the three of you hanging out, experimenting, playing with AI.
I think the two of us met on the internet, right? On Reddit.
It depends, because I definitely met you once when I was at university.
So but you didn't recognize me at the time when I was listening.
I was already famous.
But yeah, we sort of abstractly met on the internet.
But eventually we met up in Vienna. We were screaming a lot at each other, but on the internet. But in a very cute kind of way, in a very non-confrontational kind of way, and even though we might not think alike in all areas of our lives, it was a cultured exchange I would say. So that was nice.
And Peter, like six degrees of Peter Steinberger basically. I was working at an office in my town, and the company that gave me free office space in exchange for being like a mentor to the CEO had some kind of business dealings with Peter's company, PSPDFKit. And eventually he came to the office in Graz. And I think that's where we met the first time. And then also the same year we met at the conference in Istanbul. We just hung out for an entire night and that's basically where it all started.
Nice. And then how did both of you go from being skeptical about AI when these tools came out? And again, both of you at that point, by 2022, had been doing a decade plus of building complex software in different domains. What was your first reaction to it, and then eventually how did you kind of come across to the side of like, well, this thing is actually really interesting?
So for me it was, I think in 2022, I think Copilot, GitHub Copilot came out before GPT.
Yes, in 2021.
And through my previous startup stuff, I was working with Nat Friedman and Miguel de Icaza from Xamarin because they acquired the company.
Xamarin?
Yeah, they acquired the company I talked about earlier, the Java compiler thing. I knew Nat Friedman from our early startup stuff, and he eventually moved to GitHub. And then he was in my DMs in 2022, I think, and asked if I wanted to have access to GitHub Copilot, the tap-tap-tap auto-complete thingy. I was like, "I don't really care. I don't think this is going anywhere." And he's like, "No, man, it's the future. Got to try it. It's the future." So, I tried it and it was absolutely horrible.
But yeah, after GPT came out, and especially when they started providing API access, I did a lot of projects just figuring out what works and what doesn't work, not necessarily in the coding space. But eventually, once they had tool calling, that's when they became very interesting. Or function calling, as OpenAI called it back then. But it took until, I would say, end of '24, October or so, for that to actually be useful. And that's where the coding agents also became kind of interesting.
And then 2025, the Claude Code team came out with Claude Code. And that introduced the agentic search. So, basically, just give the agent a way to plow through your file system and read all your files, and that was the whole difference, actually. Like all the things that came before, like Cursor with indexing and any AST-based stuff and all of that, it just went away. And I know that the CEO of Chroma is probably mad at me for saying this, but that was the difference, that it wasn't like a dense and sparse search thing that the agent could go through. It was just give it access to your files. That was it for me. That's where it clicked for me.
I think my path was kind of similar. Because I think Copilot came out quite a bit earlier. But I know that there was a program at GitHub that gave you early access to Copilot at the time. I think it was like this maintainers group or something that I was still in. I got the feeling for Copilot that this will actually be really interesting. But not in any way in which it is now, because I felt like, oh, I've been in open source for such a long time, and now that they're doing training on open source data, at the very least this will be controversial. I didn't think of it as being productive. I felt like, oh, this is going to be a controversial thing, that they would be training on open source data, and I remember I was trying to probe it really
What are those Flasks in there?
No, no, I was trying to probe it really adversarially. So one of the things that I probed on is, will it recall GPL code? And I remember at one point I got it to spit out the
Carmichael's inverse
Carmichael's inverse Carmichael function, which was very easy because it had a very specific name. So it was very easy to get it to recall. But I also found that you can sort of type in a certain way, then it would continue putting license text on top of it, which is completely wrong. So it came from an open source GPL drop originally, I think. And so it would have been GPL code if it had done that. But it actually attributed an MIT license from a random dude. And I did this like, oh, Mr. Copilot, that's the wrong thing. And that tweet at the time got really, really popular, and then people started sharing with me.
Because I was at the time not really exposed to how much actual AI progress was being made in those labs.
Yeah.
Like I didn't come from this AI space or ML space. So I learned about it at university and it's like, oh, there's the AI winter and then nothing happens. But through this tweet and some other things, all of a sudden I recognized that there was something there. Like there's actually CEOs in certain companies who are convinced this will take off, and that's how I started paying attention to it, and I was essentially trying all kinds of stuff with the API, like can you do bug fixing things. I got really interested in it, but it didn't at all feel like the world is going to change until Claude Code.
And you also changed your stance on the whole "oh my god, this is spitting out open source code it memorized."
So because my stick for many years now has been that I really want people to share stuff. Like I think human progress comes from building on top of each other, and I'm a huge supporter of the fact that in the US you basically take knowledge from one company to another company that then competes. Like I like this pirate kind of approach to sharing.
Yeah, spreading knowledge.
Yeah, and my optimal version is like copyright shouldn't exist in a way, or a very, very limited kind of version of this. I really didn't care that it spits out GPL code and doesn't attribute. So I was like, oh, maybe this will just completely destroy copyrights, and for me that was like, if that's the outcome of it, I'm fine with it. But it was an interesting kind of thing in the beginning, that it sort of creates this license violation. I want to see what chaos will emerge from it.
And so far I think mostly what has emerged from it is a strong belief now that the system in place for copyrights has some presumptions, assumptions in the US, about how it's supposed to work. And we're all kind of ignoring that right now because we want to create a mess first and then re-regulate it probably, because at least in theory a lot of the things that we're producing right now are probably, by historic readings of the copyright interpretation, actually not copyrightable.
Yeah, that's an interesting one, but speaking of which, let's jump into today. So an interesting thing that you did recently, we talked about it just before, is, as part of your new startup, building things on top of agents. And you talked to 30 different engineering teams saying, "Hey, how are you using agents inside of your company, inside of your team?" What did you learn, from large companies to startups?
I think that the bunch of learnings I tell you, unsurprising, is that whenever people had vacation, there was more time spent on trying these tools.
And just to be clear, you talked with folks at the likes of Meta, startups.
Yeah, so
Like a bunch of different people, right?
So a bunch of different people from different European dinosaurs, people like that.
At me?
[laughter]
Well, I mean the European dinosaur would be someone like Siemens. Yeah. Or I also talked to two companies which are sort of in a critical space. And what I mean when adoption happens when people have vacation is that when your CEO or your tech lead comes and says you got to use Cursor now, you got to use Claude Code now, actually you don't get it in a way. Because you need to actually spend some time and it's like a two to three week kind of thing until it really clicks on you. And so I always felt like with the people that I knew, I had a lot of free time, I left the company in April until October, I was like I can dive into this and I felt like, how does nobody get this?
It's like catnip for
It was crazy catnip. I didn't sleep much all of this. But what happens within the company seemingly is that when there was Thanksgiving, for the Europeans a lot of it was over summer, and then at Christmas a lot of people sort of, and they also get free credits during those times. And so more and more people get
I mean the AI companies often give you generous credits.
More people went into this, and especially after Christmas, I would guess in more than half the companies I talked to, after Christmas it really exploded. And it exploded in all the ways you would expect, where all of a sudden the quality drops. And it doesn't necessarily drop because people want to make worse code, but because it actually takes some effort to stay within this.
And we have seen this in the startup ecosystem already in the summer last year. If you pay attention to the YC startups, some of them have their stuff on GitHub for some period of time and you can look at it, and at the time there were plan MD files checked in and everything attributed to Claude. So that vibe coding kind of thing was for prototypes and whatever, and they built it out. It was already out there to see. But then gradually a small version of this has been code bases with a little bit of vibe slop on top.
And an interesting part of this was how engineering teams and companies are now responding to that. With all kinds of different findings, but a lot of it has been challenges to review PRs. They're getting larger and larger and they're becoming more psychologically taxing.
Engineers specifically are having a hard time keeping up with the longer PRs, that they're more frequent.
Yeah, and also a lot of the code in those PRs is how an engineer wouldn't do it, because as an engineer you sort of get a really bad feeling coming and sorting code because you think of your future self. And the agent really does not care.
I will retell this story over and over, but I worked on an Xbox One game at the time, right around the Xbox One launch. So that was a fixed date, it has to release on that date. So I worked on the Halo Master Chief Collection. And there was a game where you had a matchmaking component and you had to start this thing. And it was an all hands-on-deck kind of situation where people had to go in and unslop the human-made slop that was the matchmaker. And it was a system with way too many states. We called it an emergent state machine because it was like 16 bools on one massive thing, and in theory it had only six valid states, but in reality it was a geometric explosion of possible states.
And that's how shitty code feels. Where it really should only be a very clearly defined system, but now in reality they're like, "Oh, config doesn't load? Let's catch it and load the default config." So instead of actually failing, it now recovers. But now your code is way more complex than it should be because instead of failing properly, it is now recovering and entering these many more failure states. And that makes it much harder to work with this code because you can also not really ask the agent to refactor it, because it's like, "Oh yeah, this could be possible, so we need to maintain this invariant."
I think it's kind of even worse than what you described about your human-made complex system, because there are moments of brilliance in agents where they spit out perfectly fine simple code. Exactly the amount and type of code you needed for that specific thing. And you as the steering engineer looking at that are like, "Wow, this is amazing. I can just sit back and not care because it's obviously doing the thing." Like 2 minutes later you have another agent running in this window and it spits out the worst horrible garbage code. But you might not notice because now you've fallen into automation bias and think your agent is doing the job well.
Do you think this might be a bit of a human bias? Because typically, onboarding a new engineer, you have a new joiner, a new grad, you review their code, and if it's terrible code, you will review the next one thoroughly until they get to the point that, "Oh, they write the code that I do." And it typically takes you 6 months or a year or something like that. But then, you know, I can trust this person.
Yes, but you don't have anything like that with agents. Agents don't learn. You can put as much stuff in the agents and they will build a memory system, but that's not the same type of learning that a human does. Obviously, humans are fallible as well, no matter, but they have some capability of learning.
And retaining that learning, right?
Yes, and they also feel pain. I think that's one of the defining things about humans. It kind of ties back to what you said. Eventually, if the pain gets too big, you as a human are incentivized to fix the cause of your pain. And in the code base, the cause is usually terrible interfaces, terrible complexity that you want to get rid of because you can no longer maintain that system.
Isn't this why, just holding on to that, senior engineers are always in demand? Because the CEO sees a senior engineer as, they just get it done. But in reality, most senior engineers who are effective, they've had battle scars. They've been burned and they felt the pain. They saw what happened when they left the tech debt spiral. So they now make all these decisions that they know will help avoid it, and of course, through this, progress goes faster.
I personally think, and your mileage may vary, but a good engineer is an engineer that says no a lot, and "I don't need this" a lot.
Mhm.
Because that keeps complexity down. If you're using agents, the exact opposite happens. You say, "Yes, I want this and that. I want this and I want this and I want this," because I don't have to type it myself. I don't have to think about it. I just give the machine a prompt and it will spit out something that kind of looks like the thing I wanted. Good enough. And that's where all the problems start.
And one thing that I also think is that good engineering is all about knowing the trade-offs that you have to make. And sometimes the right solution is actually, if you were to sit at university and learn about it, you kind of learn that you shouldn't be doing this in a way. I think Kyle Henderson had this once where he said you do the dumbest solution first until it doesn't work anymore. Because the actual problem is there's so much stuff that you need to do that if you actually do the right solution, the correct solution, all of this, you're creating the kind of complexity that kills you at scale.
And the engineer learns that, but also if you don't have that battle scar, it's very hard for you to argue correctly, because it is this learning process that gives you the authority to then convince other engineers in the engineering org that you should be doing it this way. That is part of it, that you learn that, but the other thing is also that the agents now give you world knowledge access. And one of the other things that I learned through interviewing engineering teams now is that the senior person says no, knowing something, and then 48 hours later the junior comes by and says, I talked to the agent, and I already had this inkling, but now I have all the evidence of why we should be doing it this way. Because previously, you really didn't have that ready-made access to
Someone who can tell off a senior.
Yeah. And this creates other stresses now that previously, not every team has that.
Yeah, it's like people going to the doctor with a ChatGPT printout and saying, "This is what the machine said. You better do that."
Is it fair to say, based on what you're seeing and talking, we might face a thing where it's harder for experienced engineers to say no in spite of the product manager or a junior engineer saying
It's much worse, because the product manager now comes in and sends pull requests and automated shits them.
Yeah, that's another thing Hassan. Non-engineers participating in the engineering process is a thing now.
Ask Armin how that works.
[laughter]
Ask him how does it work?
How does it work, Armin?
Well, it's hard because on the one hand, it's well-intended, right? If someone who's
What is your experience? Is this your company, talking with other people?
So, first of all, we have a little bit of this in it. We're small, and so my co-founder sometimes sends a pull request on the website. I talked to people at half that at scale, where the marketing team all of a sudden does stuff on the website. And the sales team creates ever more elaborate sales demos that sort of land up on a GitHub org. And one of the funniest ones was where the sales demo built a feature that didn't exist, but nobody noticed. Right? So this is all new, right? Because previously none of that happened.
But don't you think it's empowering? Like if your
There's a good thing to it in tool.
If your entire org, if everybody in your org can participate in the creation of software in some form, right? Previously, people couldn't do that. You had a designer who could figure something out in Figma, but they might not be able to put it into a clickable dummy demo, whatever. You might have a PM who wants to try out a feature without wasting the time of an engineer. Now you can do that. The problem is that people are now so focused on "everybody can do everything now" that they forget that you still need a process to kind of guardrail all of that.
And the integration part is the hard thing. Peter gave this idea of the prompt request, and I'm actually really warming up to this idea. Once you've demonstrated it, I no longer need your code.
And just to recap, the prompt request was him saying that he doesn't like to get pull requests, and said he would rather see the prompt, because he will run the prompt, or he will tweak it, and it will generate it in the style that
For me, it's less about I want to see the prompt, as in what is it supposed to be doing? And now that we understand, because in many ways I think the interesting part is often you don't really fully know what you wanted to do in the first place, and so the act of creating clarifies what you really want to do. And so that part is highly valuable. Often the approach and the code that comes out of it is not what an engineer with sufficient seniority would have done. So it's not like I want your prompt so that I can reclank my clanker so that it does it slightly better. But more like, now that we know what we wanted to build, it's probably faster for me to start.
Yeah, and I also kind of disagree with Peter on "I just need your prompt." I actually value seeing a terrible implementation of something. If I get a pull request, and most of the pull requests we get on the Pi repository are made by agents without a lot of human touch, let's say, then I immediately know, okay, this is going to be garbage.
[laughter]
But it's valuable garbage, because someone has put in at least a minimum amount of thought instructing their agent to create this pull request. And I get to see how a shitty implementation of what they wanted to build looks like, and I don't need to waste my own time on trying that out. So somebody else tried it out already, the naive, dumb, agent-does-the-thing, does-the-mistakes version. And that saves me time. I'm not saying I like pull requests by agents, because they're terrible and they auto close them now. But they have value. It's not just a prompt. It's on an exponential, right?
[laughter]
The speed of everything
Eventually always, because thermodynamics, but I think we're going to find out way earlier than in previous cycles that this is a bad idea. That's good news.
But I think it's going to be interesting, and I don't know the answer to this, but I read this fascinating retelling of the British Industrial Revolution and how it changed the textile industry.
The Industrial Revolution, yeah.
Yeah, and so the general thesis of that article was, every time something at the head of the pipeline got optimized, it created an incentive downstream of the whole thing to create something, right? So in the beginning, if you can weave the thing faster, then eventually you need to have yarn that can be weaved at faster speeds, then eventually everything sort of turns the bottleneck all the way down. And ultimately, the biggest bottleneck in the entire thing turned out to be what I think is actually the next bottleneck we're hitting in engineering, which is: at one point you made a shirt, and if you didn't like the shirt, you went back to the person that made it and they fixed it up for you. And so the actual thing was, if the shirt is bad, nobody cares anymore who destroyed the shirt in the process, because you're just going to get a new one. Right? The responsibility actually went from anyone in this chain to the entire factory as a whole not having to carry responsibility anymore, because we've commoditized the whole thing so much that you don't have to do this.
And if you take the engineering approach of it, a pretty significant part of running a company and running a service is running it reliably. And so you have these postmortems on incidents to figure out what went wrong in the process.
And you go back and fix the shirt.
Yeah, and the thing is, we are all running on this idea that every engineer that is in this creation process ultimately carries some responsibility. And that we're going to that person, not to blame that person, but to figure out why did you do wrong here? And so if the machine now produces stuff at 10 times the speed, the responsibility thing does not scale in the same way, because a machine cannot forget to be responsible. And I don't actually know if there is a future where you can abstract away human failure so much in how we run engineering
that now the entire company no longer cares about who signed off on a pull request or something like that. That we automated in the same way, I think, as we are sort of automating t-shirt creation. I just don't see that, but
So, here's the thing. I think one thing we software engineers or IT people underestimate is just how freaking complex the world is. And how much human squishiness is in each little nook and cranny and corner, right? So, we're thinking, "Oh, we were now able to automate that thing. Now we can automate everything, like every bit of knowledge work." But we as software engineers are so bad at becoming domain experts that we don't see all the non-machine parts that go into a workflow. And we're running through the same fallacy here again. We're seeing models doing incredible things. I'm not disputing that. For me, this is like woah. Basically, all my research in the 2000s is now null and void because transformers can do all the things. But we are overextending that to everything like we always do in software. Like we did in EdTech. Yeah, we have tablets in classrooms now. Sure. And now it's solved. Education is solved because we have now computers.
Well, in fact, I've heard, I don't know which country it was, but they're now rolling back.
Yes, it's Sweden. They're
taking the tablets out from the classroom.
out, if you do some scientific investigations into the tactics and effects on pupils, if you just throw a bunch of tablets into a classroom, close it, and hope for the best, turns out the best is terrible. So, yeah. For me, I think the biggest takeaway in the past two to three years is the hype is terrible, because it dehumanizes everything. And I want to not be part of that circus.
Well, speaking of not wanting to be part of the circus, let's talk about Pi. Which is a very popular
clown nose.
And also minimalist coding agent. Can we start with the backstory of why you decided to build Pi at a time where there were already agent harnesses around, right?
Because they were suboptimal.
So, tell me more.
Yeah, sure. I mean, so I was a believer in Claude Code, just because they kind of created that whole genre through the invention of agentic search. I mean, invention. There were precursors to that and shoulders of giants and so on, but they were the first that actually packaged it up in a really compelling package. And at the time that fit my workflow really well. It was simpler, it was predictable, so on the LLM heuristic nature or stochastic nature of being kind of unpredictable, but everything around the LLM was kind of nice and tidy and easy to understand.
You were a happy user of Claude Code, right?
Happy. I was proselytizing it. But eventually the team started dogfooding and getting more and more tokens, I guess. And kind of increased velocity and team size, and with that came more features and much, much, much more bugs. And I personally like simple tools that are stable, that I can rely on even if they have non-deterministic parts, but all the deterministic parts should be as stable as possible. But that was just not the experience with Claude Code around summer 2025.
Mhm.
So I kind of soured on that real hard.
Was it bugs? Was it unexpected behavior? Like
So they take away your control of the context. They would inject stuff behind your back, which is bad. And then your workflows that used to work stopped working because there's now a system reminder that you don't even see in the UI that will modify the behavior of the model. They would also do this to the system prompt. I reverse engineered, I mean, I wouldn't call opening an obfuscated JavaScript file and unobfuscating it reverse engineering, coming from a more low-level background, but I reverse engineered Claude Code during the summer of 2025 and built a little service where I can track the progression or evolution of the system prompt and tool definitions in Claude Code. And it's like every release it was messing with stuff. cchistory.mariozechner.at if you want to see that. And yeah, that just messed with my workflows and I don't appreciate that. If I commit to a development tool, I want it to be a stable, reliable thing, like a hammer. I don't want my hammer to break in a different spot every day.
Yeah.
That's terrible. So, that's what happened with Claude. But, again, I'm not roasting the team. Some of them are really nice people I got to know on the internet. They are just dogfooding, and that's perfectly fine. We need somebody who goes the full velocity kind of way. But I don't want to work with a tool like that.
Yeah.
Because, again, get work done with
Sounds like the move fast and break things, the break things was
Yeah.
was not for you.
No. And then I looked into alternatives, and Amp and Droid came out around that time, I think. Pretty early in 2025. I don't remember.
Amp was earlier.
It was very early.
I think they sort of spun off from the same experience of taking, because I think Amp was around when Claude Code came out.
That's for sure.
Around that time, yeah.
In any case, I looked into those harnesses, and they were super good. They were just super expensive, as well. Because none of them could basically use what made Claude Code enticing on top of it being a cool tool: the subscription. And that works in an enterprise setting, where you're paying by token anyways. But it doesn't work for the small tinkerer in the garage. While I'm not a small tinkerer in the garage in a financial sense anymore, I kind of still relate to that community, and I would like to use my subscription with something. So, I looked into open source alternatives, and found OpenCode. But while that kind of vibes with my OSS roots, it too did stuff to the context I didn't appreciate behind my back. Pruning tool results after a certain amount of tool result token output, or asking an LSP server after every single edit the model makes if there is an error. Yes, there will be an error because the model isn't done yet with its work, so the code doesn't compile. So, the server will
So, like reaching out to LSP, the language server protocol server.
Yes. So, when you go into VS Code and you type some TypeScript, you have in the bottom some error diagnostics, and that comes from an LSP server for TypeScript. And OpenCode runs an LSP server on your behalf in the background and feeds the model with diagnostics from that server on every edit. We as programmers, how do we work, right? We go into one of our more files, we edit line after line after line and only then look at the errors that resulted from that. In OpenCode's case, or in other harnesses' cases that also support LSP, the model calls an edit tool to change lines. And it would then check the diagnostics after every edit call, and that's just not smart, because now you're confusing the model with you have an error, you have an error, you have an error, and the model is like, "Yeah, I know, I know, I'm not done yet, though." It's not great. Anyways, TLDR, OpenCode wasn't for me either. I also had to fork it to modify it, which I don't think should be necessary. So, then I just thought, "How hard can it be?" I built my own little thing.
And then your own little thing is pretty minimalistic. What does it use? What's the basics of Pi?
The basics of Pi are my own abstraction over all the LLM provider APIs, because I didn't like the Vercel AI SDK for various reasons. Armin kind of wrote a blog post eventually about that as well. It's obviously good to use. Lots of people use it. It just didn't fit my old man sense of abstraction.
Well, this is the beauty of software, and especially open source. You can always build your own.
Yeah. And now with agents you can even do it faster and produce terrible conflicts after. No, so I built an abstraction over that, then I built a little abstraction for a generalized agent loop with tool calling and streaming and a lot of that. I built a bespoke little TUI that doesn't flicker, or not a lot, and then I tied it all together into a coding agent that looks like Claude Code or Codex or whatever you have. And that's it. And the extensibility comes from the fact that this minimal core has so many hook points that you can basically hook into with a simple TypeScript module that gets loaded into the same Node process. And that allows you to do things like provide the LLM with custom tools. Do your own compaction implementation. Fully revamp the TUI itself. You can modify everything in the TUI. So if you have a special workflow
TUI, right?
Yes, exactly. If you want the TUI to behave differently for a specific workflow you have, like say you're a non-techy, you can change the TUI to become whatever you need as a non-techy. I have a couple of non-techy friends that did that, because they don't need to know how to build this. They can just ask Pi to build it. And Pi will modify itself.
Oh, so this is a thing, right? So you can ask Pi to modify itself because of the extension points, and it can write code that extends itself.
And it's trivial, but it's a big unlock.
Is this what you meant when you said that for OpenCode you needed to fork it to modify it? It doesn't have this.
It has a plugin system, but there's not a lot of extension points and it was very rigid. I think they changed it recently. I think it's much more open now. I haven't kept up with it, but it might be better now.
But so I guess Pi starts as this very minimalistic thing. As I understand, the tools it has are read, write
edit, bash. It's all you need.
That's it. And then you can actually start to make it your own. What are examples of what people would add?
Pi doesn't have MCP. People just ask Pi to build MCP support into Pi. Pi doesn't have a plan mode. Armin goes, and my plan mode must be fantastically spoken, super special.
I have a plan mode.
Yeah. But he has like five implementations of a plan mode, until he realized plan mode is entirely useless.
Other people just like messing with the UI and making it their own, like a different visual style of the editor box where you enter your prompts. Trivial stuff, more cosmetic stuff. Other people have retriggered it for a full-blown RL environment for open weights models, where they use Pi as the agent that's part of the RL execution environment. So, you can do anything, really.
What drew me to it, beyond actually using the library abstraction, was in fact the custom tools part. Because one moment for me was over Christmas, again, like many people. I had some time and I tried to build other things, and Peter was talking to me in November that he's vibing without looking at code, more or less. I don't know exactly how he said it, but it was like he can do this now. So, I was like, "I want to build a thing where I don't look at the code. I wanted it to not look like slop." So, I wanted a version of it where afterwards, even though I don't really look at the code, it should look like what I would have written. And I want to make a game. And so then I basically started the whole experience with just basic Pi, like, "Well, I want to build a game, but actually before we build a game, I want you to set up the code base in a way that you can validate the changes that you're making, but also I can see them." So, like a two-pronged kind of approach. I wanted to be in the loop, but also have the agent be able to validate itself. And what sort of emerged out of that was, well, first of all, it built itself some debugging tools into the game so I can make screenshots and run a simulation and sort of dump out state, read it again, but also Pi can show images in a UI, and I added a bunch of, like, I talked with the clanker to figure out what would be interesting things to do, but we ended up having all these screenshots I can tap through quickly in the UI. Pi also has this great feature where I can revert to an earlier state in the conversation and then I can branch within the conversation. So, we built a bunch of stuff around that. And because these sessions, especially with screenshots in them, became very token inefficient very quickly. It was actually one of the other things that it was rather quickly rather good at, was having a lot of screenshots in it.
Because OpenClaw people had a lot of screenshots in their chats, and OpenClaw is using Pi. So, we had to
But having this, it felt really magical for me to actually treat the problem as: I don't know what the right way of engineering here is, but very clearly part of it is I should be in the loop so we can figure out how to do that specifically for the problem at hand. And in turn, for web projects and computer games and some of the other things I tried, they're kind of different. But very many of them sort of come down to a similar thing, where the agent interacts now with my program and it should do it the most optimal way, and I want to interact with it in conjunction with it interacting with the program. And the entire experience should be as little confusing as possible to both me as a human and to the agent. And I found that very, very fascinating, just to see how that emerges. Where your tool all of a sudden, when you launch it in this program, looks and feels different than if you launch it in the other program.
I really like this point Armin made just a few seconds ago, that AI works best when the engineer stays in the loop and the system can actually validate what changed. And this is a great time to mention our season sponsor Sonar. AI can now generate code faster than you can verify it. Sonar, the makers of SonarQube, sees this leading to a serious gap in verification. With the rise of coding agents autonomously writing code, verification is no longer a nice-to-have. While the latest coding models are extremely intelligent, they also are error-prone and they don't fully understand your code base and your context or your objectives. This is why verification must be mandatory in agentic workflows.
SonarQube provides a zero-trust multi-layered approach to code verification that is consistent and repeatable. It analyzes semantic syntax, data flows, and architectural boundaries at agent speed, acting as a critical trust and verification layer before any code reaches production. Covering 40 plus languages and 7,500 issue types, SonarQube is the most comprehensive code verification platform available. And with easy integration via MCP, CLI, and hooks, it fits right into your existing AI toolchain. Let agents move fast and have SonarQube as the independent, multi-layered verification for safe, reliable, and auditable agentic development. Head to sonarsource.com/pragmatic to start verifying your agentic workflow today.
I'd also like to talk about our presenting sponsor, Statsig. Statsig built a unified platform that enables both experimentation and continuous shipping. Built-in experimentation means that every rollout automatically becomes a learning opportunity with proper statistical analysis, showing you exactly how features impact your metrics. Feature flags let you ship continuously with confidence. And because it's all in one platform with the same product data, teams across your organization can collaborate and make data-driven decisions. To learn more, head to statsig.com/pragmatic.
With this, let's get back to the episode and to the topic of general versus purpose-made tools.
Yeah, I mean, I spent a lot of my youth on construction sites to earn money, and you don't use a hammer for all your problems at a construction site. You have a screwdriver, you have your hammer, you have your drill, you have whatever. And I think in engineering, it's kind of the same. I'm not using
the same tool for every task I do as an engineer. So now, if I use an agent, I don't want a general agent for every task, per se. I want a specialized thing where I know the performance will be top-notch for that specific task because we built the harness in a way that the agent can be most effective at this task just because of the way the harness is constructed. And that's what I wanted to enable with Pi.
That said, I'm probably the person that has the least amount of modifications in Pi. I have like two extensions that I use and they're trivial. They're basically just if you see a URL that looks like a GitHub issue or pull request thing, pull down the details via the GitHub API and display me a small little widget on top of the editor that gives me the issue title, the author account and a link to the issue. That's basically all I do.
Well, but it might work for you as a minimalist.
Yeah, I mean that's how I work on the pi-mono repository because I might have two or three sessions open in which I process an issue or pull request. That way I remember what the session was about.
But it sounds like you also made your Pi for working on the pi-mono repo a specific one, and if you went back to building games, you'd probably have to have a... I never thought of the fact that you might want a different harness for a different task. I guess we just kind of assume that most developers work on your main thing at work. You might have a side project and just experiment, whatever, but this is funny. I wonder if this is a new thing, that we could never have custom tools for a project. That just sounds crazy, you know.
Here's my intuition. I think where we are going is software that modifies itself on behalf of the user's wishes and needs. And the agents can do that now if you give them enough rope to modify themselves. And I think with Pi that is my first foray into this kind of self-modifiable, malleable thing. Just for the coding agent sector, but I think this actually can be extended to all kinds of knowledge work.
Do you agree? For specific tasks within the broader set of knowledge work obviously. Dehumanization and so on, you know. But yeah, the next plan here is actually to have an alternative user interface to the TUI because the TUI is obviously limited and the best alternative stack is obviously the web because it works everywhere and can do anything. So once I have that built out, then it really becomes interesting because then you're not limited anymore to the line-based rendering of a terminal. Now you can do really, really interesting stuff. And so yeah, we'll see how that works out.
And one reason that I learned about Pi before I knew that it was this minimalist interface is how OpenClaw is using Pi. How did that come?
I remember hanging out and then reviewing each other's blog posts and just throwing ideas at each other. And in October I started building out Pi and Peter started building out Va Relay, his little WhatsApp assistant, so to speak.
Oh, that's how it started.
Yeah. And he was in search of an agentic core he could reuse or copy. I think it started out by him taking Pi and cloning it and calling it Tao and then modifying it, but eventually he got tired of having to maintain that. So he just said, "I'm going to use your stuff." That's how it ended up being. Pi wouldn't have compaction if it weren't for OpenClaw.
Oh.
I specifically built that because Peter was crying in chat, "I need compaction." Okay, you get compaction, but I'm going to tell all my users don't use compaction. It's bad for you.
Yeah, well that's I guess the beauty of building on top of one another's open software, right?
I mean, it has pros and cons, yes. I now get to enjoy all the OpenClaw instances that think bugs in OpenClaw are actually Pi bugs. So they autonomously send me a gazillion issues and pull requests without the users probably even knowing, and I get to deal with that in my open source, so that's a negative side effect.
Well, so you're really on the receiving end of this, I guess.
I mean, just like OpenClaw itself is, which is much more exposed to this problem. I mean, they have tens of thousands of issues now and there's no way they can get a good grip on that.
But how are you dealing with the fact that you now have OpenClaw, just AI autonomously opening things on your repo as a maintainer? And do you build tools to battle this and try to close them out, or...
Build...
for OpenClaw ones, which embeds issues and pull requests into a 3D space so I can see the clusters of similar things that agents would have sent to the repository, and then I can bulk select things and close them out in...
Oh, really? So you actually have a 3D visualization?
Yeah.
OpenClaw for context at... I think it's less crazy now, but end of December to I think mid-February. I mean it was exploding obviously, but this explosion almost directly translated to... I was on this repo refreshing pull requests and the number went up.
Yeah.
It...
We actually tried to contribute and help out Peter a little bit, but I immediately gave up.
I didn't know how to do anything useful there. I was looking at this and I was like, this is a type of software engineering I'm just not used to.
I would fix two things and spend an hour on them and then 5 minutes after I committed and pushed it, some clanker comes along and just reverts my fixes. And this is not how I work.
Okay, can we talk about the name clanker?
Oh, sure. So Clone Wars, Star Wars. I actually never watched it, but kids of friends of mine watched it a lot while we were visiting them, so I kind of through osmosis got the lore. And there is an army of robots and the Jedis would call them clankers, or people would call them clankers, because when they moved they clank, clank, clank. Yeah. That's the origin of that.
Yeah. So an AI, a droid...
Yeah, exactly. But coming back to how do you deal with the influx of agentic pull requests and issues? I just auto close every pull request. Human, agent, doesn't matter. What I do is if I haven't had contact with you previously, my GitHub workflow knows about this because it has your account name in a file in my Git repository. So if you're not in there and you send me a pull request, your pull request gets auto closed.
Mhm.
And then my little workflow will comment under your pull request that says, "Hey, thanks so much for contributing. I really appreciate it. Could you please open an issue in a human voice no longer than a screen's worth of text?" And if I like it, I type, "Looks good to me," and then that account name gets put into the file. And the next time they send a pull request, they pass. And it turns out agents don't see the comment my GitHub workflow posts underneath the pull requests. So, this is a great filter for filtering out agents and keeping the humans safe, more or less, from...
This is interesting. I wonder if this might be an unavoidable future where we need a way to separate: is this coming from a human with an intent, or an AI?
I don't necessarily care. If it were actually a good PR, then if it came from a machine it's actually fine-ish. I think what's interesting in projects like OpenClaw even more so is it accumulates pull requests where actually there was no intentionality behind it at all. And so the person that dispatched the machine didn't actually care that much about it.
Or didn't even know about it.
Or didn't even know about it. And I've done open source for many years and there was a big difference between someone that's sending a pull request or an issue and they're like, "Hey, please fix this," but actually didn't care enough to even reply to questions anymore. Like this is not uncommon. And then you don't actually have to fix it. But you have to close it out because maybe it's still useful input, but clearly that person wasn't caring enough.
And with the pull requests it's even worse now because they come in so quickly that many of them cannot be merged anyways without manual resolution of the conflict. And there's a lack of back pressure mechanism. Because even I as a human, if I see there's like 500 pull requests open, I'm like, I probably will not contribute to this thing now. Because at worst I will make it worse.
Yeah. And I think previously in open source you had the people who would just send issues and be very entitled and say you're the worst person on the planet if you don't fix my little issue, but that's fine. That can be handled. And pull requests were kind of special because it needed a human to invest quite a bit of time to produce them, and you don't have that anymore.
You just have people: oh, this should be easy. Agent, please do a thing, make no mistakes, send it to this repository. And that's just not going to happen. So basically what we need are bottlenecks. I don't necessarily need human verification or verification that you're human. I just need a bottleneck that allows me to process the amount of incoming things as a human, because in order for Pi to not deteriorate into a pile of garbage, I still believe that it needs me and other capable people reviewing at least the important code.
And for that I need bottlenecks because otherwise I can't deal with...
It's the second law of thermodynamics, right? Everything degrades towards chaos and you need to put extra energy in to keep it away from this outcome. And we don't see and feel the pain of the code base anymore if we stop looking at it. And people don't feel the pain, or they feel no restraint anymore.
The issues are also interesting because on the one hand it is something great about someone doing investigation and sending you a description of that. That can be good and can be bad, but they look very similar. It takes quite a bit of energy to tell apart a good and a bad AI-generated issue request. And unfortunately most of them are not great. But some of them are actually good, and it's also kind of weird. All of it is weird.
I really don't know what the future of open source is in many ways because a lot of open source really worked because people piled onto hard problems and so they congregated around it and said, now we need to have a good database. So we're going to put all this energy on building a good database. The value of open source came from: there's some hard problems and we're going to throw our energy together and we're trying to figure out how to solve it. And now it feels like open source is all about throwing stuff up.
What really grinded me, made me so mad, was people... particularly a lot of agentic engineering right now is building more stuff for agentic engineering. So it's like Ouroboros or whatever you call it. And I see this tweet and it's like, oh, I solved problem XYZ and here is my solution for it. And you click on this thing and it's 48 hours old. That person probably never used the thing that they built.
I would like to suggest to the viewership to look at Armin's GitHub account over the last year and what happened there.
Yeah, I built a lot of the stuff, but I don't then go on Twitter and say, hey, I solved the problem. Right? I have a ton of vibe slop on my GitHub account. And I wish I could mark it differently because maybe there's some utility in it, but unless you're going to actually have that code base still be there a year, a year and a half from now and someone is still using it, the utility of that is actually not validated in a way.
And there's so many markers and metrics you can look at now for GitHub that really demonstrate this explosive growth of it. But if you were to then maybe find some other number to see how many of the things that are being created actually turn into really fundamental pieces that can sustain open source communities, that can actually deliver this value that scales amazingly, we haven't actually created many vibe-engineered projects that have become that.
But I like how you mentioned energy and how open source always worked. If we just think pre-AI again, let's say Linux, the most successful or widely used open source project. It has both an energy and a structure, and people come in with intent that they want to add something. They have a process where it goes through. There's human trust at every level. There's a little pyramid and in the end each change request goes up one level, and in the end Linus does the cut. But there's a lot of energy, there's a lot of intent. There's...
There's a lot of humans. There's a lot of humans.
And it was always about human energy. And now we suddenly have this AI, which is just tokens right now. Who knows how much they're subsidized or not, or it's just machines doing it, and suddenly they create plausible things that look like human energy, and it's hard to differentiate, and suddenly it just throws this wrench.
I actually disagree. I don't think a lot has changed for open source.
Okay.
The volume has changed.
Yes, but that's just a number. As you said, the amount of actually useful and maintained projects probably has not changed a lot.
So you're saying that the ones that were there, they're still useful and maintained.
Not even the ones that were there. I mean, there's a specific rate of new open source projects that survive longer than 2 weeks. That's always been the case, right? So now we just have more projects that die after 2 days than before. But we still have the same amount of projects that will have long-term viability just because there are humans that actually care to maintain the thing over a long time, build a community of humans that support the entire thing, build an ecosystem around the entire open source project that makes...
You're not a believer in the small book.
No. I mean, good job Meta buying that up. Super useful. No, I think at the end of the day we're kind of freaking out when we don't actually need to, because apart from the fact that I personally cannot generate code faster than the speed of light, for me, building an open-source project, and that entails not just the code but the community around it, the spirit around it, the ecosystem around it, nothing changed.
What changed is mechanical parts. I need the bottlenecks to deal with the influx of exponentially growing agents, pull requests, whatever. GitHub itself is under immense pressure because now it's not just humans hammering their infra, it's now billions or millions of OpenClaw instances hammering their infra. Yeah. Everybody complains about GitHub going down. I actually think they're doing a pretty good job. That's a lot of traffic that's coming their way since basically Christmas, since basically OpenClaw.
So yeah, I would be a little bit more optimistic. We're just in the messing around and finding out stage at the moment, and everybody wants tokens to be a KPI just like lines of code used to be a KPI. We've seen this.
Speaking of things that don't change and messing around and finding out, you wrote a tweet, or you wrote somewhere, that your biggest enemy is complexity. It's also your agent's biggest enemy. Can we talk about that?
Very simple. If I have a 600 lines of code code base and my agent can at best be effective up to a context window size of around 200,000 tokens, how much of the code can the agent see? A third, right? Right. If you manage to get all the relevant code for a task into that context window, you're probably okay. Although that is a separate project and information retrieval problem, which is not solved and which
agentic search also doesn't solve. That is, are you sure that the agent finds all the relevant code it needs to find to fulfill a thing? That's also where all the garbage code comes from, because it doesn't see all the things it needs to see. In this case, let's assume the best case: information retrieval solved, everything fits into the context, agent does a good job. Okay.
That's not the reality we're living in, because now the agents spit out so much code that they themselves cannot possibly read it into their context on a new task anymore. You know what I mean?
>> Yep. They fill up their own context window.
>> Yeah, exactly. The complexity they add is their own worst enemy, because eventually the code base will be so big and so complicated and so interconnected that the agent has absolutely no way on a technical level to ingest all the context it needs to do the new task.
And I would like to point out that the agent has learned all of this garbage from the internet and from us, because on the internet there's all our old code. While there are some pearls, there's also a lot of swine. Because we have a gazillion GitHub projects from the olden days where we just tried out things. And because instances like Linux or any other really well-maintained and well-written open-source project are minuscule compared to all the rest of the garbage.
And a machine learning model will kind of converge towards, well, simplified, the mean, right? And what is the mean then? It's not the comparatively handful of excellently engineered projects. It's all the garbage on the internet. All the cargo culting, all the trend-of-the-day kind of stuff. And that's what we get when we let the agent do all the things for us.
>> Yeah. So, we have this problem of things getting more complex, which slows agents down, which will in fact impact quality, which we were just talking about. But Armin, now that you're building your own startup, you two are building your startup now. You're working with agents, right? And they will have these things. How are you dealing with generating code, building products, balancing quality, tech debt, complexity?
>> With that?
>> Badly. Look, I think that we're coping. We're not dealing.
>> I don't know if I wrote this in the blog. I definitely have it on my slides for the conference here. I enjoyed the time from April to about October immensely. Because it felt like I can do so much, but also there was no heightened expectation. The world had not yet gotten used to this idea that everything now has to also move at 10 times the speed.
And there was a moment in time where I felt like, we worked on this VibeTunnel thing in the beginning, and it felt so much fun, because I have time now to play with the kids and I just prompt a little bit on my phone, and it felt
>> VibeTunnel was where you could sit with your phone talking with your agent on the go, when it wasn't as easy.
>> Tunnel terminal, basically.
>> Yeah.
>> And it's not that we did much with it, but it had this happy vibe, and I know that I spent too much time on the computer, but I didn't feel any pressure. But now it's like we're collectively feeling like everything has to ship faster, it has to iterate faster. The baseline that we want to achieve in terms of fidelity and everything has to be higher. And so now it feels very stressful.
>> Even in your own startup, which is a small one, right?
>> You can be the most stoic person in the world and it's still going to get at you. I'm slowly learning to work with my own emotions on dealing with this, but I find it very, very hard, because I was used to things working a certain way, and I knew how I do some stuff, and then I fell a little bit too much into the trap of giving in to the machine and actually doing things in a way that I normally wouldn't have done things.
>> It's definitely something I gently regret.
>> Gently regret, yeah.
>> And so, quite frankly, the answer is, I feel like now, with a little bit of hindsight, I learned some things that I wish I would have learned probably in November.
>> Tell us.
>> Well, a lot of it is really the recognition that there's no back channel to me or to any other engineer. Under normal circumstances there was a back channel. There was this feeling of things not being quite right in the code base. Now the change is harder, and you sort of see the complexity getting higher, but if you rubber stamp it, then what's the back channel there? And so this mechanism, this back pressure, this friction in the code base, you don't feel when you work with the agent.
>> I think there's a way to kind of measure it. If I scan through my sessions on a project from start to current date, I think the frequency of curse words increases, because the agent starts messing up more, because it itself cannot deal with the complexity of the overall project. And I would actually be really interested in whether this is measurable, because I feel it in most of my projects now that I curse a lot more.
>> But you mentioned friction in the software. You didn't say tech debt. You didn't say complexity. What is this friction? Because I don't remember us talking about this pre-AI at all.
>> So I found this ironically kind of funny, and it's kind of sad. I will not name any names, but there was what I assumed was an incident related at least in part to agentic engineering at a company, where they shipped out a configuration change that ultimately resulted in a security issue. And look, things happen. But the link that I saw on this had the social preview of that company's tagline, and the tagline was "ship without friction."
And that really gave me pause, because as an engineer, we used to talk about how you've got to get rid of all the things in the way so that you feel happy shipping stuff. But there always were changes where you really wanted to think: do you want to drop the database? Do you want to merge this migration, which might take a table lock that could potentially take you down, right? There are these moments every once in a while where you are really supposed to think, and people created checklists, or people created mechanical gates where you would have to confirm something.
There are certain things that we used to put in, particularly if you run a SaaS company, to slow things down. And in some of the best engineering teams, in order to mature a service, you have to define an SLO, you have to define your expectations, and if your service is supposed to be critical, there's some other stuff that unlocks on the sort of tree of requirements that you have. And a lot of engineers feel like, oh, this is all bureaucracy, but the reality is, if you do this correctly, then it saves you time and it makes you happier. You're not waking up at 3:00 in the morning. All of this is useful.
>> It's like friction injected to deliberately slow things down. I guess the easiest example: in any decent-size company you have services tiered based on criticality. The highest-tier software now needs to have, let's say, two or three code reviews, or an approval from a director to do a configuration change, which again all slows things down, but we know this is on purpose. By adding this friction we want you to think: do I want to pursue this, given the friction of time invested or effort or having to justify things, etc.
>> It makes you think about whether you really want to add this to the code base if you know that the end effect will be that it has to go through this entire chain of code reviews. So it comes back to saying no to yourself to avoid the pain of going through that process.
>> And then taking all the pain when you know that you have the conviction, you have the backing, you have the confidence as well, right? So typically when it's a higher-friction thing, let's say a tier-one service or a highest-tier service where a director has to sign off, when you're a new joiner on the first day and you don't know the context, you probably know that that's a pretty large ask, and you'll probably socialize it, get buy-in from someone experienced who says, "Oh, this is the right thing." You'll go with them, right? Back to human dynamics, a lot of that.
>> The thing is, there's a very delicate balance in the whole thing, because you don't want the friction to be just an accident of having created bad developer experience, right? But some things look the same, but they're very deliberate, though they may not be sufficiently documented. But there's this feeling now that you get rid of all the friction so that the agent can be very autonomous, so that you can run many of them simultaneously. A lot of it comes from that. These things are actually rather slower. And the only real time saving that you get from it is parallelism.
And so somewhere there is the trap. I feel a little bit more experienced now in managing the trap, but I don't have the solution for that either. And I will not say that there is an example code base where I felt really, really great about the stuff that I built, except for pre-existing libraries from before agentic days, where I still feel a strong emotional attachment to them. I'm much more careful about doing them than any of the other code that we have, other than Pi, to which I don't have access.
>> There's still no write access. [laughter] But there's a lot of slop in Pi, but I try to avoid it in the bits and pieces where I know that it's important code. We have an HTML export functionality where it takes the current session and just spits out an HTML file that you can then host on GitHub and whatever. I have not looked at a single line of code for that function. I don't care if it's broken, if it looks right when it comes out.
But then there is the agent loop itself, or the extension loading mechanism and all of that stuff, and that's important. And the way I deal with ensuring, or at least trying to ensure, that it has high quality is I refactor mercilessly, because that pulls me into the code base. I need to understand what I want to change structurally, not just line by line and syntactically or whatever. I need to understand what's going on to do a good refactor. And doing that every now and then, like I'm doing at the moment, prompted by wanting to add a new feature that's currently not possible with the current architecture, being in the code is the one thing that keeps the code base quality high and the complexity low. But that's against the industry wisdom of burning as many tokens as possible, token maxing, basically.
>> Yeah, that's an interesting one happening. But you just recently wrote, on the same theme, a blog post called "We All Need to Slow the F Down". Can we rehash some of it, and what triggered you to just put it out there?
>> Okay. So, basically it's this: okay, your agent can now spit out 10 times more code a day than you can, but it also means it spits out 10 times more boo-boos, errors. Even if it has half your error rate, then okay, it's not 10 times more, it's five times more. And it's still more than you would spit out. So, the rate of deterioration in your code base has now increased. And now go dark factory. Now take 100 agents that do this to your code base. What's the end result of that? So, that's the first problem, right?
You need some way to review all of that code that now gets generated to fix all the boo-boos. But you can't as a human, because as a human you're used to spitting out 1.5K LOC a day, and that's about the limit that you can actually review well, right? Your agent spits out 10 times that. No chance you can review that. And not all of the code by the agent might be important, like the HTML export thing, right? But even if the agent spits out 3 to 5K a day, you have no way of reviewing that in any meaningful sense. And then if you do the armies, yeah.
>> And then the armies, this is interesting. So you call it the dark factory. The idea being that tens or hundreds or thousands of agents, you give them a spec, they go and break it up, they organize themselves, they elect their mayor, and all that jazz. They have the QA agent, you give them roles, you give them context, and then you give them enormous amounts of tokens and spend, and the idea, or the hope, is that your software will be done in
>> Oh, something will be done. Definitely something's going to be done. First your purse, and then your... No, yeah, sure. More power to the people that make that work. I can't make it work. And the reason I think I can't make it work is because I still care about the quality of my product. And I don't care if it's built by hand or by agent. I just want the quality to be good, both in terms of how easy it is to maintain it and add new stuff to it on the developer side, and on the user side.
All the companies claiming that all of the code is now written by agents: yes, we know. Quality is garbage. We feel it in our bones when we use your product. It's garbage. So I don't want that.
And yeah, basically I think people need to turn around and say, "Hey, what are we even doing here?" We have these wonderful machines now that can take away so much pain from us by doing the stuff we hate doing, and doing that really well. Why don't we start by giving ourselves some more free time to work on the interesting bits, and delegating the stuff we know they can do to them at large, across the entire organization? Find all the things that annoy the sh out of you, and have the agents automate that for you.
And then you suddenly have time to think about what we actually want to build. What do our new users need? And if we decide to build a thing, then we can pull in the agents again and say, "And we're going to polish the sh- out of that, because now we have the time and the means and the tools to do an excellent job." But that's not how we're working. We build an army of agents and install Beads and make a big spec that hopefully will result in something basic.
But here's the thing. We talked about where the agents learned their knowledge from, right? The internet. So, garbage to mediocre. Now, if you write a spec, what's the best possible spec you can have? The best possible spec is, well, you define exactly how it should work. You give it test cases
>> The best possible spec is the software itself.
>> Oh, I see what you mean.
>> So, yes.
>> Okay, you write a spec that's not the software itself. So, that means there are a lot of blanks that need filling in.
>> Yes.
>> Where do you think the agent is going to fill those blanks in from?
>> Well, most likely from its training data.
>> Yeah. And we already identified what the quality of that training data is, right? Garbage to mediocre.
>> Well, and even before AI, don't forget, Stack Overflow had a really big criticism, because there was this thing of, well, you Ctrl-C, Ctrl-V from Stack Overflow, and oftentimes the first answer was not correct in many cases. Regex for email was a good one. You Googled regex for email, the first page was Stack Overflow, everyone just copied the first solution, and I think underneath, answer number three said it missed a bunch of cases.
>> Yeah. But here's the thing though. I'm not saying agents or humans are better. They're clearly not. But agents also don't solve that problem. And if you then don't let just one agent that's
already 10 times more productive as you do the thing that it's better at and that you as a human are better at, but 100 of those, what do you think is the outcome?
Yeah.
It's just very simple math.
Let's talk about another controversial topic, MCP versus CLI. It's coming up and, you know, right now I'm hearing a lot of people really going for CLI is the future, and I think I'm sitting with two of them, but also MCPs are also really popular inside of large companies. Especially when you talk with a bunch of people working at large companies, it seems MCPs have found a real product market fit inside of larger enterprises.
Despite what people might think, I don't actually hate MCP quite as much as it seems.
Oh, wait, we have it on recording.
Yeah, no, we don't deal in absolutes, we're not Sith.
So my fundamental challenge with MCP is that, first of all, the spec is very complex, I think, but I think this is just generally how specs happen to be, so it's a bit of the core of its time. So there's an inherent complexity in it, but if you wait to say, like, what is it really doing at the end of the day, it's authentication and it's sort of invoking some stuff. And MCP, even theoretically there's structured responses, but MCP for the most part is run some stuff, put stuff back in the context and then work off it. So it fills your context very quickly.
And Cloudflare has this code mode MCP, which I think in principle I really like. I have an MCP for testing which is a JavaScript interpreter that gives me access to the Google API, and between an MCP like this and a skill there's not a huge difference, because the skill also needs to be in a system prompt so that it finds it. But the agents are just very, very, very good at running code. And MCP is not quite running code, it's basically RAG. It's like input in and do some stuff, and maybe some state transition that the model also doesn't see, but it is in that sense just... it's a hard problem to solve, but it does solve auth, it solves a whole bunch of things.
I want it to work. I just still don't get it to work like I wish it could work, and my suspicion is still the glue has to be code execution. But because MCP servers are yet largely not defined in a way that the model actually understands them, I haven't found ways to compose MCP tools reliably. I found ways to make the MCP itself be composable by having the MCP be one tool, run code, but I haven't found ways to then orchestrate larger ones. I want it to work. And I think it has found its niche and I don't think it's going to go away.
I think it's just a victim of its own success, really. When the whole thing started, I think it was in October 2024, it was more or less a solution to get external services into consumer-facing chat apps.
Yeah, connect your emails, connect your OneDrive, connect your whatever.
Pretty much, and then IDEs also took it over because it was convenient, and the Cursors, the Windsurfs.
Yeah. But I think the origin was basically the consumer side, not the developer side, and I think that's a totally great use case. I don't want my mom to have to mess around with code generation or whatever to invoke some API or call some API and so on. So perfectly fine use case.
And then the developer side also picked it up and thought, oh, this is a great way to provide tools to my LLM. Tools as in, in the system prompt somewhere there it is: if you want to call this tool, provide this JSON payload and you get this thing back, right? And that kind of felt right at the time, because if you read Anthropic's documentation, they would say our models can deal with about 30 to 40 tools in the context, and even that wasn't the case, like at 12, 20 they would just break down, but doesn't matter. But there was still like, yeah, this can work if you kind of keep it small and contained and very specific to a use case. And then people started building MCP servers that would just basically map an entire OpenAPI spec into a gazillion tools.
Yep.
And that's where it all fell apart. So that's the first problem: very bad MCP servers from big corporations that thought, "We need this now. What's the fastest thing we can build? I'll just push the OpenAPI spec of our APIs through this thing and make it an MCP server." That's garbage. The second problem is that it's inherently non-composable. If you want to combine the MCP tool outputs of two different servers, they need to go through the context. The model itself needs to do the data transformation, the composition of multiple pieces of data fetched through an...
Compare this with a CLI, it's a pipe, right?
Exactly. The model only sees the end result and it is super free in how it massages that data. And that's also the idea behind code mode. Basically, it's a hack. It's basically, okay, we now have MCP, we know it doesn't work for this specific use case where we have multiple sources of truth, data, and you want to combine them but don't pull that through the context. So let's build code mode. And code mode is basically: we take all the MCP servers, we expose them as functions in TypeScript, and then the model can actually just write some code that calls the MCP servers and then does the composition in the code. It's like, how many indirections do we want here? We can just let the model write the code. We don't need the MCP server.
And then the third part is, David from Sentry is a big proponent of MCP because it's the auth thing. And honestly, that's again for me super valid, but the model itself kind of doesn't make sense anymore.
I think there is a world for MCP too, which is ironically maybe based more on... I saw there's a company called Stainless which basically generates SDKs out of OpenAPI specs. And I'm really warming up to the idea of, like, maybe it is an MCP that's entirely based on auth plus, like...
Libraries.
Libraries, or like directly HTTP requests against OpenAPI specs, because you compose it together there. And I think one of the things that's also kind of underappreciated, and you see it if you see Pi do its stuff because it's kind of transparent about the tool calls that it does, it's kind of magical at times how creative agents get at large outputs. Like for instance Pi, when it runs a program in bash and it produces too many lines of output, it actually only reads, I don't know what the cutoff is, but it reads the first couple and says, like, oh, if you want the rest of the file, it's 20 megabytes large and it's in this file. And then the agent is like, oh, 20 megabytes, that's too much, I'm going to grep on the file. Right? And they get really ingenious in how they're interacting with it, and an MCP takes that away. The question is, how would you define MCP in a way where it wouldn't take that away? Where it still has all of that magic and capability. And I don't really know the answer, because I think it's hard, but auth needs solving and composability needs solving, and I think there's a bright future for that kind of stuff.
And also, like what Mario said, if coding agents wouldn't have become so popular, then the idea of code generation, code running, for non-code-related problems probably wouldn't have taken off quite as much either. But the most capable personal agents, OpenClaw being a good example of it, are just coding agents hidden from you. And then naturally some random person who is not a programmer is going to say, how am I going to do this? And the model doesn't say, install this MCP; the model says, okay, I can write a Python script that does it. And so you naturally have, in sort of the crazy space, the adoption of more code execution, and in the compliant enterprise space you don't have that. There's a different path that...
And personally, I don't think that models are going anywhere else other than code generation going forward for any kind of agentic task. I think that's mostly a function of there being a lot of training data for code generation, and code generation being a very easy means to control computers. So I don't see a different paradigm there coming out of the model labs anytime soon. So I think, taking that as the assumption of where the future is going, we just need to figure out how to make code generation work within an enterprise setting, with all of the other enterprise things that entails.
So let's do a fun one, trying to predict a year out, which is hard. But then in 2027, knowing some of these basics, just again from first principles, where do you think these coding agents might be and the software engineering workflow might be? This is just speculation again. We know we cannot predict the future, but where do you think there'll be a lot of focus in the coming year, where we might, in an optimistic case, see some results in tools and how we work, and what's working, what's not working?
I have no idea. I honestly have no idea. I could make up something that's probably not going to happen. I think the self-modifiability thing is obviously something I believe in. I think we will see more of that.
Self-mutable software.
Yeah. Yeah. Including the tools themselves with which we build the software. And I think that will expand not only to the tech sector but also to non-tech applications of agentic tools.
So, is it dog years, which is times seven? Is that how it works? So that's basically the model I have right now of how this stuff works. When you ask me what's going to be in a year, it's like seven years. Right? And to me that makes it incredibly hard to have any sort of predictions about the future, because it's still not one year. Maybe now it's one year from people starting to use Claude Code, but it feels like it is much, much longer. Much more time behind. More time has passed.
And I think right now the closest that I can imagine is going to be: we know that code execution and code generation and this harnessing around it, this is going to be it, because reinforcement learning gets more of that data. And my strong hypothesis is that as more and more people are starting to wake up to this, that you can do interesting things with agents, there will be a societal recognition also of how much more dependent you are on basically two companies. And I think we'll have a conversation about that part. We should have a conversation about that part, particularly as Europeans, because we don't really have these labs over here. And so I hope we have that conversation. But my best guess is that we'll wake up to the fact that we are now... I mean, engineering teams are already now telling me that they have code bases that they don't think they could maintain anymore without a machine. My guess is that one of those companies will be public, and...
And all of a sudden...
And it will be expensive, and I think that might actually dominate, or at least become a conversation that's much bigger than the question of are you using Pi or using Claude Code or something like this.
I also see, we've seen this with, was it Missus? The new Claude model? Oh no, Spot, the new GPT model. They will only give this to select partners. So now we are seeing a split in who can get the best intelligence.
Yep.
Or the perceived best intelligence.
It'll be interesting dynamics. So both of you are working on popular AI tools. You're building a startup where of course you're using AI, and it's also around agents. How do you both keep up to date?
I've just seen things, and it's not as easy to get me on a hype train as it used to be. But that comes with age. It's definitely easier not being in San Francisco, because I think that would just drive me crazy. I hear so many things from my peers over there and I'm just like, yeah, I'm not going to go to San Francisco, thank you.
So having a peaceful environment around you where it's not all about tech might be helpful.
It helps having a kid.
Yeah.
It helps just going outside, climbing trees, going ice skating, and then looking back at what you did just half an hour ago and being like, why would I do this? It's just stupid.
I'm, to the detriment of maybe people that are trying to stay in contact with me, I got very good at not muting notifications, not reading emails, and that has in part become necessary, I think, over the last year or so. But it actually turns out that the passage of time sometimes clarifies stuff a lot, because if it's really necessary, it's going to reach you again. I have an unhealthy Twitter addiction, which I'm not particularly proud of. But in terms of a source of interesting things, that is still a thing. But I try to now sort of consume it in the form of: if it's really, really important, it will stay in the discourse for quite a while, and I just wait it out. And if it's there two, three weeks after it originally happened, then there's probably something to it. And I don't need the three-week head start necessarily.
But honestly, it's really hard. It is really hard to deal with this, because there's a genuine excitement in it. And I feel like my more than 20 years of experience in the space of software engineering tells me a lot of stuff. But at the same time, it hits you in certain ways where you felt like there would be grounding and there would be something to build on and a strong foundation, and now it feels like, well, seemingly everybody else doesn't care about that foundation anymore. So maybe you don't need the foundation, and for quite a while it sort of works, and that is sort of weird, and I don't know.
I kind of feel like since we've been funemployed in 2025 when all this started, we had a head start. I see all the excitement the two of us and Peter had in April last year...
As a wait.
No, no, but nobody else at the time kind of shared that excitement that much. And then the Christmas break came and now everybody else has that excitement that we had in April, right? So now they're learning the ropes. Now they are catnapping themselves into immeasurable amounts of lost sleep and terrible code bases, and I think it will self-correct because it's not sustainable.
Yeah, we did see this as well. I did a deep dive in The Pragmatic Engineer in early March, when a lot of people who were very excited in January about AI started to use the new models, what they can do, and they went all in at work or on side projects. In about two months' time a lot of them were like, hang on. It introduced all this complexity. It has these things. I'm not going as fast as I thought I would be, etc. So I guess there's just a natural thing with anything new, right? A job, anything. You have a honeymoon period where you've got the blinders on, which you should, by the way, and then you start to realize and maybe overcorrect. But there's a natural thing where, in general, it just takes time to see the outcome of your decisions.
So I'm not worried about all the dark factory and all the "software is dead" and "SaaS is dead" and all that. I truly believe this is just part of the hype machine that will self-correct.
Yeah. As closing, what's a book that you would recommend and why?
Code by Petzold.
Classic.
I just love it. It's just such a great read. It's also for non-techies, and it's the first thing I recommend if anybody asks me, "What's your job?" I'm pointing at that, and it's like, it has much less to do with computers than you think.
And I read recently Breakneck, which I unfortunately forgot the author of. It sort of goes a little bit into an exploration of how China works, or maybe how Europe and the US are different, and I found it at least thought-provoking.
Well, Mario and Armin, thanks a lot for this conversation. It was great to have it in person. Thanks for having us.
Thank you.
>> This was a really fun conversation. Thanks to Mario and Armin. The idea of self-modifiable software really grew on me. Mario said how Pi doesn't have MCP support, plan mode, and many other features that devs would want from it, but you can build it into its own code. So far it's working. Pi is popular because it modifies itself. I wonder if and when this concept of self-modifying software thanks to AI will spread outside of just the dev tool.
I also liked how we talked about the observation that agents don't feel pain, but humans do. When a code base gets too complex, the human engineer feels the issues this creates. And this tech debt is what pushes refactors and rewrites. But agents simply do not do this. They just keep adding to the complexity. And if a code base where devs regularly feel the pain of the code base and do something about it, the quality will probably be also better.
And finally, the MCP versus the CLI discussion, this was a good one. MCP is more about offering tools for AI through context, and CLIs allow piping one tool after the other. Both Mario and Armin are more the fans of the CLI, but in all fairness, MCP has its use cases, for example, inside larger companies. The right tool for the right job.
Do check out the show notes below for related Pragmatic Engineer deep dives that go even deeper into related topics. If you've enjoyed the podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating for the show. Thanks, and see you in the next one.
Article published
