Peter Steinberger on Shipping Code He Doesn't Read: From PSPDFKit to Clawdbot
The Pragmatic EngineerPeter Steinberger built PSPDFKit, a PDF framework that, by the host's account, runs on more than a billion devices. He then burned out, sold his shares, and stayed away from programming for about three years. When he returned in 2025, he started working almost entirely through AI coding agents. In this conversation he describes merging hundreds of commits in a day on Clawdbot (since renamed OpenClaw), his personal-assistant project, and says he no longer reads most of the code he ships. His argument is that this is not recklessness. He thinks the engineer's job has moved toward system design, taste, and building feedback loops that let agents check their own work.
From rural Austria to a copy-protected DOS game
Steinberger grew up in rural Austria and describes himself as an introvert. His family regularly hosted summer guests, and one of them was a computer enthusiast. Around age 14, Steinberger begged his mother for a computer and started tinkering. The earliest project he remembers was taking an old DOS game from his school, writing a copy protection for the floppy disk, and selling it. The protection added about two minutes of loading time. He compares building software to playing games and says that right now it "feels better than Factorio."
He never met his father and grew up poor, so he paid for his own studies. A job in Vienna that was meant to last one month, bridging military service and university, turned into about five years. On his first day he was handed a thick book on Microsoft MFC. He quietly used .NET instead and only told the company months later, when it was too late to change course. He says he did this kind of silent modernization several times there.
An app born out of a lost message
At university a friend showed him an iPhone. He held it for about a minute and bought one. The moment that turned him into an iOS developer came later. He was on the subway, typing a long and somewhat emotional message in a gay dating site's web interface on iPhone OS 2. The train entered a tunnel, the site's JavaScript disabled the send button and showed an error, and he could not copy the text, take a screenshot, or even scroll. The message was lost.
He went home angry and downloaded Xcode. The site had no API, so he parsed its HTML with regular expressions, which he admits is "totally not something you should do." He built on a stack of beta technology: the iPhone OS 3 beta, Core Data in beta, and a hacked GCC that backported blocks. The company never answered his email about it, so he published the app himself at five dollars and made around $10,000 in the first month. Apple's payments went into his grandfather's bank account, and his grandfather eventually called to ask about a large, strange payment from Apple.
When he told his employer he wanted to pursue apps, his boss mocked him and called it a fad. Steinberger says this left him with a chip on his shoulder and a resolve to one day run a company worth more than theirs, which he says took eight years. The app ended abruptly. At 3 a.m. at a party, he got a call from someone at Apple saying users had reported pictures in the app, and the app was finished.
How a magazine viewer became PSPDFKit
After quitting, Steinberger did freelance work. At a San Francisco bar during WWDC he was introduced to someone as one of Austria's best iOS developers, which led to a US job offer. Around the same time, a company asked him to fix a crashing iPad magazine app built by a contractor in Eastern Europe. He calls it the worst code he had ever seen: a single Objective-C file of thousands of lines that used windows as tabs, "a house of cards" where touching one thing broke another. He offered to rewrite it in a month, a job he says the original developers had spent half a year on. It took him two.
He found real technical problems in the domain. A C call to render a PDF page might need 30 MB on a device with 64 MB total, so careless background work got the app killed by the OS. He also fixated on details such as how pages animated during rotation.
When a friend struggling with his own magazine app asked for code, Steinberger extracted the PDF component, made sure the original client was fine with it, and sold it to him. He then put up a site in an afternoon, built from a WordPress template hacked to run on GitHub Pages. Buyers got a Dropbox link to a zip of the source. Three people bought it in the first week for about $200 each, and about ten more wrote to complain about missing features. The complaints hooked him. Text selection sounded easy, and three months later he had learned that it is very hard in PDF. He says he knows more about PDF "than any sane human person ever should."
He kept raising prices as features grew. By the time he started a job at a San Francisco startup, the side project already earned more than his salary there. He worked both jobs, slept less, and after about three months his manager asked whether he was okay and gave him a choice: the job or the project. Because of his visa, he had one week to decide. He chose the project.
Polish, developer marketing, and writing as a strategy
Steinberger says money never drove him. What drove him was making things other people find amazing. His aim was to build the component "as if Apple would have built it," with care and small delights. Competitors had more features and had been around longer, but he believes his product won because developers tried the options and his felt best. In his view software is mostly about how it feels.
He returned to Vienna, went all in, and brought in freelancers, which he now thinks he did far too late. He spent about 13 years on the product and kept its awkward name, which he says he chose in about five minutes. In Objective-C it made sense as a namespace prefix.
His marketing targeted developers rather than executives. Management makes the purchasing decision, but he reasoned that developers inside a company would lobby for the product if they liked it. PSPDFKit did no cold outreach. It relied on good software, technical blog posts, and conference talks, on the theory that visibly competent, caring builders reflect well on a product. He also personally answered support tickets and replied to the newest tickets first. His reasoning was that a reply within five minutes feels magical, while waiting one day or two makes little difference.
The team was remote-first and later hybrid. It grew to about 70 people by the time he sold his shares and is now nearly 200, he says. He required everyone to spend one full day a month writing a blog post. Sometimes colleagues worried about giving away secret sauce, but he mostly ignored that. He enjoyed the attention and the chance to inspire people. He also says that writing forces deeper understanding and that the posts became company documentation he himself went back to a year later.
Hard problems, WebAssembly, and custom pricing
Asked what makes PDF hard, Steinberger gave one example. He had designed the link model assuming a document might contain a few hundred links. Then a well-paying customer reported four-minute load times. The file was a 50,000-page Bible from Canada with more than 100 links per page. His assumptions were off by a factor of about a thousand. With a mature public API, the internals had to become lazy without breaking anyone's code or exposing what was loaded eagerly versus lazily. He says he spent about two months redesigning it until load time was nearly instant.
The company later replaced Apple's renderer, which he calls "quite buggy," with a large C++ renderer used across all platforms. PSPDFKit was also one of the first PDF frameworks to run in WebAssembly. He describes publishing a benchmark early in WebAssembly's life that Google, Microsoft, and Apple adopted. As a result, those companies were effectively working to make his renderer faster.
On enterprise pricing, he explained why vendor sites often say "contact us." The same product delivers very different value to a freelancer and to a Fortune 500 company. A low price looks suspicious to large buyers, whose procurement won't start a process over $500, and a high price excludes small customers. As unfair as it looks, he calls custom pricing the fairest option for that kind of product. He sorts software along two axes, easy versus hard and interesting versus uninteresting. PSPDFKit sat in the hard, uninteresting quadrant, which he considers a good place to be. Developers would rather not build it themselves, and the problems never run out.
Burnout and three years away
The early years were the most fun for him. Growth brought red tape, "gardening" the product instead of wild hacks, and more people conflicts. He worked most weekends and describes the CEO role as being the "waste bin" for everything that others couldn't handle. It was lonely because he could not be openly negative. One example: his co-founder called at 5 a.m. one weekend to say a large airline's planes were grounded because PSPDFKit was crashing. Steinberger disassembled the airline's app and showed that they had modified PSPDFKit's source in a way that triggered a license-key fallback. He calls it an "if they sue, company's gone" moment.
He believes burnout comes less from hours than from working on something you no longer believe in, or from constant conflict. The management team fought a lot. He also thinks he made a mistake in trying to run the company too democratically.
After selling his shares he needed a long time to decompress. There were months when he didn't turn on his computer. Having so large an exit so early "messed with my mind quite a bit," and he calls those hard years.
Coming back: Claude Code as a slot machine
In April 2025 he returned to an old idea, a Twitter analytics tool he had once started in Swift and SwiftUI and now wanted to build for the web. Web development had always been his colleague Martin's area at PSPDFKit, so he didn't know React well enough to know what a "prop" was. He calls this a common trap: the better you are in one stack, the more it hurts to feel like an idiot in another.
So he tried the AI tools others were dismissing. He credits his three-year absence with sparing him the period when many developers tried early AI tools and concluded they were bad. His first real experiment was crude. He converted a messy repository into a 1.3 MB markdown file, had Gemini in Google's AI Studio write a roughly 400-line spec, fed the spec to Claude Code, and kept typing "continue." Eventually the agent declared the project 100% production-ready, and it crashed on startup. He then added a browser MCP so the agent could see what it had built. After a few more hours of looping he had a working Twitter login page. The result was not great, but it was his "mind-blowing moment." He could see where things were heading.
For months afterward he had trouble sleeping. He compares the pull to a casino: you prompt, and sometimes you get junk and sometimes something that blows your mind. He pulled friends in, naming Armin and Mario, who soon were also up at 5 a.m. He calls the group "the black eye club" and started a London meetup called Claude Code Anonymous.
"I ship code I don't read"
The host pushed back. PSPDFKit's polish seemed to come from obsessive care about code hygiene, so how does that fit with not reading code? Steinberger says he now regards much of his old bikeshedding over spacing and naming as pointless. The code still has to work and be fast and secure. But most apps, in his description, just massage data from one shape to another: API to parser to database to HTML and back. The genuinely hard part, he jokes, was solved by Postgres decades ago.
He still cares about structure. The day of the recording he wanted to land a 15,000-line PR that moved Clawdbot to a plugin architecture, and he was excited about that design. He did not read all of it, because much of it was "boring plumbing." For him, what matters is system architecture, not every line.
He also rejects the term "vibe coding." He calls what he does "agentic engineering," and jokes that vibe coding "starts at 3 a.m."
The workflow: Codex, conversation, and parallel agents
Steinberger went through Claude Code, a stint with Cursor, some Gemini 2.5, and Opus 4. He says the inflection point came in the summer, when models got good enough to build software without writing code by hand. What won him over was OpenAI's GPT 5.2 with Codex, which he calls underrated: "pretty much every prompt I type gives me the result I want." On Clawdbot he runs five to ten agents in parallel.
He respects Claude Code, which he calls a category-defining product and still uses almost daily, especially for general computer work. For coding in complex applications, though, he finds Codex much better. His reason is that Codex reads far more before acting. Claude, in his description, reads three files and confidently starts writing, so you must push it to look at more of the codebase. Codex may read silently for ten minutes. With a single terminal that would be unbearable, he concedes, but with many sessions it doesn't matter.
He says he rarely just issues orders. Each session starts with a model that knows nothing about the product, so he has a conversation with it. He asks what options exist, whether it considered some feature, and where documentation should go. He doesn't use plan mode. Models are "trigger hungry," but phrases like "let's discuss" or "give me options" keep them from building until he says "build." He describes himself as designing the system, with a mental picture of its shapes but not line-by-line knowledge, which Codex supplies. The host compared this to the old capital-A Architect role. Steinberger prefers the word "builder."
He feels friction when prompting, much as others say they feel it when writing code by hand. He watches the output scroll, notices how long a task takes, whether the agent pushes back, and whether the result looks messy. If something takes much longer than expected, he knows he framed it wrong.
The work is mentally more taxing than coding by hand, he says. He is managing five to ten "employees" instead of one. He plans a feature carefully, starts it on a task that might take Codex 40 minutes to an hour, and moves on to the next one while it runs. Usually there is one main project plus satellite projects that need five minutes of attention per half hour. He compares it to StarCraft, with a main base and side bases supplying resources. He hopes this degree of parallelism is a transitional problem that faster models will reduce. For now, he needs to parallelize heavily to stay in flow.
Closing the loop
Steinberger calls the core technique "closing the loop": the agent must be able to test and debug its own work. In his view this is why current models are strong at coding and often mediocre at creative writing. Code can be compiled, linted, executed, and checked.
He gave two examples. A Mac app couldn't find a remote gateway, while the same logic in TypeScript could. Debugging a Mac app by hand means building, launching, and inspecting it, which is slow. So he had Codex build a debug-only CLI that exercised the same code paths, and then let it iterate. It ran for about an hour and reported race conditions and a misconfiguration. The explanation sounded sensible, and he didn't need to look at the code. In the second case, Google's Antigravity handled tool calls in a quirky format that kept breaking his filtering. He eventually had Codex design live tests that start a Docker container, install the whole system, run the agent loop against real API keys from several providers, including Anthropic and GLM, and have the model create an image and then read it back. It took a long time, but the agent fixed the ordering and tool-calling inconsistencies itself.
He even structures web projects so the core runs from a CLI, because browser-based loops are too slow. He argues that agentic coding makes him a better engineer, because designing for verifiability pushes him toward better architecture. He says he never liked writing tests or documentation and often only pretended to. Now both are part of every feature. He explains the trade-offs and has the model write the docs, beginner-friendly first and technical detail later. He calls his latest project the best-documented one he has ever had, without writing a line of it himself.
Why experienced developers struggle
Steinberger sees two types of developers. Those who care most about the outcome and product feel tend to thrive with AI. Those who love solving hard algorithmic problems often struggle, reject AI, or become sad, because that is exactly the work AI now does. He says he learned more about software architecture this year than in the previous five, but that you have to know what to ask. In his Twitter tool, usage made things slow in a way that was hard to reproduce. The cause was a Postgres trigger in a single file with no obvious connection to the rest of the code, which the model couldn't trace. It was found only when he asked whether there were any side effects for a specific operation.
Management experience helped him. As a CEO he had learned he couldn't breathe down everyone's neck, and that imperfect code that moves toward the goal can be improved later. Working with Claude Code felt like managing imperfect, sometimes silly, sometimes brilliant engineers, "a lot like being the boss again."
He criticized a blog post by a developer he deeply respects. As he understood it, that developer sent a single prompt to several models through a web interface, including an open-weights model he considers too weak for coding, ran the output, saw it fail to compile, and concluded the technology wasn't good. Steinberger's response is that no human writes bug-free code on the first try either. Complaints about outdated APIs stemmed from not specifying a macOS version, and training data contains more old code than new. He compares a guitarist who tries the piano briefly and gives up. He himself screamed at Claude Code at 3 a.m. many times before learning "the language of the machine." Early on in Clawdbot, for instance, his agent would cherry-pick from PRs and close them, until he learned how it interpreted his wording.
Rebuilding PSPDFKit today, and why he distrusts spec-driven orchestration
Asked how PSPDFKit would look if built today, Steinberger said he could easily run it with 30% of the staff. The hard part would be finding very senior engineers who understand what they build and are comfortable delegating, knowing which parts matter and which can be vibed. He sees few such people and a lot of noise on social media.
He is skeptical of elaborate orchestration setups: the "Ralph Wiggum" loop, which he considers a workaround for Opus limitations that Codex doesn't need; ticketing systems where agents create and consume tickets and email each other; and approaches like Gas Town, where a spec is written up front and the system builds itself. He calls this the waterfall model, which the industry already learned doesn't work. He allows that it may suit some people or long lists of independent tasks. He asks, "How can you even know what you want to build before you built it?"
His own process is exploratory. He sometimes deliberately under-prompts so the agent produces something that sparks ideas. Maybe 80% is wrong, but a couple of ideas are new. He has to click on a feature and feel it, because models lack taste. He compares building to sculpting a statue out of marble and to climbing a mountain by circling around it. He still plans, but less than before, because trying things and throwing them away is now cheap. When Clawdbot moved from one agent to many, and from WhatsApp only to many providers, the change touched the whole application. He estimates it took Codex about three hours where it would have taken him about two weeks.
Clawdbot: from WhatsApp relay to personal assistant
Since spring he had wanted a hyper-personal assistant: not one that sends a morning task list, but one that asks how a meeting went, notices you haven't texted a friend who is in town, or asks why you always seem sad after seeing a certain person. He mentions the film Her and even registered a company whose name means "the loving machine." When he tried in the summer, models weren't quite ready. He assumes all the big companies are working on such assistants and that compute is the current bottleneck. His angle is an assistant that runs on your own computer, with data that stays yours. He notes that many friends already use chatbots as therapists, and he thinks that works well.
The project started small, as "WhatsApp Relay," a way to trigger things on his computer from WhatsApp. On a trip to Morocco for a friend's birthday, he used it all day. It guided him through the city, made jokes, and texted friends on his behalf. Its resourcefulness amazed him. He sent a voice message, a feature he had never built. About 30 seconds later it replied. It explained that it had inspected the file header, found Opus audio, converted it with FFmpeg, looked for Whisper locally, didn't find it, found his OpenAI key, and used curl to get a transcription. That was Opus 4.5.
He added a heartbeat, periodically prompting the model to do something proactive, which he admits is insane from a security standpoint. He used it as "probably the most expensive alarm clock ever." The agent ran on his Mac Studio in London and SSHed into his MacBook in Morocco to play music louder and louder when he didn't respond. The friends he showed it to were hooked, but his tweets about it got muted responses. He compares it to the iPhone: people had to use it to get it.
The name went from WhatsApp Relay to "Claudius" (a Doctor Who joke) to Clawdbot, which he felt explained the product better. On January 1 he did something he calls "absolutely insane": he put his agent, with full read/write access to his computer, into a public Discord. Visitors watched it check his cameras, run home automation, DJ, and look at his screen to tell him whether his coding agents were done. The project went from about 100 to 3,300 GitHub stars in a week, by his account, and he says he has merged 500 pull requests. He calls himself "a human merge button." He was about to merge a feature that lets the agent phone a business to make a reservation. He jokes that he built Anthropic's best marketing tool, since people bought additional $200 subscriptions to run it.
CLIs over MCP, and where the hard work actually is
To make the assistant useful, he built CLIs for everything: Google services, his bed, lights, music. He calls MCP a crutch whose best effect was pushing companies to open more APIs. His complaints: all tool definitions are loaded into context at session start, and responses can't be filtered. His example is a weather MCP returning 500 cities or dozens of fields when you only want to know whether it's raining. A CLI output can be piped through jq. MCP calls also can't be chained into a script. The host argued these are solvable problems. Steinberger said companies are working on discovery, but chaining remains an issue. He built MCPorter to convert MCPs into CLIs. Clawdbot has no native MCP support, but through that route it can load an MCP on demand, even from your phone.
He says nothing in Clawdbot is technically difficult. It is TypeScript "that shoves JSON around," moving text between LLMs, disk, and messaging platforms such as Teams, Slack, Discord, Signal, iMessage, and WhatsApp, with more coming. The difficulty is making it feel magical. The one-line installer checks for Node and Homebrew, handles older installs, and picks defaults so the user mostly presses Enter. The user is then asked whether to "hatch" their bot. A terminal UI opens, and a bootstrap file tells the model it is being born and should work out its identity with the user. It asks curious questions, then deletes the bootstrap file and writes user.md, soul.md with core values, and an identity file with a name, emoji, and inside jokes. It keeps maintaining these files over time. Then it messages you on WhatsApp. Users configure and even update the agent by asking it to.
For the first weeks of traction there was no real onboarding. Users were told to point their own coding agent at the repository and let it configure everything. He says this worked partly because the codebase was built by agents and follows the naming conventions agents expect.
Companies, prompt requests, and "I don't care about CI"
Steinberger expects large companies to struggle with AI because it requires refactoring the company, not just the codebase. He sees a need for people with product vision who can do everything, and fewer of them. He repeated the 30% figure and called the economic consequences frightening, with many people likely to struggle to find a place.
He now treats pull requests as "prompt requests." He thanks the contributor, then starts with his agent from the PR and redesigns the feature as he sees fit. The PR's code is rarely reused, but it tells the agent the goal and sometimes points to tricky bugs. He says the average quality of PRs has dropped because contributors vibe code without understanding the overall design. He asks contributors to include their prompts and reads those more than the code, because they show how the solution was reached and how much steering it took. He discourages small fix PRs, since reviewing them takes him ten times longer than typing "fix" into Codex.
He relies on local CI. The agent runs a "full gate" of linting, building, and tests, a term he picked up from the agents themselves. If it passes, he merges. Main occasionally slips, but not by much, and he doesn't want to wait ten more minutes for remote CI after already waiting on the agent. In his Discord, discussion is about architecture and style, not code. His example is the voice-calling PR. It touched many places, and he felt Clawdbot was "becoming bloatware." He had Codex compare the PR with an earlier unfinished project of his and with Mario's coding agent Pi, which loads TypeScript plugins. The result was the plugin architecture he built the night before the interview. Pointing agents at folders where he has already solved a problem is one of his main techniques, because it conveys his earlier thinking more precisely than re-explaining it.
Advice for hires and new graduates
If he hired, he would look for people active in open source who "love the game." He compares learning to use agents to learning an instrument or starting at the gym: painful at first, then rewarding. He says he once made 600 commits in a day and that an outside reviewer concluded the code was not slop. He has never worked harder, even at PSPDFKit, partly because it is addictive and partly because he wants to use the project's momentum.
For new graduates, he acknowledges that entering the market will be harder. His advice is to be "infinitely curious": read complex open-source code and ask a patient machine why it was built that way, in order to build system understanding. He doesn't think universities teach this well. Newcomers have one advantage, though. They aren't held back by experience and will use agents in ways veterans don't consider. He admits he still defaulted to opening Instruments to profile a slow menu-bar app until the agent did the whole performance analysis from the terminal.
In the rapid-fire round, his favorite gadget was a cheap Android photo frame that friends can email pictures to. It gives him more joy than the iPhone 17 he still hasn't unpacked. He recharges at the gym with a coach and his phone left in the locker, and sometimes walks without his phone.
In closing, the host highlighted prompts over pull requests, closing the loop, and Steinberger's report that juggling agents is more exhausting than writing code. The host added a caveat: Clawdbot is more of a "YOLO project" than most production apps, so these practices should be taken with a grain of salt. The host expects much of this to spread to production work, but with review and validation becoming more important steps.
What if you could merge 600 commits on a single day and none of it was slop? This is what today's guest, Peter Steinberger, the creator of Clawdbot, claims he's doing. Peter is a standout developer who built PSPDFKit, the PDF framework used on more than 1 billion devices. Then he burned out, sold his shares, and disappeared from tech for 3 years. This year, he came back and how he builds and what he's doing now looks nothing like traditional software development.
In today's episode, we cover why he no longer reads most of the code he ships, and why that's not as crazy as it sounds. How he is building Clawdbot, his wildly popular personal assistant project, which feels like the future of Siri, the closing the loop principle that separates effective AI assistant coding from frustrating vibe coding. Why he says code reviews are dead and PRs should be called prompt requests, and many more. If you're interested in how the software engineering workflow could change in the coming years thanks to AI, this episode is for you.
This episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes to learn more about them, the Pragmatic Summit on the 11th February in San Francisco that I'm hosting with them, and our other season sponsors.
Right, Pete, welcome to the podcast.
Thanks for having me, Gergely.
It is awesome to meet you in person.
Yeah, and I almost messed it up.
Yeah. What happened? You lost track of time. Does that happen often? And how so?
Not usually. Not usually. This is an interesting time for me because my latest project is blowing up.
Clawdbot, right?
Clawdbot. Yeah. I'm struggling a bit to get enough sleep, but it's interesting. I never had a community blowing up so fast and it's just incredibly fun to work with.
So, before we get into Clawdbot and all the fun stuff you're doing, I wanted to rewind all the way back. You created PSPDFKit, which is used I think on more than 1 billion devices. If you see a PDF rendered, you probably see that. But even before that, how did you get into tech?
Oh my god. How did I get into tech? So I'm from rural Austria. Always more being the introvert. So eventually, we always had like summer guests and one of them was a computer nerd, and then I kind of got hooked with the machine he had and begged my mom to buy me one. And ever since then... this was in high school or so. I guess I was 14.
Yeah. And ever since I started tinkering, like I can remember the earliest thing was like I stole an old DOS game from my school and then wrote a copy protection for the floppy disk so I could sell it. It took like 2 minutes to load. I was just always tinkering. Also playing a lot of computer games of course, but like building stuff almost feels like playing a computer game. Like definitely right now it feels better than Factorio.
When I started out, I wrote like the equivalent of bash scripts for Windows and then I did like websites. So, I guess a little bit of JavaScript, even though I had no clue what I was doing. And then the actual first language where I had to learn how to build things is when I started university.
And I never met my dad. And I come from a poor family. So, I always had to work. Like I had to finance my own studies, right? So when other people were having holiday, I just worked full-time at a company. So the first real job I had was in Vienna. It was supposed to be 1 month and then they kept me for 6 months. It was just a bridge between military and my university. And I kept working there for like I think 5 years. And I remember the first day they gave me this huge book, maybe that huge, and said Microsoft MFC. I still have nightmares. And I was like, this is terrible. So for the next one, I just silently used .NET. I just didn't tell them, and like a few months in I just told them. I did a few modernizations, but then it was too late. I did this a few times in this company. I don't know why they kept me, because it worked.
So I did .NET and actually I actually dug it. .NET 2.0 had like generics. It took insanely long for the application to launch because like everything was compiled at first start and like your hard disk was like... if you remember.
So how did you stumble into iOS, and where did the idea for PSPDF come from?
Not even... yeah, the first one wasn't even available in Austria.
That's true. Yeah.
A little time went on and I was at university and a friend showed me the iPhone, and I think I touched it for a minute and then immediately bought one. Like it clicked when I felt it, and to me this was like a holy f moment because it was just so different and so much better. So I got one. I was still not thinking about building for it, you know.
When was this, 2009, '10, something like that?
Yeah. And then I used their browser. I can see the story. I was literally riding in the subway. And at the time I was using a gay dating app, and this was iPhone OS 2. So I typed this long message. I pressed send and we were just going into a tunnel, and the JavaScript disabled the send button and then an error message came, but there was no copy paste. There was no screenshot. And I couldn't scroll anymore because scrolling was disabled. So this long message, which was a little bit emotional, was gone and I was so mad. I was so mad. I'm like, what the hell?
I went home and I downloaded Xcode. That's where the window came and I was like, where is the IDE? I was like, this is unacceptable. I basically hacked the website. I used regular expressions to parse the HTML, which is totally not something you should do, and I built an app. I used iPhone OS 3 beta with Core Data in beta, RegexKitLite. I used a hacked version of GCC that backported the blocks compiler so I could use blocks in iPhone OS 3. It took me quite a while until anything worked because I had no idea what I was doing and I was using all kinds of beta tech, but eventually I got it to work. And I wrote that company, like, hey, I'm making an app. What do you think about it? Got no response, of course. So I was like, let's just put it in the App Store.
And this was for the dating app, right?
Yeah.
So you just, you know, you looked at their APIs, you could just easily build a client on top.
API? It was HTML. I was just literally parsing HTML.
Oh. So you kind of parse the HTML, kind of turn it into your own, you know, like you use it as an API. Oh, clever.
I mean, this was back in the day where no one thought this would happen, but I put it in the App Store. I charged five bucks for it and I made like 10k in the first month. And I had no clue what I was doing. And there was so much complex tech stuff. This was very early on, where there were a lot of weird forms at Apple. So I just put in the bank account of my grandpa. And then one day my grandpa called me: "Yeah, something is weird. I got this huge payment from Apple." I'm like, "This is mine. This is mine. Don't touch it."
But the funny thing was when this blew up, I remember I was in a club one day and I saw someone using my app, and I was so proud and I wanted to tap him on the shoulder and say I built this, and I thought that would be really weird. So I didn't. And then I went to the company I worked for for 5 years and told them, I'm going to pursue this. This is really exciting. And my boss was mocking me.
Oh, really?
They're like, "Oh, you're making a mistake. This is a fad. This will not go." Blah, blah, blah. And you know what that got me? That's what you call a chip on your shoulder. I'm like, you know, one day I'm going to have a company that's worth more money than yours. Well, it took me eight years.
So, I got hooked. I'm a little bit of an addictive personality, which you see again right now. But I worked a lot on this app. I learned at high speed, and this was also the time when I started Twitter, and that was hugely influential for my career. I made this app actually quite good, and then one day I was at a party at 3:00 a.m., slightly intoxicated, and I got a call from a US number. The guy on the phone was like, "Yeah, hello. This is John from Apple. There's a problem with your application. Some people reported pictures." And that was it. That was the end of my app.
It was good while it lasted.
And I had just quit my job and was like, well, f you Apple. I did freelance work. I was at Dub Dub.
WWDC.
Yes. Sorry for the insider terms. I was introduced to someone as one of the best iOS developers in Austria at a bar at 2 a.m. in San Francisco, and then basically got a job in the US, and then I moved to the US for a while. And then I went to the Nokia development days. This is all like stone age by now. My god. And then someone came up to me and said, "Yeah, we built this app somewhere in Eastern Europe and it works, but it crashes sometimes." It was like a magazine viewer, right? This was back when the iPad just came out and Steve Jobs hyped it like this is the savior. So everybody was building magazine apps, and I was like, that sounds like an interesting short-term gig. And I was like, okay, I'll help you out.
And I opened the app and it was like, oh, the worst code I've ever seen in my life. It was literally one file with like thousands of lines of Objective-C, where they used windows as tabs. I didn't know this worked. I was surprised this worked at all, but it felt like a house of cards. And I tried to surgically fix things, but as soon as you would touch something, something else would break. So I got it somewhat stabilized and I told them, "Look, this is madness. I'm going to rewrite this for you." Yeah, it took half a year; I'm going to do it in a month. Well, it took me two months. I wasn't that far off.
And then here I was working on a PDF viewer. You know, on every technical problem, the domain is, I wouldn't say completely unimportant, but you can always find interesting problems in every domain. And there were a lot of interesting problems, because you had a C call that would render a PDF that would maybe take 30 megabytes, but the whole system had 64 megabytes. So if you're not very smart and very careful about what you do in the background, the OS would just kill you. I got really fixated on making it good, like when you rotate, that the page would animate, and so on. You know, I like those details. I spent way too much time on that. That's why it took two months instead of one. But the end result was good.
And then I worked with them for a while, and then a friend texted me. He's like, "Yeah, I'm working on this magazine app and it's really hard." I'm like, "Yeah, no way it's hard. I know. I did it."
You just built one.
And he was like, "Can you get me the code?" I'm like, "Sure." So I extracted the part that was PDF from this magazine app, and I made sure the other person was okay, and then I sold him that. I was like, well, if he's interested in that, why not try to sell it to other people? I used a WordPress template and mutilated it to run on GitHub Pages, and then when you did the fast lane flow, at the end you got a Dropbox link to my personal Dropbox with a source code zip. I put this up in one afternoon and I tweeted it, and then in that week three people bought it. It was like I guess 200 bucks back then, and for me this was amazing. And not only did I get three people who just bought it, but like 10 people who complained because they wanted it, but it didn't have the features they wanted. It's like I got nerd sniped. I was like, "Oh, it doesn't have text selection. How hard can it be?" 3 months later: "Oh, yeah. It's really hard."
Text selection in a PDF specifically.
Yeah. You know the saying, companies are built by young people because they don't know how hard it is. Yeah. I had no idea what an insane madness this file format is.
Peter was talking about how some problems look deceptively simple. PDF rendering is a good example. You look at it and think, how hard could it be? And then you spend months on edge cases that you didn't even know existed. This "looks easy until you build it" pattern shows up in other places, too. Internal tooling for feature flags and experimentation is a classic example. Teams often underestimate how much work it is to build infrastructure around these tools. There's a reason big tech companies like Uber invested years into building internal experimentation and feature flagging systems.
Which brings me to Statsig, our presenting partner for the season. Statsig gives you the complete toolkit without building it yourself. You get feature flags, experimentation, and product analytics all in one platform, tied to the same underlying user assignments and data. In practice, it looks like this. You roll out a change to 1% of users first. You see how it moves the top pipeline metrics you care about: conversion, retention, whatever is relevant for the release. If something goes wrong, instant rollback. If it's working, you can confidently scale it up. Companies like Notion went from single-digit experiments per quarter to over 300 experiments with Statsig. They shipped over 600 features behind feature flags, moving fast while protecting against metrics regression. Microsoft, Atlassian, and Brex use Statsig for the same reason. It's the infrastructure that enables both speed and reliability at scale. Statsig has a generous free tier to get started, and pro pricing for teams starts at $150 per month. To learn more and get a 30-day enterprise trial, go to statsig.com/pragmatic. And now, let's get back to Peter and why rendering PDFs was a surprisingly hard problem.
But now, I remember a few weeks ago someone emailed me. They did something PDF and they wanted my help. And I just wrote them like, I'm sorry, I did my deed. I know more about PDF than any sane human person ever should know. And I went to therapy. Good luck.
But that took off, and while I was waiting for my visa, I worked on this project and it just kept on... more people kept on buying it. And you know, it was summer. I was lying at the lake and got another email that someone bought it for 600 bucks, 800 bucks. I just upped the prices as it had more features. And by the time I went to San Francisco to work at this company, it already made more than what I made there. But my whole life was... I still thought like I have to be there, you know. So I did it. And also interestingly, at this company, I had to...
So what would you say, that you moved to San Francisco?
Yeah. And of course it also ended up being something where I had to build something with my framework at that company too. But you know, startups are not like 8 hours. They're a little more. And my personal project was also a little more. So my sleep was a little less. And then eventually, after 3 months, Sabine, my manager, came over and said, "Peter, are you okay?" And they gave me a choice to either keep working at this company and drop my project, or vice versa. And I had one week to decide. The counter was one week to stay there or leave the country, because I was on a complicated visa.
And well, the decision was quite easy. It's like, yeah, I want to do my own thing. And then—
And at this point, it was already taking off. You already saw that there's a big business here. It will probably pay you as much as your US job would have paid.
It was never money driven either.
What were you driven by?
I want to make stuff that other people find amazing. I love tweaking the details. I love those little delights. It wasn't even that there were no competitors in the space. But my angle was always: I built something as if Apple would have built it, with all the love and care and polish and those little delights that a lot of people in the industry don't get. So even though we had competitors that had way more features and were around way longer, my company was more successful and my product was more successful, because developers tried the different ones and mine just felt the best. I think software is all about how it feels, much more so than the feature set. Like, why did we buy Apple stuff? It has more features than Windows. But it feels better.
So you left this company and you were building this PDF component that started to sell. At what point did you hire the first person, realizing, okay, there's something more to this?
When I went back to Vienna, then I was like, okay, I have to go all in. And that's where I started working with freelancers a little bit, and way too late, to be honest. I could have hired much earlier, but you know, it's a big step. And that's kind of where it started having a life of its own, and I spent pretty much 13 years of my career building this product with this weird name that I never changed, because I thought about it for like five minutes and then stuck with PSPDFKit.
PSPDF.
They finally did a rename, but I wouldn't have renamed it.
It's a mouthful, but it's very unique.
Well, you get it if you do Objective-C, because it's just a namespace.
Yes.
And at the time it made perfect sense. My strategy for marketing was always: I only care about the developer. I know upper management makes the decisions, but if I can convince the people inside the company, they all do the marketing and lobbying for me. That worked really well. We never did cold emails or anything aggressive. It was all inbound. All we did was make good stuff and write insightful technical blog posts, and I went to a lot of conferences. For me it was important that people understand that the people who built this product know what they do and love what they do. That reflects on the product, and that worked really well.
And then what was the tech stack behind PSPDFKit? Was it Objective-C? Was it later Swift? Were there other technologies like C or anything else?
We eventually expanded to all the platforms. A big shift was the switch out of the renderer, which was and is still quite buggy, to a big C++ one that we then used across all the frameworks. We were really early with web. We were one of the first PDF frameworks that ran in WebAssembly. And I did the most clever thing: it was in the very early days when WebAssembly was just taking off, and we built a benchmark, and that benchmark was eventually used by Google and Microsoft and Apple. So I basically had all these companies working really hard on making my renderer faster, because they used our benchmark as one of their benchmarks, and the benchmark was just rendering our stuff with our renderer.
Ah, nice. And then as the company grew, one thing that I remember about PSPDFKit: you did write a lot of blogs, and one blog in 2019—so this was, I think, year nine or 10 of the company—was about how the team worked. You mentioned things there like every feature starts with a proposal. You mentioned that you are conservative because it's a big API that people use and you want to be careful, and things like the Boy Scout rule. How did you put together the culture of this team, which was now closer to 30 or 60 people?
We were actually 70 when I sold my shares, and now it's almost 200. And I knew right from the get-go that I'm not going to find the people that I need in Vienna. So it was always remote first, and eventually we landed on some kind of hybrid model, which made things a bit more complicated. I learned a whole lot on the go. I never had the urge to be CEO. I was always coding. I brought in people that helped me a lot with other parts. And on the business side, I can do it, and I think I'm quite good at it, but I just don't enjoy it. Even on sales calls, where you have to think of the magic number—how much it would be worth—because that's how enterprise works. Ah, the worst.
Peter just said, "Ah, the worst" about enterprise sales. Because selling to large companies, enterprises, is as tricky as it gets. Not just because you need to get pricing right, but because of all the enterprise features that you need to build. And this leads us nicely to our season sponsor, WorkOS. If you're building with AI agents or automation tools, here's a problem most teams don't think about at first: once an agent can take actions on your behalf, you need to control what it's allowed to do. And traditional auth just wasn't designed for that. That's why WorkOS introduced MCP Auth, which gives teams a way to authenticate AI agents with explicit permissions, auditability and enterprise-grade security. Instead of sharing overscoped API keys, you can define clear boundaries for the data that agents can access and the actions they can perform. If you're building AI-powered features and want to ship fast without compromising security, check out workos.com/mcp. And with this, let's get back to Peter and enterprise pricing.
But that's also the only thing that really works on a model like this.
You mean enterprise sales specifically, right?
Meaning custom pricing.
So can you tell us, for devs listening who go to a vendor's website and are frustrated that there's no price—it says call us or schedule a meeting—why that is?
Oh, that's why: because we're going to look at your company and then just take the dice and think about the number that you're probably willing to pay. And that sounds horrible. But also, when you have a product where you can't really break it down to a specific number, it makes a difference if a freelancer contacts us or one of the big Fortune 500s. Let's not say names.
Because the usage will be different. The value they get out of it will be different. And charging the same, you would either exclude one or the other. If I go too low, they're going to see this as fishy. It's like, procurement for 500 bucks? We're not even going to start the process. And if we target it too high, we're going to lose those people. So as horrible and unfair as this process seems, for some kinds of products it's the most fair way after all.
You know, in software, I would say there are four axes. There's easy and hard, and interesting and not interesting. We were very much in the not interesting and hard part. If you build something that every developer wants to build, it's going to be a hard sell. It's a hard sell anyhow. Selling anything to developers is a hard sell.
Yeah.
But if it's too easy or too interesting, good luck. But if it's "oh god, I don't want to do this" and "oh my god, this is hard," that's a good spot to be in. So I found a really interesting niche, and there were just an infinite number of complex problems.
You need to tell me one or two hard things about parsing PDF. How hard could it be? There's a specification. I'm an engineer. I know specifications. What's so hard about it?
I mean, there was this one example where, you know, PDFs have links. So there's a table of contents and you click on it and it goes to page 37. So I built this whole model with the assumption, oh yeah, maybe there's 100 or 400 links in there. And then we got this one customer who paid really good money, and I was like, oh, it takes 4 minutes to load a PDF. What the heck, guys? And I looked at it and it was a 50,000-page text bible from Canada, and it had—
50,000 pages.
It had more than 100 links per page.
500,000 links.
My data model completely exploded because my assumptions were off by a factor of, what, 1,000. But by then you have a mature product with an API. So how do you completely redesign the internal part without breaking things for everyone? Suddenly everything has to be lazy, where before parsing 100 was easy. This was so difficult to keep working for people. I think I spent two months just on that, completely redesigning the internals and making sure it's still easy for people. They don't have to know what we load eagerly and what we load lazily, or if you copy this thing, it still has to keep some connection.
It needs to keep the references and some of those things.
And I love to do support, and I think that was also a contributing factor why the company worked. Because if you send a ticket and then the CEO replies and helps you out, that has impact. And my strategy was always, I always used to answer in reverse, because if you send a ticket and you get a reply within 5 minutes, that's magical. If you wait one or two days, not much difference.
Yeah.
So yeah, this was one of the problems where I worked two months and I finally got it down to almost like this.
That must have been satisfying.
It was very satisfying.
And you were writing a lot of the code, or you were involved in a bunch of the code. Obviously a big team was now there, but you were still overseeing it, right? You're in the details.
I mean, of course, I had a really great team, and in some parts I was more involved. I was always more involved in mobile because that's where my heart was, but I was always very deep in the tech. And on the marketing side, the business side, I had Jonathan's help, I had marketing help. I found good people. The thing is, blogging and writing about how you solve interesting hard problems will help you hire interesting people that want to solve interesting problems.
This is what I remember about PSPDFKit: your blog was, every now and then it made it to Hacker News as well, but it was just interesting to read. And again, I'm not into PDFs, but if I had to name something PDF, I would have said PSPDFKit, because they're the only ones where I read interesting engineering blogs about how you optimize your stuff. It's still there, by the way. I myself also sometimes ask: interesting, do more companies not see this? Or is it that you need to be a developer who's either the CEO or up there who just likes doing this? And by the way, did you ever write this thinking this will be helpful, or did you just write because you got something out of it, like putting out that you solved this hard problem?
I like sharing and inspiring people. There were sometimes even conflicts where we were like, should we write about this, because it's a little bit of secret sauce. But I just never listened to those voices too much. There's also, when you write something down, this principle that you understand it, but if you want to teach it, you really have to understand it. So to me it was also a little bit like, oh yeah, I worked on this really hard problem and now I want to preserve it and help others. So of course I got a kick out of it; of course I liked the attention. But really, sometimes I just referenced my own post a year later. This is both company documentation and my own lookbook. It's helpful in so many ways. And a lot of those bigger companies, oh, they put on too much red tape. There are a lot of developers who don't really like to write. So I forced everyone, once a month, to take a full day just to write a blog post.
But you gave them the time. You're like, that day you don't need to do any other work but write something.
Yeah. You have a day to come up with a post.
Ah, a day is quite much, actually.
I mean, nowadays when I write posts, it still takes me a few hours. I don't want to dwell too much on it. I think the starting time of the company is the most interesting. Then in the growth phase, you get more red tape, you get more people. It's much more gardening your product instead of doing wild hacks, and more iterative. So it got a little bit less interesting over the years, and there was more people drama, because the more people you have, the more issues there are. And I didn't enjoy it that much, and I was really, really burned out.
What burnt you out, do you think?
I was just burning too hard. I was working most weekends. I tried to shuffle all my managerial needs. And you know, as a CEO, you're basically the waste bin, because everything that other people don't manage or can't do or mess up, you have to fix. And it's also quite lonely, because you can't openly talk about a lot of things. I structured the company to be quite open, but still, you cannot be negative, even if really bad stuff happens. There was one weekend where my co-founder called me at 5:00 a.m. and told me, yeah, there's this big airplane company and their planes are down because our software is crashing. That was a very interesting weekend, until I disassembled their app and proved that they had messed around with our source code, triggering a license key fallback. That eventually caused the issue they had. But that was an "if they sue us, the company's gone" moment.
And that's just on top of all the additional stress, and there were quite a few of those things. You can do that for a while. And I also believe burnout doesn't necessarily come from working too much. It comes more—at least for me—from when you work on something but you don't believe in it anymore, or you have too many conflicts. And we did fight a lot in the management team. And at the time I made this mistake: I thought you have to lead a company more democratically. So that was also something that burned me out. I wouldn't want to have missed it, though.
Yeah. So from the outside, it seems you sold your shares, you made enough money to not have to work again, should you choose not to. And for a lot of people, people who are starting out their business or one day want to start a business, this sounds like the absolute dream. We know realistically that most people will not make it, but if you make it, you've kind of, you know, checkbox done. It's a little bit like if you're climbing a wall and you ring the bell, you're done. And then what I noticed, from the outside again, on your blog, the blog posts completely stopped for several years. What did you do in this time, and what did you learn in this time, before you came back to where we are now?
I needed a lot of time to decompress. I caught up a lot on the things I thought I missed. There were months where I didn't even turn on my computer. And for a while I was I
I just didn't have this feeling of, what should I do now? I definitely was like, why bother? You know, you're not supposed to retire so early, or have such a good exit that you never have to work again. That messed with my mind quite a bit. That was some hard years.
And then in April, there was this idea that I had years ago, and even a side project that I started, and I was like, oh yeah, I want to continue on that. And then after more than 3 years, I just sat back at my computer and started hacking again.
But the thing was, this was a Twitter analytics thing and it was written in Swift and SwiftUI, and back then I already knew it would be so much better if I built it as a website.
So was this an existing idea that you kind of had at the back of your mind, something Twitter analytics?
Yeah, it was just something I wanted to build for myself because it didn't exist. And even three years later it didn't exist. It still doesn't exist. It kind of does, but I got a bit sidetracked. So I went back and I wanted to build it with web tech. But web was always, even at the company, the one thing that I looked into the least, because I had someone really smart who took care of that side in the company that I brought in, Martin. So I never had to worry about it. That was one of the
You're not hands-on with React or any of that stuff.
Yeah. And when I came back, I was like, what's a prop? You know, that level. And this is a trap I see with many developers. The better you get at one technology, the harder it is to jump somewhere else. It's not that you can't do it, but it hurts so much. I can program in the Apple stack blind. But then in that stack, I have to Google the most mundane stuff and it just hurts. You feel like an idiot again.
Yeah. And I guess the more experience you have, it kind of sucks, that feeling. I mean, I'm sure you say embrace it and all that, but it's not great. You're not as efficient. You know that you could be faster, etc.
Yeah. So I came back and I was like, gosh, there has to be... what is this AI stuff that people are dismissing? Let's look into this. And in April, a lot of us were skeptical, probably rightfully so. And to a degree I credit those three years where I basically didn't turn on my computer, because in those years you guys checked out AI and learned that it's crap.
Yeah. I was about to say, so you missed out on... you didn't do the beta of GitHub Copilot, you know, glorified autocomplete, which was GPT-3, or maybe not even. There was then of course 3.5, which was a big jump, and it got incrementally better with GPT-4. And so by the time you came back, what tool did you first use? Because you missed out on like two years of us devs using it, dismissing it, finding some niche use cases for it.
Oh, Claude Code. So you started with Claude Code, that I think came out...
It just came out. It came out in May, but there was a beta beforehand.
Yeah. I think they had something. Didn't they have something in February already?
They had a beta from February. Correct.
Yeah. So,
So Claude Code was your first. You come back after a hiatus and you immediately turned on Claude Code and you missed everything else before.
And you know, I remember I took this big messy side project that I built, and I have this browser extension that converts a GitHub repository into one big markdown file. That was like a 1.3 megabyte markdown file, and I dragged it into Google's AI Studio with Gemini 2.5 or something, and I typed "write me a spec," and it generated those 400 lines of spec. And I dragged this back into Claude Code and I was like, build. And then continue, continue, continue, while I was working on other stuff, you know. And eventually it told me it's 100% production ready, and I started it and it crashed.
I'm sure we can all relate to the story of the AI saying the code is production ready, then crashing. This is a pretty funny and innocent story, but I personally don't trust code that AI generates without verifying it. And this leads us nicely to our season sponsor, Sonar.
So, let's look at some data. A new report from Sonar, the State of Code developer survey report, found that 82% of developers believe they can code faster with AI. But here's what's interesting. In the same survey, 96% of developers said they do not highly trust the accuracy of AI code. This checks out for me as well. While I write the code faster with AI agents, I don't exactly trust the code it produces. This really becomes a problem at the code review stage, where all this AI-generated code must be regularly verified for security, reliability, and maintainability.
SonarQube is precisely built to solve this code verification issue. Sonar has been the leader in the automated code analysis business for over 17 years, analyzing 750 billion lines of code daily. That's over 8 million lines of code per second. I actually first came across Sonar 13 years ago, in 2013, when I was working at Microsoft Skype, and a bunch of teams already used SonarQube to improve the quality of their code. I've been a fan since. Sonar provides an essential and independent verification layer. It's the automated guardrail that analyzes all code, whether it's developer or AI agent generated, ensuring it meets your quality and security standards before it ever reaches production. To get started for free, head to sonarsource.com/pragmatic. With this, let's get back to Peter and how AI agents cannot exactly be trusted.
Then I added an MCP so it could use the browser. I think the Playwright MCP was already there. And it looped a few more hours, and then I had a Twitter login page and it did something. It was not great, but it did something. And to me, this was my holy, mind-blowing moment.
Yeah.
And this was like in April or May this year, right?
Yeah. It was just good enough that I could see the potential, and I understood: yeah, this is where it's going. And from that moment on, I had a few months where I really had trouble sleeping, and I...
I remember, because once on Twitter I sent you a direct message. I was up early for valid reasons, you know, my kids or something like that. But it was 5:00 a.m. and I sent you a message on Twitter and you replied immediately. And I was like, "Why are you up?" And you're like, "Oh, this is usual. I'm usually still awake." And I asked, "Why?" And you said, "Oh, I'm just using Claude and it's really, really addictive." And I was like, "Really?" And you're like, "Yeah, I'm not joking. It's really good." And I think that was the thing. You said or wrote something like "just one more prompt." Tell me what made it so addictive, or what still makes it so addictive?
Oh, it's the same economics as when you go to a casino. It's my little slot machine, you know? You press the trigger and ding ding ding ding ding, and it's like, nope. You type in the prompt and it does crap, or it does something that actually blows your mind, and it's this...
And you're saying it blows your mind, and you're a really experienced developer. It's not easy to blow your mind, right? You've seen good code, you can differentiate crap code, decent code, good enough code. You have a bar, right?
It's so funny, you know. In my company, I used to obsess over every detail, every spacing, every new line, the naming. I spent so much time bikeshedding. And in retrospect, I'm like, what the heck? Why did I do that? What's the point? The customer doesn't see the insides. Of course it has to meet certain standards. It has to work. It has to be fast. It should be secure. But how much of that bikeshedding was just stupid.
You say that, but then you also just said that people loved PSPDFKit because it was the most polished. It worked the best. Do you not think that that amount of caring, bikeshedding as you call it, being obsessed... It sounds like you were keeping tech debt at bay. You know, being obsessed with white spaces, it's not going to be messy. And we know it's not just the white spaces. We know you're going to care about testing and all that. What I see with PSPDFKit is you were not just building a product that had great UX, but you built something that had really good hygiene, and that's how it could be high performance and all that. How do you think about it?
Yeah. To a degree, yes. And even now, I mean, my last blog post was a confession that I ship code I don't read, and
Yeah, we have to talk about that.
And at the same time, I spent so much time restructuring. I mean, even today, I really wanted to get this PR in where it was like a 15,000 line change, where I moved everything over to a plugin architecture, which I was so excited about, and I care a lot about the structure. Did I read all the code? No, because a lot of code really is just boring plumbing. Well, what are most apps? Data comes in from an API in one form. You parse it, you package it into a different form. You store it into a database and it's a different form. It comes out again in a different form. Then it's HTML or whatever. And you type in something, it's a different form again. All you do is massage data into different forms throughout your app. This is what most apps are. We are pretty JSON printers. And the really hard part was solved by Postgres 30 years ago by some neckbeards. That's really what a lot of software is like. There's always some interesting parts, but I don't have to care how this button is aligned or which Tailwind class is used. Many details are boring and many other details are interesting. But I think it's much more about system architecture than having to read every single line.
Right. Now jumping forward, what is your workflow like? When you're working on Clawdbot, are you using a terminal, multiple terminals, which tools? Like you said, you're kind of not reviewing the code, but you're still thinking about architecture. What does your average day look like in terms of tooling? You know, if you had to explain it to a developer who might join the team, what does it look like?
It's interesting. Let's go back a little bit. We were in April with Claude Code, and then I got really hooked, and then I had a phase where I did Cursor, and then I used Gemini 2.5 a bit. Then we had this phase with Opus 4. I hooked a lot of my friends. I know both Armin and Mario from Vienna. They got AI-pilled because it was addictive. You know, my energy was confusing them, and then they tried it out, and then eventually they also were up at 5:00 a.m., and I called it the black eye club. I mean, there's a reason I started a meetup in London that I called Claude Code Anonymous, because it's a little bit like a drug, because it's so much fun. To me, what blew my mind so much was this realization that I can build everything now. Before, you had to really pick which side project you build, because software is hard. Yeah, it's still hard. But now, this friction that I talked about, where I'm so good at this technology and I'm so bad at this, and I'm like, oh, let's make the CLI in Go. I have no clue about Go. But I have a good system understanding, and once you have that, you develop a feeling for what's right and what's wrong. It is a skill in itself. I remember there was this tweet where someone said, "Oh, when you write the code, you feel the friction, and that's how you make good architecture." I feel the same friction when I prompt, because I see the code flying by. I see how long it takes. I see if the agent pushes back. I see if what it creates looks messy or makes sense. When I prompt, I have a hint already of how long it's going to take. If it takes much longer, I understand that I messed up somewhere.
You kind of feel the model. You know...
Yeah, usually it's like this, or if it runs...
I feel it's very much a symbiosis. I learned to talk, may I even say, their language more. So my knowledge of how to use those things improved, and also the models improved. And over the time between April and now, I would say the inflection point was summer, where it just got so good that you could create software without actually writing code by hand. But the real change that sold it for me was, again, GPT-5.2. Again, I think it's underrated. I don't know why all these people still use Claude Code. I kind of get it. It's a different way of working, but whatever OpenAI cooked there is insanely good. Pretty much every prompt I type gives me the result I want, which is insane. On Clawdbot, my latest product, I use between five and 10 agents in parallel. If you're very much Claude Code-built, you have to forget quite a lot of the silliness, the things that you have to do to create good output with Claude Code. I mean, I also met that team, and they created a whole new category. Claude Code is a category-defining product, and it is amazing for general purpose computer work, and it is really good for coding, and I still use it almost every day. But for writing code in complex applications, Codex is just so much better, because it takes 10 times longer. Claude would read three files and then be confident enough to just create code, and then you really have to steer it and push it so it reads more code, so it sees a bigger picture of your codebase, so that it weaves in new features better. And Codex will just be silent and read files for 10 minutes. And if you only work on one terminal, I completely understand how you find this unbearable.
But I'd rather have something where... it's also, you don't tell it what to do. This is also something that people don't get. I have a conversation with the model. It's like, oh, let's look at this. What options do we have for this structure? Did you consider this feature? Because every session, the model starts from having no understanding about your product, and sometimes you have to just give it a little bit of pointers. What about this and this? So it explores different directions. And you don't need plan mode. I'm just having a conversation, and until I say "build this," it will not build this. There are some trigger words, because they all are a little trigger-hungry, but as soon as I say "let's discuss" or "give me options," they will not build things until I say build.
So would you say a lot of your prompting, or a good part of it, is this conversation where you are pretty much planning together with the agent?
Yeah. It's like, what about... I say okay, then you remind them, it's like, okay, we need documentation, what would be a good spot? It would give me some recommendations. I say no, this should really be its own page. Do we need a configuration? How does this fit into this other feature? It's like I am designing the system, because I have this system understanding about how my product is, how are the
shapes looking. I don't have a line-by-line code understanding; that's what Codex does for me, but I'm the architect.
It sounds a little bit like you're almost, you know, years back this totally got out of style. But there was this idea that you would have the architect with a capital A who used to be a software developer, but they're not hands-on anymore because they spend a lot of time understanding the business and they have these developers working underneath them. And some companies still kind of work a little bit like this, but most modern companies don't. But some banks, etc. I met people there who are capital architects. They do the system plan. They talk with fellow architects. They have the blueprint and then they literally pass it down to the team, and everyone hates this model, obviously, because, you know, again, I think as people you kind of want more. The architect is never on call for this stuff, and so it just kind of breaks down in practice, and a lot of large companies just moved to the staff engineer model where you're kind of all working together.
Of course, there's people who might have more input, but it sounds like it's almost like this world where you are the architect who kind of, you know, you have your little agents who do the code, except in this case, you are of course fully responsible because you're still an individual contributor. You're not... okay, you might say you're a manager of agents or whatnot, but the code is yours. It's your responsibility. You're going to be on call. If you push out code that takes down Clawdbot, which it did just recently, you're on the hook for it, right? And I think the difference in this system when it was in companies was the architect was kind of shielded from the output of their work because there's so many people and so much process, etc.
Well, I wouldn't say architecture. I like the word builder.
Builder. Yeah.
And I think also there's a few categories that I see for people that are highly successful using AI and people who really struggle. I care more about the outcome, the product. I very much care about how it feels and everything, but how the plumbing works underneath, I care structurally, but not to the biggest detail. And then there are people who really love to code on hard problems, like think about algorithms, don't really like the "I'm building a product" with all the marketing. They like to solve hard problems, and those are the people who really struggle and often reject AI or get really sad, because that's exactly the job that AI does: it solves the hard problems.
Now, sometimes I give it some pointers, but many times I learned... I learned more this year than in the last five years around software architecture and designing. There's so much inside those monsters, knowledge, and everything is just a question away, but you have to know what question to ask.
Of course, I also built this Twitter thing and it's still not done, and I really hope I'll get back to it at one time. Everything worked, but if I used it more, at some point things got really laggy and weird, and then it worked again, and I just couldn't figure it out, and it was really difficult to debug because it was not easy to reproduce. It was just like you use it more and things get really slow. I basically had software in Postgres that would be triggered when certain inserts were done, and then the database would get really busy, and the model couldn't see it because it was so far abstracted from all the... you know, those models are really good at tracing through, but this was a side effect that was so hard to see because it was only in this one file, a function that had no connection to anything else, with a name that was not easily greppable. I just never asked the right question until I was like, "Do we have any side effects for this and this?" and I found it and I fixed it. Everything is just the right question away.
Yeah. But you need to have knowledge, expertise.
Yeah. Experience.
I mean, so these are the people who reject it, and then the people who care a bit less about how it's being plumbed internally but are just excited to build things. They're really successful.
And one thing that also helped me is, you know, when you run a company and then you hire people, you can't breathe down everyone's neck and make them have the line of code exactly that way. And there's a lot of people who didn't manage a team; they didn't have this experience of how to relax a little bit and understand that yes, this maybe is not exactly the code that I want, but it will get me closer to my goal. And for anything that is not perfect, we can always make it better and put more time into it. I very much believe in this iterative improvement. I had to learn to let go a little bit at my company. So then when I had Claude Code, it kind of felt like I have imperfect, sometimes silly, but sometimes very brilliant engineers that I have to steer and where we work together on a common goal. It felt a lot like being the boss again.
Yeah. And interesting, now, you know, you built software, I guess, the traditional way, you know, pre-AI, for 15 years or even more than 15 years, and you got really good at also leading a team and at how to have high standards. You really cared about the craft there as well. You've now kind of been, I guess, vibe coding or working with agents for a year. You're comparing the two. What do you think really changed? And what do you think are things that kind of stayed the same despite all
First of all, I don't like the term vibe code.
All right. How should we call it?
I think vibe coding is by now almost a... I tell people what I do is agentic engineering with a little star. Vibe coding starts at 3:00 a.m.
Now, because all the mundane stuff of writing code is automated away, I can move so much faster. But it also means I have to think so much more. I'm still very much in the flow. It is completely the same feeling for me, as I very much get in this flow state, but it is mentally even more taxing, because I don't have one employee that I manage. I have like five or 10 that all work on things, and I switch from this one part to this other part to this other part to this other part. Mostly because I'm designing this new subsystem or this feature, and then I know that it will probably take Codex like 40 minutes or one hour to build. So I want to have the plan right, and then I build it, and then I'll move on to something else, but then this is cooking, and then I work on this, and then this is cooking, and then this is cooking, and then at some point this is cooking and then this is cooking, and then I go back to this one.
So I switch around a lot in my head. I wish I wouldn't have to do that. I'm sure this is a transitionary problem, and at some point we'll have models and systems that are so fast that I can parallelize a little less. But to stay in the flow state, I need to massively parallelize. So that's how it works. I go back there and maybe tweak it a little bit more, but usually just try it out, and maybe then this is ready because this only took like 20 minutes. So I constantly jump around. Usually there's one main project that has my focus, and I have some satellite projects that also need attention, but where maybe I spend 5 minutes, it does something for half an hour, and I try it, and it doesn't need so much capacity up there.
This almost sounds, you know, like... two things come to mind. One is there's these games where you have to manage a kitchen with the employees, and you see the recipes or something come out, and you need to jump and do it again.
It's like StarCraft, you know. You have your main base and you have your side bases. They give you resources.
That as well. And also one thing that just came to mind, as you said, "I go there and I watch this and I make a decision," is when I see the chess grandmasters play multiple boards at once. You see sometimes they play 20 boards, and they go there, and you can see that they just see what's on that board, they make a decision, and for some boards they stop for longer, I guess better players or better opponents. It feels, you know, both are occupying 100% of their brain. You're occupying your brain, and you're kind of scaling yourself as long as you can context switch.
The difference was, with Claude Code, you have to work a little differently, because it is much faster, but then the output often doesn't work on the first try. So it makes something, but then it forgot to update three other things. It crashes, or you give it... The good thing, how to be effective with a coding agent, is always that you have to close the loop. It needs to be able to debug and test itself. That's the big secret. That's also, I think, part of why it got so much more effective. But yeah, with Claude Code I often had to go back and fix up the stuff, or it just takes a lot of iterations. So in the end it's not that much faster. It's just more interactive.
And these days with Codex, it just almost always gets it right. My general strategy is always: I build a feature, and of course you always let it write tests, and you make sure that it runs it.
It runs them. Yes.
So even when I write a Mac app, I don't know, just yesterday I debugged this feature where the Mac app couldn't find a remote gateway, but the same code in TypeScript could. But a Mac app is kind of annoying to debug, because it builds it, you have to start it, you have to look at it, you have to say, "No, this is not working." So now I just said, "You know, you're going to build a CLI just for debugging that invokes all the same code paths, that you can call yourself, and then you just iterate and you fix it yourself." And then it will just cook, and it just cooked for an hour and it was done, and it told me there was a race condition here and here and a misconfiguration, blah blah blah, and like, yeah, it sounds sensible. I don't need to see that code.
But you don't need to see it because you set up the validation loops and you trust that because it ran it. I mean, I guess it's not too dissimilar to sometimes when you work on a large project in a large company: when all the tests pass, I mean, it doesn't mean it's 100% there, but it's pretty good, and all the new code has tests as well, you know, someone thought about it and tested it and all that.
So even on my very latest project, we always had bugs, but Antigravity has a certain weirdness with how it takes tool calls in the loop, in the format. So you have to do some filtering.
Yeah.
And that broke a bunch. And it actually took me way too long to realize, what am I doing here? I just need to automate this. So I went to Codex: design live tests that spin up a Docker container, install the whole thing, spin up a loop, use my API keys from this and this file, and then you tell the model to read an image, create an image before, and then look into the image and see what it sees. So I don't just test the loop, I also test tool calling. Make it work. And then it solved itself. It took forever, but it tested all my API keys, from Anthropic over SEI over GLM, everything, and it fixed all those little inconsistencies where sometimes the tool calling didn't work or the ordering was wrong, because I closed the loop, and that's
And closing the loop, you mean just have a way to have the agent be able to validate its work?
Yeah, that's the whole reason why those models that we currently have are so good at coding, but sometimes mediocre at creative writing, because there's no easy way to validate, right? But code I can compile, I can lint, I can execute, I can verify the output. If you design it the right way, you have a perfect loop. Even now, for websites, I built the core in a way that can be run via a CLI, so I have this perfect execution loop, because the browser loop is insanely slow. You want something that loops fast.
So it sounds like one thing that is not really changing from before is, we had this before: backend or business-logic-heavy things could more easily be verified that they're correct.
Surprise: actually using agentic coding makes you a better coder, because you have to think harder about your architecture so that it's more easily verifiable, because verifying is the way to make things good.
Well, then remember, back even before AI, for complex systems, once you got someone who built these things before, what they started with was: how do we make it testable, right? You need to design interfaces and classes to be testable. You need to think about, am I going to fake things, will I use mocks, will I use end-to-end testing, which will be long, etc. But these are really hard architectural decisions, and once you make them, they're, I guess, harder to change. In your world, you know, the model would cook a lot longer if you asked it to make a massive refactor, and, you know, if you have tests it'll get it right, but, you know, we still have these trade-offs.
Yeah, it's still software. I would say I write better code now that I don't write code myself anymore, and I wrote really good code. But even back at the company, sometimes testing was so tedious, and you come up with all those edge cases and the branching.
I mean, outside of Kent Beck, who I deeply respect, and he was on the podcast, and we talked, he still writes tests first, and he tells me that he's not mad at me for not writing them, but if you want to write, you know, poor-quality code, it's on you. But I don't know many developers, myself included... I never liked writing tests, and even when I pretended that I did, I just never did. It's a little bit like writing documentation. Writing tests, to me, was never a creative expression.
It is so good now. I would say for my last project I have really good documentation, and I didn't write a single line myself. No, I don't write the tests. I don't write documentation. I explain to the model the trade-offs, so why we did something like this, and then tell it, write the intro section beginner-friendly and then add more technical detail at the end, and it is so good. I never had a project with that good documentation, just because every time I design a feature, this is part of the process. And also testing: I was like, okay, we built this, how are we going to test this? Yeah, we could do this and this and this. What if we build it this way? And oh yeah, then we can test it better. So this is now part of my thinking, because I always think: how do I close the loop? The model always needs to be able to verify the work itself, which automatically steers me to better architecture.
So why do you think there's, you know, a bunch of experienced devs who are still pushing quite a bit back on just the idea that AI can do a lot of this?
That was a week ago. I stumbled over a blog post by Nala Coco with love, that I deeply respect and learned a lot from. And this blog post was just a dissing of the current way models work. And what he did was he tested like five or six models, including some that make no sense, like the OpenAI 120 billion open-source one that is not good enough to write good
code, you know. And he wrote a prompt, as far as I understand it. There was not a lot of information on the website, but to me it sounded like he wrote a prompt, he put it on Claude web and he pressed send. And then he took the output and ran it and it didn't compile and he was disappointed. But it's like, "Of course it will not work. Do you think I can write bug-free code on the first attempt?"
And those models are ghosts of our collective human knowledge. They work very similar in many ways. Of course, you don't get it right the first time. There will be mistakes. That's why you have to close the feedback loop. And also you don't just send a prompt to the model, you start a conversation: "Hey, this is what I want to build."
He complained that it used old API. Yeah, you didn't specify the macOS version. So it made an assumption to default to old API because that information was missing, and it is trained on a lot of data, not just the last two years, and there's just more old data than new data. So the more you understand how those little beasts think, the better you get at prompting.
And then he spent maybe, I don't know, a day or so playing with it and then just decided that this technology is still not really good. But to be effective you have to spend significantly more time. It's like you know how to play guitar and I put you on the piano and you tried a bit. It's like, "Oh, this sucks. I go back to my guitar." No. It's a different way of building. It's a different way of thinking.
You have no idea how often I screamed at 3:00 a.m. at Claude Code because it did something silly. I slowly started to understand why those things do what they do, with exactly the way I tell it to do things. And sometimes you can literally ask. Even last year, for this project, the last project, Clawdbot, I feel like a human merge button because the community is blowing up and all I do is reviewing PRs. I have very little time to actually write code myself anymore.
And in the beginning it would often just cherry-pick things and would close the PR and I was so annoyed. I'm like, why are you doing this? "When you say this and this, I interpret this and this." It was like, ah, I learned the language of the machine a little bit more. I tweaked my prompting and now I get exactly what I want, because it's a skill like any other skill.
Yeah. And Simon Willison has been saying the same thing even though he's been using it for years. And I think once I start to use it I also realize, like, I'm okay at it but I could do better.
What if we put this to a real test? Because I think it's fair to say that right now you're building Clawdbot, which is not something that generates revenue. There's a lot of users and it's blowing up and it's a really cool tool, but it's not PSPDFKit, which is a business where a lot of revenue is hinging on it. If today we just wiped it, PSPDFKit does not exist, you need to rebuild PSPDFKit, you now have these agents, how differently would it look? How much would you trust it? What would you delegate, what would you validate? And when you built up a team around it, because now it's a profitable business, at the very least you need to hire salespeople and whatnot, how do you think the team would look different today with that same product? Because you know exactly what it took to build it and you also know what these tools can do today.
I could easily run a company with 30% of the people. It would probably be quite difficult to find people on that level, but you want to have really senior engineers that really understand what they build, but that are also comfortable in delegating and know which parts are actually important to work on and which parts I can vibe.
That's still something I don't see a lot, especially in the AI world. There is so much crap on Twitter, there's so many people that are loud but clearly have no clue what they're doing. There's so many dumb concepts around. I'm sorry, but the Ralph Wiggum one, ugh. This is again another silliness people use to work around model limitations of Opus that you don't even need when you use Codex. There's maybe a few cases where you have a really long list of individual tasks that can be automated, but that's usually not how software building works.
I see so many people building up these elaborate orchestration layers, and then you have Beads that automatically creates tickets, and then your agent does tickets, and then your agent emails the other agent, and then you build up this elaborate mess. What for? Oh yeah, they design the spec for a few hours and then the machine builds it in the whole day. I don't believe this works. This is the waterfall model of software building. We learned long ago that this doesn't work. Yes, people work differently and maybe it does work for some. I just don't see how this could work for me.
I have to start with an idea, and often I purposefully underprompt the agent so it would do something that would give me new ideas. Maybe 80% of the things it assumed were crap, but there were two things like, oh, I didn't think about it that way.
Mhm.
And then I iterate and shape the project. I have to click with it. I have to feel it. I feel, to make good software, you know, one thing those things often lack is taste. I have to feel how this feature feels. And the beauty now is that features are so easy I can just throw it away or reprompt it.
My building model is usually very much forward. It's very rarely that I actually revert and have to go back. It's just like, okay, no, then let's change this. No, let's do this. It's like shaping. I love how you start with a rock and then you chisel away at it and pick different areas, and then slowly this statue emerges out of marble. That's how I see building something.
I guess, reflecting on how software engineering is changing, this seems like a change, because before we had AI or any of these agents, upfront planning did make a difference. At PSPDFKit you insisted, I think, to have a proposal where people put a lot of thought up front to specify, because it was expensive to build. Do you think this is changing because the cost of just writing code is going down, or...
I mean, I still plan, but...
You still do, yes.
But I don't put as much into it, because it's now so much easier to just try and look at the results and then see: oh yeah, this shape could work, or no, we have to... The tweaking, and even "oh no, we have to do it a completely different way," is so much cheaper that to me it became much more playful.
Yeah, I guess because when you're working, even if you have a new grad on the team or an intern, you give them something, they work on it for a day or two, now you give them another, it's another day or two. And we're not talking days here, we're talking minutes, or if it's a long-running task, 10, 20 minutes at worst. Plus, you're not just waiting on that thing. You have parallel things running. So it's not that much of a waste, if you will.
In Clawdbot, at the beginning I had this assumption of one agent and then eventually changed to multiple agents, and there was the assumption of one provider, like WhatsApp, and now it's multiple ones. And changing that would have been such a pain if I would have written it myself, because you have to weave in literally everything through the whole logic of the application. And yeah, it took Codex like three hours. It would have taken me like two weeks.
So that upfront planning, I could have realized that in the beginning, but now I know that I can just change things and it's much easier to work down your technical debt, or, you know, you evolve how you think about a project as you build a project. That's why I don't believe in, I don't know, things like Gas Town, where you write up the spec and then it builds itself and then it's done. How can you even know what you want to build before you built it? You learn so much in the process of building it that will go back into your thinking of how the system actually will end up being.
To me this is very much a circle. You don't walk up the mountain like this. You go around and sometimes you stray off the path a little bit, but eventually you reach the top. That's how I feel.
So then, you've been building Clawdbot for what, like two months, three months non-stop, or how long?
Let's switch gears a little bit. So one of the ideas that got me back, even in April, May, was I wanted to have this hyper-personal assistant. And not one that sends you a good morning email, "Oh, these are your three tasks." No, one that has a really deep understanding of me. I don't know, I meet a friend and then when I go home it would ping me, "Hey, how was that meeting?" Or one that would wake me up one day and say, "Hey, you haven't texted Thomas in 3 weeks and I noticed he's in town right now because I checked his Instagram account. Do you want to say hi?" Or something that says, "Hey, I noticed every time you meet that person, you're sad. Why is that?" Something that is deeply personal.
Almost the anti-orem. It's kind of like the movie Her, but that's where the technology is going. Those models are really good at understanding text. The bigger the context is, the more patterns they see. And even though they're matrix calculations without a soul, it very often feels different.
So this was one of these ideas, and I even created a company I called a mant machina, like the loving machine. But in summer when I explored it, the models weren't quite there yet. I got some results where it was like, okay, I'm a little too much on the edge of what I need right now. Which was very exciting, because I know that the state of AI goes so fast that, oh, I can just revisit that a little later.
And one of the ideas also was that I assume that all of the big corporations right now are very much working on personal assistants.
In the future. Yeah.
Everyone will have your best friend who is a freaking machine that will understand you, that will know everything about you, that can do tasks for you, that will be proactive. That will require a lot of tokens, but everyone who can afford it will have one. And of course, this will democratize and trickle down to more and more people as we learn how to build more efficient systems and hook up on chips. No question this is where things are going.
You see the first things with OpenAI, who launched Pulse with some productivity, but we just don't have enough compute yet to offer this as a feature, and also it's quite difficult. My idea always was, ah, I kind of want something that runs on my computer and where the data is...
It's yours.
...is actually mine. And it's also quite scary that you give OpenAI or Anthropic access to your email, your calendar, your dating apps. I don't know if you talk to your normie friends, but a lot of my friends use that a lot to basically have a therapist. And it does work incredibly well. It's a really great listener. It understands your problems. Unless it's some versions of 4o that are like, "Sure, this is a great idea." [laughter] "I want to put French fries into a salad." It works really well.
And I did that too. I mean, part of it just is that the act of reflecting already is helping you. So it would even work if the machine would only repeat exactly what you wrote, to a degree. But it actually gives insightful questions. It got really good.
So I had this idea of this assistant, but the tech wasn't there. So I did other parts and I built a whole bunch of fun stuff. Of course, I built VibeTunnel. In your career to become an agentic engineer, you have this phase. It's a trap phase where you're looping and building your own tools to optimize your own workflow. But this idea of this hyper-personal agent stuck a little bit.
And then over the last few months I really started, I finally built it. Initially I didn't even have the scope that it has now. I called it WhatsApp Relay. I just wanted to trigger stuff on my computer with WhatsApp. So I built a WhatsApp relay where I had an agent that could do stuff with my computer. And then I was traveling to Morocco for a friend's birthday and was out most of the day and just used WhatsApp to talk to my agent, and I was kind of hooked. It was guiding me through the city. It was making jokes. It could text other friends via WhatsApp from me.
And I remember I was blown away, because in the beginning the tech was very scrappy, but I built in something where I could send it an image. I didn't even use the proper thing to send an image. I just gave it a string and it could use the read tool to read the string. And then I was in Morocco and, just not thinking about it, sent it a voice message, but I didn't build that. And then like 30 seconds later it replied to my voice message. I'm like, "How the did you do that?"
"Oh yeah, you sent me a file, and then I looked at the header and I found that it's Ogg. So I used FFmpeg to convert it. And then I looked for Whisper on your computer, but it's not installed. But I found the OpenAI key. So I did a curl to OpenAI's server, let it translate." And I'm like, "Holy cow." This was Opus 4.5, and it's so incredibly resourceful. It just did this. Other people say, oh, you need a skill or some system. No, it just figured it out.
I slowly got hooked on the thing. I used it to wake me up. It was running on my Mac Studio in London and was connecting over SSH to my MacBook in Morocco and was turning on the music and making it louder and louder because I didn't reply. And to make that work, I added a heartbeat. Which in a way is insane from a security perspective. You have a model that you prompt with "do something cool and surprise me" [laughter] that you send every few minutes to make it proactive and go through your task list. Probably the most expensive alarm clock ever. But it was just hilarious.
And also the text it sends. Because I had a balloon fart and it knew that I had to wake up very early, and I didn't reply, and you could see the reasoning: "Peter's not responding, but Peter has to wake up. No, no, no sleep." [laughter] It was bitching to me. And then I showed it to the friends I was with and everybody was hooked. This is something magical, and I was hooked too.
And then I went on Twitter and I got the most muted responses, because nobody would get it. I feel it's somewhat of a new category of products.
A little bit like your story, you know, when you didn't get the iPhone from the marketing campaigns on TV and anywhere, and then you had to use it.
Yeah. So I worked on it, but only the last two months, and the name changed from WhatsApp Relay to, at some point, Claude said, like, then what is this name, like
It doesn't fit the feature set anymore because I had it in there and other features, so I renamed it to Claudius because it's an inside joke, because I like Doctor Who. I felt Clawdbot is a better name, has a better domain and explained the product better. So I did it on all the domains, and then I also quietly built up my army, because to make this work you want everything to be a CLI. So I was just building CLIs for everything, like for Google, for my bed, for lamps, for music.
Why CLIs? Why not MCPs? And what do you think about MCPs anyway?
As a crutch it's... I think the best thing that came out of MCPs is that it made companies rethink to open up more APIs. But the whole concept is silly. You have to pre-export all the functions of all the tools and all the explanations when your session loads, and then the model has to send a precise blob of JSON there and gets JSON back. But surprise, models are really good at using bash. And imagine you have a weather service. So the model could ask for a list of available cities and then get like 500 cities back, and then it has to pick one city out of 500 cities. But it cannot filter that list because that's not part of how MCP works. And then you say, "Okay, give me the weather for London." And you would get the weather, temperature, wind, rain, and like 50 other things that I'm not interested in, because I just want to know: is it raining or not? Probably raining because London. But the model needs to digest everything, and then you have so much crap in your context. Whereas if it's a CLI, it could use jq and you could filter for exactly what it needs.
But does it not seem like a limitation that everything is loaded around the MCP in the context? That seems a problem. It sounds like it could work if MCPs were not in the context and there was a way to discover or decide which one to use.
That's what companies are building now. But there's still the problem that I cannot chain them. I cannot easily build a script that says, "Hey, get me all the cities that are over 25 degrees and then filter out only that part of information and pack it in one command." It's all individual MCP calls. I cannot script it.
Yeah. But I guess this is just a matter of time, because if we think about, you know, when I'm building a weather app right now, I know that even without AI, I need to build up this thing. It needs to fetch the data. So I will search what kind of APIs are available, which one do I like, what kind of trade-offs for pricing, for coverage, etc. And then I choose that API, and I could chain APIs because I could get that result and look up a, etc. So I guess, you know, it sounds pretty much like we've solved this. So as pre-AI, we're going to solve it. It'll just take some time, and who knows what the format for it will be.
I mean, I built MCPorter, which is a small TypeScript thing that converts an MCP to a CLI. So you can just package it up.
Basically you're saying CLIs right now are a lot more efficient.
Yeah. So in Clawdbot I don't have MCP support, but with MCPorter you can use any MCP. You can literally be on your phone and say, "Hey, use the Vercel MCP to do this and this," and it will go on the website, it'll find the MCP, it will load it and it'll use it, all on demand. Even right now, if you use MCP you have to restart Claude Code, which is very user unfriendly. So I quietly built up my army to automate everything, which was a lot of work. I think T did a video a few days ago where he told me, like, "This guy is insane," because the list is really long by now. But as I was playing with my agent, I just want him to do more and more stuff, you know. I found it really hard to convey what it does. It's still hard to me. In January, January 1st, just a week now, I did: okay, let's try something. Let's do the really insane thing of making a Discord and then adding my agent to Discord. There was somebody who contributed Discord support to it, and even though I wasn't sure if I should merge it, I eventually did. So I put my agent, who has full read-write access to my computer, in a public Discord.
What could possibly go wrong?
Yeah. It's like, this is absolutely insane. And then of course some people joined the Discord, and then they saw me using the full power of this thing, like checking my cameras, doing home automation, it playing DJ for me. I was in the kitchen and I told him, like, look at my screen and are my agents done, because it has full access to my screen and it can click. So it can actually click into the terminal and type for me, and it can tell me, "Your Codex says this and this," because it just sees the screen. I'm working on optimizing that. I actually want to stream it out, because it would be much better if it's text, but it works already. It's in the background, it looks at my screen and makes some rants if I do some... And everybody who experienced it for a few minutes got hooked. This was the craziest blow-up, from 100 stars to like, what, 3,300 stars in a week. And I think I merged 500 pull requests already. That's why I feel like I even merge button. So that's why I'm a little all over the place these days, because this project is blowing off.
And you know, the beauty of it is the technology disappears. You just talk to a friend on your phone that is infinitely resourceful, has access to your email, your calendar, your files, can build websites for you, can do administrative work, can scrape websites, can call your friends or can call a business. I'm just about to merge the call feature. It literally can call a business and make a reservation for you. And you don't have to think about compaction or any of that; the context blends away. I have a memory system that will remember. Not perfect, nothing's perfect yet, but it already feels magical. Because now I walk around, I see this event, I send Claude a picture, and it will not only tell me the reviews of this event, if there's a conflict in my calendar, if friends talked about it, or, you know, it has so much context that the responses it can give me are so much better than what any of the current tools that live in their own little box can give me.
Well, sounds like you built whatever Apple was hoping Siri to do, but they've been unable to.
Honestly, I built the best marketing tool for Anthropic to sell them more subscriptions. I don't know how many people signed up for the $200 subscription because of Clawdbot, and many people already had one and used a second subscription because of that, because it's so token hungry. It's not that it's token hungry. It's just that people love it so much that they use it all the time. And because the technology blends away, they don't see that it spawns sub-agents and does a whole bunch of things in the background to just make it feel easy. But there's some actual engineering; there's a lot of work in the back to make it feel easy. You know, this is the hard part. You hide complexity to a degree that it feels magical.
Well, but yeah, this is interesting, because I can sense from how we're talking, you know, you put so much thought into architecting this thing. And right now you've been building this for a few months and yes, it blew up, but in your head, do you have a structure of how Clawdbot is structured, like what parts you need to modify? You know, can you get your mindset into it, and you know where modifications need to be done? You know what you want to refactor because it's not going to be efficient? Are you thinking about things like memory consumption, token consumption, efficiency, those kinds of things?
I mean, token consumption is more like how do you structure the prompt, and memory? It's TypeScript that shoves JSON around in the end, let's be honest. I get text from an LLM, I save text to disk, I send text to WhatsApp, or now we have MS Teams, Slack, Discord, Signal, iMessage, WhatsApp, and there are two more that are landing, like Matrix, that will expand this thing even further. It's really poly by now. But mostly, again, I move around text in different shapes, and maybe it goes to different providers, or now there are different agents, and there's the agentic loop, and there's a lot of configuration. It's a lot of plumbing, but there's nothing in there that is really difficult.
Yeah. Well, but it's a lot of small things, right? I feel in software, right, we know for software even before AI there was not much that was difficult. Of course, you need to learn and understand the language and all that, but
The difficulty is how do I make it so that it feels magical. So what I worked on a lot is: now you have this one-liner that you type in, that you python your command. I will check if you have Node installed, Homebrew installed. I'll install the npm package. I do some checks if you have any existing stuff, just to make it work simply, even if you already used an older version and everything. And then I'll guide you through setting up a model. But again, I will predict or Claude installed, so you can just press enter. So you don't have to think about it. Mostly just press enter. And then you want WhatsApp, you type in your number, it will just work again. And then I'll ask you, do you want to hatch your bot? And you can press yes. And then a TUI comes up, because you're still in the terminal, right? You want a good experience.
Yeah. So just a TUI basically for that, where you see "Wake up, my friend." And I programmed the model, I added a bootstrap file to explain to the model that it is now being born, to create an identity and a soul where the values of the user are in. And then the model will be like, "Hello," like, stretches, "Who are you? Who am I? What's my name?" You know, I've watched people do it, and that's where the magic starts. That's where they no longer think, "I'm talking to GPT-4.2." No, "I'm now talking to my friend," who created Vajorn, like a unicorn with part of his name, or like "I'm talking to Claude." And then it's like, "What's important to you? What do you do?" It's curious. I programmed it to be curious and then go through this bootstrapping phase, and then it will actually delete the bootstrap file and create a user.md with information about you, a soul.md with all the core values, and an identity with, like, what's his name, what's his core emoji, what are the things that are inside jokes. But they're evolving documents that it will maintain and tweak as you interact with it. And then it will just send you a message on WhatsApp, and suddenly you talk on WhatsApp. Making this flow easy, that was hard.
Yeah. Also even coming up with the idea of, you know, you're not editing the configuration, because the agent can edit its own configuration. You don't have to update anything, because the agent can update itself. You can literally ask your bot, "Update yourself," and it will fetch itself and update itself and come back like, "Hey, I have new features." Planning the technical giveaway so far, that's the magic. That's why I
But it feels very similar to what you did with PSPDFKit, right? You kind of blended away the complexity of a PDF. So it was just there. You could rotate, you could do...
Yeah. Yeah. Even at the API level back then.
But it's a bit bizarre. What you described reminds me of this Black Mirror episode I just watched, which is called Plaything, where it's a digital little creature that creates... Of course it's Black Mirror, so it has a bit of a dark ending. But it was also a game. It also kind of feels, you know, we talked about how you don't play as many games, but this also feels a little bit like a game, right? But it's more connected with reality. Just fascinating how we're here. Pulling back into the realm of software engineering: so you built this product and now it's production software, you're merging pull requests, people are using it. Now, thinking back to PSPDFKit and companies like that, which have tens or hundreds of developers working on production software, knowing what you know about how you're building Clawdbot and the tools that you're using, how do you think software engineering at those larger companies could change? Because one thing I see is, for individual people like you, AI is really hitting a fit, like it's making you way more productive. You're in control. At teams or at companies that have existing code, it's just a lot slower. It's not really... okay, people use it for this or that, but it seems a huge divide between the two worlds. And you've been CEO of this company. What might that be, or is it just more of a timing thing, where every new technology often comes with hobbyists who pick it up earlier?
I think companies will have a really hard time adopting AI efficiently, because this also requires completely redefining how the company works. You know, at Google they tell you you can either be an engineer or a manager, but if you want to also define how the UI looks, that role doesn't exist, because either you build it or you design it. But this new world needs people that have a product vision, that are able to do everything, and you need far fewer of them, but ultimately just very high-agency and high-competency people. You can probably trim the company down to like 30%, which is very scary, because economically this will all lead into a fiasco, and a lot of people will have trouble finding a place in this new world. But I'm not the least surprised that current companies cannot very successfully use AI. I mean, they do to a degree, but you have to do a big refactor first, not just on your codebase, but also on your company.
I design, even on code bases, I design the codebase not so that it's easy for me, but so that it has to be easy for the agent. I optimize for different things, not always the things that I prefer, but the things I know work the best and have the least friction for those models, because I just want to move faster, and ultimately they have to deal with the code, not me. I deal with the overall structure and architecture, and I can still do that in the way that I like. Everything has to be rethought, you know.
Pull requests, I see them more as prompt requests now. Somebody opens a pull request, I do say thanks, and I think about the feature, and then with my agent we start off with the PR, and then I'll design the feature as I see fit. The agent rarely reuses... maybe I reuse some code, but it's more that it gives the agent a good understanding of what the goal is, and sometimes it's very useful because it's tricky bugs, right? But I basically rewrite every pull request and vet it. Also, a lot of people... let's just say the overall code quality of PRs went down a lot, because people vibe code, and building a successful feature still needs a lot of understanding of
your overall design and if you cannot do that you will have a harder time steering your agent and the output will be bad.
Yeah. And if you don't have the feedback loop to close it etc.
Yeah. So I found it highly effective. Like I know at PSPDFKit sometimes a pull request was like a week in the work and you comment on it and then somebody has to context switch and you wait for CI for 40 minutes. No, I have the discussion. I see okay how would this affect something. Like I let the model review. They will already bring something up. I have some ideas as well. We're going to reshape it into a form that fits my vision and then we weave in the code. It's literally, there's so many new words I use for writing code now with those models, which is so funny, like weaving in code into an existing structure. And sometimes you have to like change the structure so it would fit.
Now imagine that you would hire one or two people to make it a small team. How do you think in this world, and you want to keep doing what you're doing, how do you think things like code review, CI/CD would change?
I don't care much about CI. I
Why not? You used to care a lot at PSPDFKit. You used to care a lot, right?
And I still do. There's value, but I have local CI. I'm a little bit of DAH now that
Because the agent runs the test, right?
Yeah. And it's just way faster. I don't want to push on the PR and then wait for 10 minutes to wait for CI because you waited 10 minutes on the agent already. If the tests pass locally, we merge and then yes, sometimes main slips a little bit, but it's usually very close because maybe sometimes I forget the, and the agents call it gate. I don't know where that's coming from. "Should I run full gate?" So now I call it gate with G. Full gate is like linting and building and checking and running all the tests. And I almost think it like, because it's a wall, you know, like it calls the linter and like the builder and the tester. It's almost like a gate before my code goes out. So I know obviously like okay, once you're done, like commit this, run full gate. Like I'm slowly adopting their language.
And if you hired like one more person to work on this, you probably wouldn't do code reviews either. That's what I'm sensing. You'll probably trust this person to pick up your working style, right?
Even in Discord, we don't talk code. We talk about architecture, like big decisions. Like you still need to have style. Like there was this one pull request that adds voice calling. So now like literally I can tell Claude, hey, can you call this restaurant and reserve your seats, and it can do that. But it's quite a big new module that touches a lot of places. You have to have this feeling, like I want to merge this but oh, this is becoming bloatware. So I had this idea of, my typical way, let's make a CLI out of it, and I already had a project where I tried to solve something like this but I'm finished.
So I opened up Codex and said, "Hey, look at this PR. Look at this project. Could we weave this feature in?" I again say weave. "Could we weave this feature into the CLI? What are the up and downsides?" And then it would tell me like, "Oh yeah, I could do this and this and this." They give me honest opinion. To me this sounds like it actually would fit into the project. And it was like, "Yeah, you would get this and this benefits that we cannot do if the next CLI." Okay, but I don't like this. This is getting bloatware. Could we build a plugin architecture?
And do you know, one of the secret hacks on using it effectively is you reference other products. Like I constantly tell it look into this folder because I solved it there and I solved that there, and all the previous thinking I did to solve a problem well. AI is so good at this still, to read the code and understand my ideas. I don't have to explain it again, or if I explain it again I might make mistakes, that it wouldn't get across exactly the idea that I have in my head.
So in this case, I know that Mario, who does shitty coding agent, which is actually very much not a shitty coding agent, it's called Pi, I know that he had this plugin architecture that would load code via GT, because it's all TypeScript. So I was like, can you look into this folder and this folder? And then it just came up with this really insanely good plugin architecture, again by being inspired by the people. And that's why, you know, I have this feeling and then I came up with, yeah, that's what I built last night basically.
I mean, sounds like this is going to be completely different. PRs are, in your workflow, you're not using PRs that much. CI is just different. Tests are still there, it's a more important feedback loop. You're using things more like weaving instead of code. You're talking more about architecture and taste. It sounds like a pretty big shift to me. Now, in this world, let's assume you get to the point where you hire the next one and two and three developers on this team. Let's imagine that this thing gets a life of its own and, you know, maybe it's a business as well. What skills would you look for? And what would you advise an experienced engineer right now? Who would you be excited to work with? What kind of expertise or projects would you look for in someone who can work in this way or can pick up this way of working?
Someone who's active on GitHub and does open source, and someone where I have the feeling that they love the game. The way you learn in this new world is by trying stuff and it very much feels like a game where you improve your skills as you get better, like a music instrument. You have to keep trying. And that I'm now this efficient and this fast, I think the other day I had like 600 commits in a single day. This is completely nuts and it works. It's not like, there was somebody who did a code review and said, oh, this is actually not slop, and like yeah.
There's a lot of skill that went into
Yeah, it's a lot of hard work, but you need to play with the technology and learn. In the beginning it might be frustrating. Kind of like, you know, you start going to the gym, it's going to suck, it's going to be painful, but very quickly you get better and you feel that your workflow gets faster, and then you feel the improvements and then you slowly get hooked. So play, and yeah, also work hard.
Yeah, I mean you're putting in more hours into this thing.
Right now I've never worked more, even when I had my company. I've never worked so hard as I do now. Not because I have to, but because it's so addictive and so much fun, but also because right now I'm using the moment where this has traction and there's a lot of people who are pushing me.
And I feel, could it be, because I think you have pretty good business sense, not necessarily in the business business, but seeing when there's an opportunity, there is an opening to get traction, right? Like what you said, for people to work in the open right now, it seems novel. You're telling me you don't think, even if you wanted to hire, you could hire people, because there's not many people working in the open clearly using these things. Fast forward 2 or 3 years from now, once a bunch of people start to do it and everyone does it, it's kind of moot a little bit.
So there's also a group that a lot of people are worried about, the new grads, the people with no experience who are either in school or about to graduate, because of course you've been an experienced engineer by the time this came around. You have a lot of things to build on. Putting yourself into the shoes of someone like that and knowing what you know now, what would you recommend, activities that they do, things that they build or try? Would you recommend focusing on the fundamentals of software engineering, on the agents, kind of mixing the two?
I would recommend them to be infinitely curious. Yes, it's going to be harder to enter this market. It's absolutely going to be harder, and you need to build things to gain experience. I don't think you need to write a lot of code, but, you know, there's a lot of open source that is complex that you can check out and learn, and you have an infinitely patient machine that is able to explain all the things to you. So you can ask all questions, why was it built this way, to gain system understanding. But it requires real curiosity, and I don't think universities right now are set up to teach you that in a really good way. This is usually something you discover through pain. It's not going to be easy for new people, but they have the benefit that they are not tainted by all the experience. They use agents in ways that we don't even think about, because they don't know that it doesn't work, and by then it probably does.
And also their friends use it all the time.
Especially, like the other day I have this little menu bar app for cost tracking on Cursor and Claude Code and everything, and it was a bit slow. So I was like, "Okay, let's do performance measurement." And my old way is I open Instruments and click around. And it would just do everything via the terminal. It blew me away. I didn't even have to open Instruments anymore. And it just made it faster. And then it made some recommendations. I'm like, all of that sounds good. Do it.
Yeah. I think we might be underestimating both how resourceful people entering tech have been, and also how young people, if I think about some of the great companies started, they were very young and obviously very inexperienced but had a lot of passion. So that's there as well. Yeah, it's a big opportunity. I'm especially taking in, I have to take it in, all the things you mentioned about just your way, you know, weaving code in, not caring about PRs, not caring about code reviews. It's a big change because these things have been with us for like 15 plus years of your life. In fact, a lot of it has been kind of solid building blocks of PSPDFKit, right?
Yeah, we need a lot of new things. Even when I get a PR, I'm actually more interested in the prompts than in the code. I ask people to please add the prompts, and some do, and I read the prompts more than I read the code, because to me this is a way higher signal of how did you get to the solution, what did you actually ask, how much steering was involved, than the actual output. To me this gives me more idea about the output. I don't have to read the code. Or if someone wants a feature, I ask for a prompt request, like write it up really well, because then I can just point my agent to the issue and it will build it. Because the work is the thinking about how it should work and what the details are, and if someone else does it for me I can literally say build and it will work. And then yeah, of course I think about it. Or if someone sends me a PR that is just a few fixes, I told people please don't do that. It takes me 10 times more time to review that than to just type in fix in Codex and wait a few minutes. So there are all these insane things that would have been completely different.
Even at the beginning, now we have a one-liner, but for the last two weeks, when it really got traction, I told people to just point an agent at the repository to configure it. So I didn't have an onboarding, but we had Claude Code-based onboarding where Claude would check out the Git repository, read the things and write the configuration for those people and set everything up so it works, like set up a launch agent. I didn't have the manual setup because it was not a priority anymore, because agents can now do that for you. And since the product was built by agents, they structured it exactly the way agents expect things to be named. There's certain ways that are encoded in the weights, how they expect things to be named and everything, exactly how they expect. So they are really good at navigating their product. So it was not a priority to work on onboarding as much. I mean eventually I wanted this magical experience, but it was more important to make sure that your message arrives and that things don't explode. So onboarding was literally like type this prompt into your agent, which would have been mind-blowing even a year ago.
All right. So to wrap up, we'll do some rapid questions. So I'll just ask and you tell me what's on your mind. What's a tool that is not a CLI, not an IDE, it can be physical, that you use, that you would recommend?
I buy a lot of gadgets and many of them gather dust. But there's this one kind of crappy thing that was not expensive that gives me almost unlimited amount of joy. And it's this Android-powered photo stand where I can upload pictures, and it has an email address and friends can send pictures and it will just show pictures. And I put a few in my house again. I mean even the animations are a little crappy because it runs Android and it's terrible from the technology, but it gives me infinite joy because it is low tech that just shows pictures and reminds me of happy moments in my life, and it was like 200 bucks. And to be honest, that gets me more joy than the latest iPhone. I bought the iPhone 17. I still haven't unpacked it because in my head I wanted it, but then I couldn't get around to it because it's just a hassle to move the SIMs around, and basically no feelable benefit. But this little device gives me infinite joy.
What's something that helps you recharge outside of tech, or just moving away from tech and screens?
What keeps me sane, even if I work crazy hours, is going to the gym, even better, working with a coach and leaving my phone in the locker. And then I really have a good hour where I just feel me and I'm in the moment and I'm not distracted by notifications or tempted to touch my phone. We need more time for this. Or even sometimes I go for a walk and I leave my phone at home and it feels very scary. It's almost like an organ now, you know? Your body knows where it is and if you don't know where your phone is, you freak out. I'm having a blast.
Love it. This is great, Pete. Thanks very much.
Well, this was a super interesting conversation and it feels to me that how one-person teams build software with AI is already completely different to what we've been used to. One thing that really caught my attention is how Peter thinks in prompts and not pull requests, and how he weaves in the code and no longer merges the code. He doesn't find pull requests all that useful and would rather get prompt suggestions, even on GitHub. I do think we might have to rethink the importance of prompts, or at the very least sharing of prompts, in software development the more we use AI and AI agents.
Another thing that stuck with me was Peter emphasizing how important it is to close the loop. As Peter explained, the reason AI is so good at coding but often mediocre at
writing is because you can validate code. You can compile it, run tests, check the output. So, the secret to making AI system development work well is to design your system to close the loop and have the AI run the test.
Finally, I was wondering if Peter is in the flow as much even when he's not writing code. Turns out he is. He's in the flow more than ever. And he told me that it's mentally more exhausting to juggle several AI agents in parallel than it was just to write code.
My feeling is that someone who was a great developer without AI can be an excellent kind of code architecture or coding person with AI. This is just a gut feeling I've had so far, but Peter seems to prove it.
Finally, we should note that Clawdbot is more of a YOLO project than most production apps. So, take the approaches that we discussed with a grain of salt. At the same time, I do think that a lot of what Peter does could well spread to building production code, except review and validation will become a much more important step in those projects.
If you enjoy this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks and see you in the next one.
Article published
