Matt Pocock on Grilling Your Agent: Why Old Software Fundamentals Matter More with AI
The Pragmatic EngineerMatt Pocock built Total TypeScript into a multi-million-dollar course business. More recently he created a set of AI agent skills, including "grill me" and "wayfinder," that spread quickly among developers. In this conversation on The Pragmatic Engineer podcast, he explains how he got into tech without a computer science background, and why his search for better ways to work with AI coding agents led him back to decades-old books. His argument is that AI has taken over much of what he calls tactical programming. That leaves humans responsible for strategic thinking, and the classic literature on software design turns out to be a practical way of steering agents.
From voice coach to self-taught developer
Before he became a developer, Pocock spent six years as a voice coach. He started as a singing teacher while at university, set up his own small company, and later did a master's degree in voice. After that he taught Shakespeare at drama schools, taught accents, and coached public speaking, including a couple of large engagements teaching consultants to deliver speeches. He says he left because doing the work at a decent level meant living in London. He tried it for about two years and hated it. He wanted to go back to the countryside where he grew up, so he taught himself to code in order to have work he could do remotely.
He had already been writing software for his students. There were small flashcard apps, and his first project was, in his words, the most ambitious thing he has ever attempted: a web audio analyzer that displayed a spectrogram of a student's voice so he could see which resonant frequencies were present. He admits it "ran terribly," but it made his lessons a little better. He looked at job postings, decided his JavaScript and CSS were enough to get started, quit his job, took a couple of months off, and was hired around 2017. He notes that getting a job in the UK was easier then than it is now.
He calls his communication background an "unbelievable advantage." In interviews he could come across as a reasonable person, unlike some candidates straight out of CS degrees. He had very little technical knowledge at first but could explain technical ideas to people, and once his knowledge grew he moved up quickly across several companies. His first employer was a tiny company with a couple of senior developers he found inspiring. One of them, who wore sandals and had lived on a canal boat, had Pocock install CentOS 6 on his Windows PC on his first day, because that was what their application ran on. That company ran into financial trouble, and Pocock then moved through several agencies.
TypeScript as a coordination tool
He was already arguing hard for TypeScript by his second job. The agency was building a learning management system for a car manufacturer. The front-end team was small and full of bugs. The back-end team, based in Portugal, kept changing API contracts without warning. The front end adopted TypeScript to keep the two sides in sync. Pocock says their velocity rose so much that they overtook the back-end team, and eventually people were moved off the front-end team because it was working so quickly. He describes this as his origin story with TypeScript.
Open source, XState, and the Stately job
In his fourth job he worked on a complex application where two people on a video call could move around a house together in real time. It involved a lot of state that had to cross the network boundary. He used XState version 4, the state machine library, and considered the project a success. He then began building tooling to make XState more type-safe, including a CLI. That caught the attention of David Khourshid, XState's creator, and Pocock joined the core team.
He says the work put him in contact with developers at a level he had never seen, naming Khourshid and Mateusz Burzyński (Andarist). When Khourshid raised funding for Stately, a company betting on state charts and visual programming, Pocock was hired. It was his first job paid at American rates, which he says changed his life and how he thought about money and flexibility. At Stately he also did advocacy work. He still thinks state charts are incredibly useful for certain kinds of work, although he has backed off his belief in them somewhat, especially now that AI is here.
Three days a week at Vercel, and the Total TypeScript launch
His advocacy caught the attention of Vercel's developer education group, which Lee Robinson led at the time. Pocock joined under Jared Palmer. While trying to force XState to be type-safe, which he calls "a mostly impossible job," he had learned many advanced typing tricks. He also missed teaching. He began posting two-minute TypeScript tips on Twitter and found they spread in a way nothing he had done before had. One Sunday he recorded 13 to 15 tips and scheduled them over the next few weeks, and his follower count went from about 4,000 to about 10,000.
That made him confident a course could work. He negotiated an unusual contract with Vercel: three days a week, initially for three months. Vercel had wanted him full-time. He admits it is slightly embarrassing to say, given that it is many people's dream job, but he treated Vercel as a stable 9-to-5 while he tested the course idea. During his time there he wrote some of the initial documentation for Turbopack and flew to San Francisco for Next.js Conf, where it was announced.
Around two months in, it became clear he couldn't stay. A pre-release sale of Total TypeScript earned roughly 30 to 40 times his Vercel pay, and he describes that pay as already very good. He says he loved working at Vercel and would probably go back at some point, but he had no real choice.
Building Total TypeScript and its business model
Pocock says he tries very hard not to work weekends and that the "996" culture "turns my stomach." He wants a life where he spends most of his time with his family, and he asks listeners to read his later decisions in that light.
He built Total TypeScript with Joel Hooks, who is known for Egghead and has worked on Kent C. Dodds's courses. The full course came out around January or February 2023 and, by his account, reached seven figures in revenue very quickly. That revenue was split between him and Hooks, and expenses came out of it. The host notes that he later publicly shared a milestone of $2.5 million in total revenue. For Pocock, this was the high-leverage work he had wanted since his singing-teacher days: do the work once, step back, and keep earning while spending time with his family.
He wants his income to come only from products people buy. He avoids sponsored content and has been trying to take down his GitHub Sponsors page. Many buyers pay from company education budgets, a market Hooks pushed him toward. Pocock says he had no idea how much money was "sloshing around" in the industry. They also offer an extended refund policy.
What AI did to technical education
Asked how AI has affected education, Pocock separates teaching into two layers. There is the "what," meaning syntax and knowledge, and the "why," which he calls wisdom. The first is usually the vehicle for the second. He argues that knowledge is now very cheap to acquire. He even has a "teach" skill that teaches you whatever you need. Wisdom, he says, has gotten no easier to learn, and you will hit the same problems without it whether or not you use AI.
He says Total TypeScript revenue has gone down, and he thinks people teaching that kind of material will struggle. It took him a long time to find his place in AI. He felt he lacked the credentials of an OpenAI researcher. Starting in late 2023, he made courses about building AI into applications, which seemed like a natural move for a front-end developer. He is proud of that material, but he did not see the returns he expected and decided it was the wrong bet.
The December shift: tactical versus strategic programming
The turning point was what the host calls the winter break when everyone came back "AI-pilled." Pocock associates it with Opus 4.5 and OpenClaw: people had time off, used the models heavily, and saw that something had changed. His conclusion was that agents had become good enough to delegate to. You could build structures around them so they handle the tactical work (syntax and knowledge) while you handle the strategic, long-term thinking.
He borrows the tactical/strategic distinction from John Ousterhout and says AI has "largely eaten tactical programming." That left room for a course aimed at the strategic layer. At the time he was also experimenting with Geoffrey Huntley's "Ralph loops," which repeatedly run an agent toward a goal. The host recalls a popular video of Pocock's that contrasted traditional top-down planning, where plans break once implementation reveals new work, with a Ralph loop driven by a markdown file of tasks that grows as the agent works through it.
Pocock says his two years building agents into applications were good preparation. Building those apps forced him to think at a lower level than harnesses usually expose: how data gets in, what shape it has, whether it goes in the system prompt or the user prompt, and how to compact or clear it. When he moved to Claude Code, which had been out for five or six months at that point, those questions felt familiar. The processes he built looked to him like finite state machines, much like his XState work, with agents firing events that send control back to earlier states. He ran experiments on his own projects, including a custom video editor he describes as a large codebase and several open-source projects. Sometimes he dropped his process and went back to the default setup, and he says he noticed "a huge difference."
Skills as a distribution mechanism
He settled on skills as a way to share these processes. He describes a skill as "just a folder of markdown files." Some are model-invoked, meaning the agent decides to use them. Others are user-invoked, meaning only the user can call them. He had seen skill sets like Superpowers, Claude Code plugins, and GStack, so he published his own and moved on to other work. When he checked back, the repo had more stars than anything he had ever made, with almost no promotion. At the AI Engineer conference in London in April he gave a talk called "Software Fundamentals Still Matter," which he says is now around 1.2 million views. By his count the repo has about 230,000 stars, making it the second most-starred skills repo and somewhere around the 20th to 25th most-starred repo of all time. He says this was the second time in his career he felt that kind of momentum, after the TypeScript tips.
His design goal is the simplest possible set of skills that people can easily audit and tinker with, which makes them more likely to be adopted at work.
"Grill me": getting the agent to interview you
The host describes being grilled by the skill while building a simple API endpoint. The endpoint needed to tell an authenticated caller whether an email address belonged to a newsletter subscriber. The skill asked 35 questions: whether to use a bearer token or put credentials in a JSON body, how strictly to enforce rate limits, and so on. The host found it annoying and impressive at once. It had been a long time since they'd had such an involved design discussion, and it forced them to make decisions and sometimes go research the options. The host adds that as long as people keep learning while using AI they will be fine, and that trouble comes when the learning itself gets outsourced.
Pocock says "everyone's got a grill me story." The skill is tiny. It just tells the agent to interview you relentlessly about the topic. He got the idea from someone named Tariq who works on Claude Code and suggested that having the agent interview you produces better results. He says the skill has "weird emergent behavior": models start thinking outside the box and suggesting ideas. It reminded him of design discussions with the senior developer at his first job and with Andarist at Stately. That led him to ask how to "tickle the right latent space" so the agent behaves more like a real senior developer, since better conversations lead to better output.
He argues that people underestimate the communication gap between themselves and the agent. The belief that you can "just trust the model" ignores the fact that no model, "even Mythos," can read your mind. If you just ask for code, you get something misaligned with your values, because the agent doesn't know what you consider important. So grilling is about more than implementation details. It establishes what is in scope, what is out of scope, and what matters. In his words, it is the agent getting to know you.
Smart zone, specs, tickets, and the night shift
Turning that conversation into code raised the problem of context size. Pocock credits Dex Horthy's idea of the "smart zone" and "dumb zone." Every token competes for attention, and as context grows the model loses connections and makes more mistakes. Pocock estimates the smart zone at roughly the first 150,000 tokens of current frontier models. He says the window size doesn't matter, because what counts is the raw number of tokens and attention relationships. Beyond that, performance slowly degrades.
The question became how to spread work larger than 150k tokens across several sessions. Ralph loops were one answer: tell the agent to make the smallest change that moves toward the goal, then clear context. State survives in the codebase and file system, not in the model. To make this more stable, he defined two kinds of documents. The first is a destination document, which he used to call a PRD and now calls a spec, and which says when the work is finished. The second is a set of tickets, one per session. His skills turn a grilling session into a spec, then turn the spec into tickets. A spec might cover 30 or 40 tickets, and an implementation loop works through them.
The loop is meant to run with the user away from the keyboard. Pocock describes "day shift and night shift": plan during the day and let agents work overnight, so you wake up to clean code to review. He was tired of constantly switching between terminals, which he admits he still does to some extent. He wants long planning sessions, a couple of hours of autonomous agent work, and focused blocks of his own time for other tasks. He says the December models were the first good enough to delegate to that way.
Wayfinder: grilling across many sessions
Planning can also outgrow a single context window. Pocock's example is "build me a Stripe clone," which he says cannot be planned in 150k tokens. He asked what each grilling session needs to work well: its own purpose, what has been decided so far, and what other sessions are in progress. His answer was a map. The map is a central record of decisions, with milestones and a "fog of war." Each grilling session reveals more of the map, and he describes the overall structure as a directed graph that you walk until you reach the destination.
Each session is a ticket on the map. Tickets aren't limited to grilling. Some are prototyping, research, or arbitrary tasks like provisioning infrastructure. He says he has had maps with 50 to 100 tickets. He also uses Wayfinder for non-technical work such as course planning and building a garden office, and says he is interested in how well these skills transfer to other areas of life.
Why agents are good at software, and bad at what isn't text
The host brings up Hillel Wayne's research comparing software engineers with traditional engineers. One difference Wayne noted is that software, unlike physical materials, behaves the same way every time you run it. The host wonders whether LLMs bring back the kind of variance other engineering fields have always dealt with.
Pocock's explanation for why agents do well in software is that everything is text. The inputs are code, docs, and instructions. The outputs are code, test results, type-check results, and lint output. Anything that isn't text is, in his words, "just garbage from the agent." A one-shot UI demo can look impressive, but if a hover animation looks wrong, it is hard to get that to the agent. You could record a video, but he says vision isn't good enough yet. He guesses that other professions will get better results if they can express their work as text, for example through simulations or something like a linter for architectural diagrams. He is currently connecting all the services he uses to agents.
Rediscovering the fundamentals
The host asks about the claim that AI requires setting aside prior knowledge. Pocock says he believed that at first. He tried spec-driven development, a term he has mixed feelings about because it covers too much, on the theory that "English is the hot new programming language": keep editing a persistent spec and let the agent update the code. He says the results were worse than coding by hand and weren't improving. Each cycle of changing the spec made the code worse. "You're not supposed to look at the code," he says, but he did, and it was garbage. He reasoned that a bad test suite gives the agent bad signals, just as it would a human.
He then opened The Pragmatic Programmer, which had been sitting on his shelf still in plastic. Its section on software entropy helped him frame the problem: agents produce software entropy faster than ever. He says almost every line read as if it had been written for today, citing "don't outrun your headlights," working within feedback loops, programming by coincidence, and tracer bullets. Because the book has been around for about 25 years, he figured its concepts were probably in the models' training data.
Tracer bullets addressed a specific failure. Even with Ralph loops, agents would build the whole database layer, then the whole application layer, then the whole React component library, and only connect them at the end. By then, decisions in the database had already affected the front end without any feedback. A tracer bullet, or vertical slice, builds a thin path through every layer first so feedback arrives immediately.
When he used these phrases in prompts, the agent began repeating them in its reasoning, saying things like "Okay, I'll turn this into a tracer bullet." He calls these "leading words": short phrases repeated a few times in a skill or prompt that change the agent's behavior. He began going through other books for more. From Ousterhout's A Philosophy of Software Design he took "deep modules," which he calls "a massive one" for him. The host compares leading words to professional jargon, which makes communication faster and reduces misunderstandings between experts.
Ubiquitous language and "Memento-driven development"
Pocock took the idea further. If established terms work, what about terms for his own application? Agents are very verbose, and he singles out Opus 5. So he turned to Eric Evans's Domain-Driven Design and its "ubiquitous language." He built a variant of grill me, which he admits is "terribly named" grill with docs, that develops a domain language during the grilling session. He says the difference is "night and day."
His example comes from an app with "ghost" and "real" lessons. When a ghost lesson inside a ghost section inside a ghost course becomes real, the section and course also have to become real. With the agent's help, that became the "materialization cascade." He says agents are good at coming up with such terms, and he has a separate domain modeling skill for it. When the vocabulary also appears in the code, the agent can find relevant functions with a simple grep. He has worked DDD into every part of his setup and says he already has The Mythical Man-Month too.
His overall framing is to imagine a coworker who wakes up every morning with no memory, like the protagonist of Memento. A human can work around a bad codebase by building up memory over time. An agent starts fresh every session, so the codebase has to be optimized for a permanent new starter. He argues that software fundamentals have always aimed at this, and that AI is mostly emphasizing rules developers already knew they should follow but often didn't.
How to learn strategic programming, and what juniors face
Asked how an engineer can find the fundamentals that matter, Pocock says strategic programming has always been hard to learn because the feedback loop is so long. Someone who leaves a job after six months may never see their strategic mistakes, which can take nine months to surface. He compares it to mixing music on a desk full of sliders, such as the number of deployable units (turn it up for microservices, down for a monolith), except that "you can't hear what's wrong until 9 months later." He thinks AI shortens that loop because the extra code brings mistakes back sooner. His advice is to treat code as the environment the agent works in, keep improving it, think at that level constantly, and read the books to gain the vocabulary.
The host wonders whether faster fixes also make lessons less memorable, since painful incidents, like being burned by a lack of idempotency, are what teach people. Pocock raises a harder question: if strategic knowledge is now so valuable, why would a company hire someone without it? He mentions that Uncle Bob, in a recent interview, suggested hiring juniors and treating them like agents, delegating tactical work until their mistakes come up. Pocock calls that a huge waste of money now that tactical work costs less than minimum wage in many countries. He says he doesn't know the answer. He only knows that strategic understanding is worth more than ever.
On a listener's question about convincing non-engineering stakeholders to invest in fundamentals, he says the same question could have been asked ten years ago about tech debt. He suggests measurement. Start with observability across every agent in the organization, tracking success and failure rates, which he says would have felt invasive for human developers but is fine for a paid service. Have someone analyze the data, find which repos do better, and spread those lessons. Organizations should share a common set of skills so people can contribute and A/B test different workflows. His own setup includes a loop that runs an "improve codebase architecture" skill every morning and proposes an improvement he can turn into tickets with one click. He floats spending something like 20% of time on "the factory that builds your software," and adds, with a laugh, that you might keep that work quiet until it has shown results.
Moving from local to cloud agents
The host quotes Pocock's tweet: "I'm moving away from my local dev setup. Makes zero sense to me now." He explains that people keep asking how to make skills collaborative, such as tagging a colleague into a grilling session. Every developer now has something like a hundred terminals, and he thinks those should be available to the whole organization, inside tools like Slack, Discord, Teams, or Linear. He doesn't work with a team himself, but he chats with his remote box over Discord, including on the train to this recording, to build course material and fix bugs students report. He can port-forward to see the dev server, and he views an expensive laptop running this work as wasted compute. The remote box is always on, so it can run schedules, including a morning stand-up where the agent plans his day using his Discord chats. The only thing he now does locally is debug the remote bot.
He adds that local setups have their own problems, like piles of Git worktrees and several Docker containers per worktree, and the cloud lets you provision resources on demand. The host mentions that some companies with platform teams have moved full dev setups into the cloud and see most developers choosing cloud agents voluntarily, though front-end work is often cited as an exception. Pocock suggests tunneling to the dev server might solve that, but says he hasn't tried it.
When to plan and when to course-correct
The host asks whether fast agents make upfront planning less necessary. Pocock says it depends on the size of the work and how hard it is to undo. If a large feature goes wrong, the wrong code stays in the agent's context and shapes everything after it, and realigning afterward is expensive. So align first. For a five-line change or moving a button a few pixels, align afterward. His rule is to "shift right" wherever possible. His video editor has a feedback button that creates a GitHub issue. An implementation agent picks it up, a review agent checks it, and he reviews the result at the end. His decision tree: if it is small enough to align afterward, skip grilling. If it fits in one session, use grill me. If it spans many sessions, use Wayfinder. The host compares this to how tech companies used PRDs: skip them for trivial work, share them for team-sized work, and require sign-off for large projects.
On the criticism that this is just waterfall again, Pocock points out that he does a lot of aggressive prototyping before writing specs. Agents make it cheap to produce three or four versions of something, pick the best, and iterate, and he calls prototyping essential to spec writing. The host cites Grady Booch's view that waterfall was only ever a problem when planning took a year and implementation took three. Pocock says he doesn't mind criticizing waterfall as a cautionary tale anyway, because cheap labor makes agile a better fit for agentic work.
TDD: wrong problem, useful evidence
Pocock has a TDD skill and recommends it, but says he has been rethinking it. As he describes it, TDD is designed for small working memory: a failing test tells you where you were after a coffee break. Agents have much larger working memory than humans, though not infinite, so he thinks TDD is "aiming at the wrong problem" for them. What agents do need is feedback loops, and writing the failure first is hard for an agent to fake. Even when he doesn't use strict red-green-refactor, he asks the agent to "provide proof" that a change works and would fail without it, which he calls TDD evidence. He also notes that agents often write tautological tests, such as asserting that a constant equals its own value. He still recommends TDD because of the confidence it gives humans, but says he is "starting to see the counter arguments."
Tech debt at any size
Responding to Jared Friedman's claim that large codebases no longer have to live with tech debt, Pocock replied that now you can have it "even in a tiny code base." He defines tech debt as anything that makes a codebase harder to change over time, and a good codebase as one that is easy to change without cascading failures, for example because it has a solid test suite. Agents can't think strategically, so they easily make codebases worse. Automated review, where one agent implements and another enforces standards and removes tautological tests, helps. But he asks how you would know whether the review agent itself is doing a good job. His conclusion is that this is a constant battle and the same conversation the industry has been having for 20 years, now with "this new elephant in the room."
Distance from Silicon Valley, curation, and advice for juniors
On living in the UK, Pocock says he has no ability to predict the future and no privileged access. He is "just a person in the field" and focuses on what works today, which he thinks has kept his scope narrow. He probably could do more from San Francisco, but he would have to live there, and he prefers being near his parents and raising his son in the countryside.
On education, he doesn't think the way people learn has changed much. An agent teaching you everything sounds attractive, but what learners want is curation. He sees knowledge as a dependency graph and his job as turning that graph into a sensible linear path, which he considers strategic work AI isn't good at. He says his pivot from tactical TypeScript content to the strategic layer is "working okay" for him, but he can't speak for others, many of whom he knows are not doing as well. He says the industry shifted in about seven months, faster than ever. He calls his own success mostly luck: Joel Hooks had to push him into AI, it took about three months of trying and failing before it clicked, and he has made mistakes too. If it hadn't worked out, he says he would simply go back to engineering, which he also loves.
For people starting out, he says he would love to be a junior now, and imagines going back to singing teaching just to build tools for students between lessons. He would use agents as much as possible, since that is how work is done now. He thinks his skills help because grilling gives you a senior developer to talk to and keeps you thinking about deeper ideas. His old spectrogram tool, with its "six nested for loops," would have been far better with an agent pointing them out. The key, he says, is caring about how you produce code as well as the code itself. It is a great time to be a "navel-gazing programmer," and the people doing well now are the same curious, adaptable people who did well ten years ago.
Gardeners and introspection
The host quotes Lauren, a software engineer, saying every team needs a gardener who watches incoming PRs and pulls weeds like creeping lint suppressions. Pocock had replied that teams need only gardeners, and adds with a laugh that you probably need a few other people too. He describes developers as "Ralph's platform team," people building the environment agents succeed in. Diagnosing entropy before it becomes a problem, and building loops where agents improve code based on bug reports and feedback, may be the essential skill now.
Asked what makes a great engineer, he points to Lars Grammel at Vercel, who is building a "software factory" to handle the flood of issues on the AI SDK. If he had to choose one word, it would be introspection: looking at how you work and putting it into words an AI can use. That is what his skills and automations are, he says. He keeps asking how he could do something better and "how could I encode this into this strange animal that I have in front of me."
He closes by recommending three books: The Pragmatic Programmer; John Ousterhout's A Philosophy of Software Design; and roughly the first three chapters of Eric Evans's Domain-Driven Design, for ubiquitous language, domain modeling, and encoding that model in code. He says he is less of a fan of the rest of that book.
It's annoying how damn it bro.
Everyone's got a grill-me story. It just has this weird emergent behavior where the models start thinking a little bit outside the box and they start throwing ideas at you.
I guess the idea of the tracer bullet is like a tracer bullet that leaves a mark.
I just started using these phrases in my prompts when I was talking to the agent and I started noticing that it was saying those phrases back to me. It was saying, "Okay, I'll turn this into a tracer bullet." This is what I call a leading word, where you lead the agent just with a simple phrase that you repeat. How do I go about and find those fundamentals that matter and go back to?
It's really tough because strategic programming has always been really hard to learn. I think of learning strategic programming as kind of like you've got a huge mixing desk in front of you with loads of these different sliders. You turn it up, you've got more microservices. You turn it down, you've got a monolith. It's kind of like you're mixing some music, but you can't hear what's wrong until 9 months later, until the mistakes come and get you.
How do you convince non-engineering stakeholders that investing in software fundamentals is important? You need some sort of metric for figuring this out. I think the first step to this is
I got the grill of my life when building a pretty simple API endpoint using the grill-me skill. It asked me 35 questions. I kid you not. It was intense and annoying and it forced me to think more. Today's guest is the creator of this popular skill, Matt Pocock. Matt is a developer turned educator, well known for his Total TypeScript series and now for his AI skills and educational videos. Today we cover Matt's unusual path into tech after years of being a voice coach and building his own DIY coaching software. Matt's popular skills, grill-me and wayfinder, and why these skills became so widespread. Taking inspiration from decades-old programming books to build better software with AI, and many more.
If you want to understand which software engineering fundamental approaches remain very useful when working with AI agents, this episode is for you. This episode is presented by turbopuffer, vector and full-text search built on object storage. It's fast, cheap, and extremely scalable.
This episode is presented by Linear, and I wanted to take you back in time to remind you how we used to get work done. Back when every line of code was written by an engineer like you or me, a tracker's job was to keep people in sync without slowing people down. Linear was built to be fast and low friction, and you could tell. In last year's The Pragmatic Engineer survey, Linear was the most loved tracker tool and Jira the most disliked one for its sluggish performance. And data coming from The Pragmatic Engineer audience showed how Linear started to gain traction against existing tools, especially at startups and mid-size companies.
And since then, Linear grew up. They added all the stuff that larger companies need to manage work: projects, initiatives, roadmaps, and customer requests. And large companies started to switch. For example, healthcare company Oscar helped move 600 engineers from Jira to Linear. OpenAI started with 100 seats and moved all 3,000 staff without any mandate. Coinbase, Cash App, Brex, and Ramp are all on Linear. Many of them saw Linear as a way to consolidate a single tool that brings planning and building together.
So now let's fast forward to today. When you have AI agents inside a company, those agents need context to work well. They need access to things like specs, customer requests, history. Oh wait, these are all already in Linear. So when agents arrived, Linear became the ideal context layer. Today, 80% of enterprise workspaces in Linear have adopted agents. You can use agents like Codex, Claude Code, Linear agent or your own agent. Coinbase and Ramp both built their own internal agents and describe Linear as a place that their agent goes and picks up the context before starting work. See how it works at linear.app/pragmatic.
Matt, it's great to have you on the podcast.
Great to finally be here. I'm a huge fan. I've watched so many of these. I feel like this is like the Tiny Desk of being a software engineer. You know what I mean? This is big stuff. So, I'm glad to be here.
And it's also great to reconnect, 'cause about a year ago, we had lunch at Microsoft Build as well, which was really fun. But now it's good to jump into this. And with this, I wanted to ask about your background. Unlike many people in tech and on this podcast, you didn't start out to study computer science, right?
Absolutely not. So for six years before I became a developer, I was a voice coach. I was a singing teacher working in London and working in Exeter, where I went to university. I was teaching accents. I was teaching singing. I was teaching voice. I did a master's in it. I spent a lot of time thinking that was what my career was going to be. You know, I didn't have any inkling of tech, didn't sort of think about it at all. I sort of ran my own website and stuff, but yeah, so I did that for a long time and it's been an extremely important influence on my life and I think my personality as well.
Can you get a bit deeper? Where did the voice come from and what do you do as a voice coach? Who are people who came to you for help and what kind of help?
So, I started as a singing teacher. I was in a band and stuff at university. I sort of had a bit of experience doing singing and so I set up my own company kind of at university and doing that stuff, and it was people who just wanted to sing better, who wanted to use their voice for choirs, who wanted to just do it as a hobby. It wasn't anything particularly professional. Then I went and did a master's in it and I started going to drama schools to teach people Shakespeare and stuff, and getting people in who wanted to do public speaking. I did a couple of big gigs for consulting companies, you know, going and teaching them how to deliver speeches and how to talk better. It was wild, you know, and the reason I got out of it was because I realized in order to do it at a decent level, you had to live in London. I didn't want to live in London. I tried it for like two years. I just hated it. I hated it. I didn't grow up in London. I wanted to get back to the countryside and where I was from. And that's what I did. And so I learned how to be a developer. I was essentially self-taught, in order to have something I could do remotely.
So basically you were looking at professions that you could do from outside of London that had a career or perspective or future.
Exactly. And I'd sort of taught myself how to build stuff and just sort of build basic stuff in JavaScript because I was interested in making my lessons better for my students. So I'd actually made sort of little flashcard apps. The first app I ever built was the most ambitious thing I've ever attempted. It was like a web audio analyzer, so I could analyze the spectrogram of your voice to see which resonant frequencies were happening, whether your T1 and T2 were properly balanced and things like that. Extremely in-depth, ran terribly, but actually, you know, made my lessons that little bit better. And so I was doing pretty hardcore stuff terribly straight away. And I realized, okay, I started looking at job postings and I thought, well, I could do a bit of JavaScript, I could do a bit of Sass, I could do a bit of bits and bobs. And I just jumped into it. I, you know, quit my job, had a couple of months off, and eventually got a job. This was about 2017, where it was a little bit easier to get a job in the UK than it is now. And I just went from there.
I guess in some ways you were also lucky because that was the peak. That was a time where demand was so high for engineers that people had boot camps with a few months of experience, and I think people got chances from a lot of places if they had the drive and the motivation and the smarts, right?
Yeah. And because I had this history of talking to people, that was an unbelievable advantage, right? I could actually go into an interview and sound like a reasonable person instead of someone who'd come straight from a CS degree who maybe didn't have those skills. So I had this bizarre ability of having zero technical knowledge, or very little in the beginning, but the ability to explain technical knowledge to people, right? And so basically all I needed to do was increase my technical knowledge a little bit, and I was very passionate about it and that increased quite quickly, and then it sort of seemed to be an unfair combination because I just rose through the ranks very quickly in various different companies, and I don't know, I felt different from the other software developers I was working with. Does that make sense?
And then how did you step up on the ladder? So you decided, I'm going to do this. You taught yourself. You went to some interviews. I'm assuming it must have been a small company, right?
Yeah. A tiny company with a couple of really inspiring software developers who worked there. Basically a guy, I won't say his name because he likes his anonymity, but basically a guy who lived in sandals, who lived in a canal boat for a long time, like long hair, proper hardcore, you know. It was around the time that Microsoft bought GitHub. I remember him coming in almost in tears.
Yeah. Microsoft hater.
Yeah, absolutely. You know, classic. I remember the first thing he got me to do was set up CentOS 6 on my Windows PC.
It's a pretty hardcore Linux distribution.
Really hardcore Linux distribution, because that's what our application was running on in the cloud or something, you know. So, really lovely, wonderful guy and someone who taught me a lot straight away. And so basically that company ran into financial troubles and so I had to move to an agency pretty quickly, and I got a higher job there. From there, nine months later I moved to another agency and then another agency. So just sort of bouncing around different agencies, and then I was working in open source, which is kind of the next part of the story.
And with the agencies, what tech stack were you using at the time?
Yeah, it was TypeScript, it was React.
Oh, was TypeScript already back then?
Well, I was pretty hardcore on TypeScript already, almost as in my second job. I think I was doing, you know, presentations on how important TypeScript was. We were working for an automobile manufacturer building a learning management system, right? You know, classic boring agency stuff, right? And the front-end team at that time was pretty small and we had a backend team in Portugal, right? So classic front end/backend split. The backend team were racing ahead, and at the time I joined, the front-end team was really slow. We had a ton of bugs. The backend team kept changing their contracts without telling us, and we thought we need something to link us up a bit better. TypeScript felt like the obvious thing, and once we shipped it, our velocity just went, you know, we were faster than the backend team, and eventually they took people off our team because we were so quick. So yeah, that was my history with TypeScript. That's kind of my origin story with it.
How did you get into open source? Was it at work? Was it on the side?
I'd been constantly playing around with open source on the side and I was interested in different things. By then I was into Twitter. I was sort of looking at people online and thinking that's someone I want to emulate, someone I want to look at. And there was a guy who crossed my radar called David Khourshid, who's the state machine and TypeScript guy on Twitter. A lovely, lovely guy, and I owe a lot of my career to him really.
And I was working on a project, this is I think in my fourth job, where we needed a state machine. It was a very complex application where you were on a video call with someone and you could navigate around a house in real time together using some sort of Matterport integration, and there was a lot of linking up that needed to be done across the network boundary, a lot of complicated state.
And so I used a library called XState at the time, XState version 4, I think. That was a resounding success. And so I wondered, okay, how can I make this more type safe? And so I started to sort of build some tooling around it, have a fiddle, built a sort of CLI that constructed around it, and that got me the attention of David, and I became a member of the XState core team. So I started contributing issues, started having discussions about the future of the library, and it brought me into contact with just a level of developer that I'd never seen before. David and another guy called Mateusz Burzyński, called Andarist on Twitter. These are the most talented developers I've ever seen. Like, this is another level.
And eventually David wanted to form a company out of it. He wanted to make a big bet on state charts and visual sort of programming as the future of development. They got some funding, and that was my first job where I was being paid American money basically, and it was a huge step up for me.
Yeah, which, as we know, is quite different from a European or local, even UK, company paying, because I also covered some of it in the trimodal nature of software engineering compensation, where US companies, especially in Europe and also in the US, think about compensation differently and value generated differently, right?
Totally. It changed my life, you know, in terms of the way I was thinking about money and the way I was thinking about flexibility, and it meant I was working on something I was passionate about. And I started, while I was there, doing a bit more advocacy for it, because obviously the company's very small. I was doing a lot of development, but also I wanted to be an advocate for it because I believed in it, you know, and I still think state charts are an incredibly powerful primitive for certain kinds of work. I've sort of rowed back a little bit on my belief in them, especially in the AI age, but I was doing a bit more of that, and that got me the attention of a couple of guys at Vercel, because at Vercel at that time, Lee Robinson was the guy in charge of developer education there. They had this incredible team: Delba de Oliveira, Lydia Hallie, both of whom are now at Claude Code.
Lee himself, and I got a job there under Jared Palmer as like the...
Wow. That Jared Palmer.
That Jared Palmer. Yeah. Who's a really good mate actually.
And he's the one who later moved to GitHub. He started or spearheaded stacked diffs, or stacked PRs, and now he's at Cognition.
Yeah. He went into GitHub, shipped stacked diffs, left, refuses to elaborate, and is now at Cognition. Exactly.
Yeah. But he's also an industry legend.
Yes. I mean, he's a great guy, and I worked under him for not very long at Vercel. So I was only there about 3 months. From there, I got a funny contract at Vercel, because I'd already been floating this idea of sort of TypeScript and thinking about TypeScript and thinking about maybe making educational material for TypeScript.
I had this urge while I was at Stately, the XState company, to teach stuff. You know, I'd been teaching for six years before. I'd been not teaching for four or five years at that point, maybe six years. And I thought, I need to get back to this. Like, I miss it, you know, and I love making stuff. I love making content. I love teaching people. And so that's what I started doing. And I started doing it for advanced types. I'd got in contact with a lot of crazy typing tricks, a lot of really advanced TypeScript stuff, while I was trying to force XState to be type safe. Very, very hard job. I think a mostly impossible job. And so I made a couple of tips. I made these two-minute
tips, posted them on Twitter, and they just went, you know, in a way I'd not felt before. And I realized, okay, there's a market here. And so, one Sunday, I just made like 13, 15 of these two-minute tips. I just queued them up over the next few weeks, and my follower count went from, you know, 4,000 to 10,000 or something, you know, it's just you suddenly felt that there was huge interest in this, right?
Exactly. A massive wave of something was, you know, some combination of the way I was speaking, the material I was delivering that was clicking in a way that I'd not felt before. And that's only really happened twice in my career. So, I was already floating the idea of a course and I knew I could do it well. I knew I could do a really great course if I just had the right audience and if it clicked. And so, I went into Vercel. I got a contract there for only 3 days a week for 3 months initially, which is very unusual. Is that what you wanted, or is this like how, you know, Vercel was probably testing the water, see how it goes?
I... Vercel wanted me to be full-time straight away.
You knew that there's this other thing. So, let me kind of hedge my bets if I'm able to do, right?
Vercel was this weird backup to what I... which is wild.
Which is wild to me.
For most people, this would be the dream job, right?
I know.
So, it's a little embarrassing to say because obviously it's so many people's dream job, but I went into it going, "Okay, I need a stable 9 to 5 for 3 days a week while I test this other thing out." But I mean, just to be fair, I think this is sensible, right? Like at this point, if we just go back to where you are, like you've been a voice coach for a good part of your career, let's say six years, and let's say now for 5 years, you've been building software. You love doing it. You think you're good at it. You think you might be able to teach, but who knows, right? And at that point saying, all right, let me take a gamble and like do this thing that might or might not work out, whereas if you can pull it off when you have something stable and it gets traction, now that's different, right? You know, a lot of engineers have aspirations, ideas, especially because with software engineering you can work remotely, you can take your idea, build a company, and they're thinking, all right, should I just plunge, should I quit my job, should I not quit my job? So in some ways I guess this is one model that is kind of unique, if you're able to pull it off. I mean,
It was the most bizarre thing because it became very clear very quickly that I couldn't stay at Vercel basically. So, we had about two months into my work at Vercel. I was there actually over a very tumultuous time because I was there when they released Turbopack. I actually wrote some of the documentation, the initial documentation for Turbopack, and met some of the team
which was a lot faster build system, right?
Yeah, it was a build system essentially. At the time they were trying to rival Webpack, what they were working with, and I was there initially when they were building the docs. I flew out to San Francisco. I was there for Next.js Conf when they announced it, you know, big, a really fun experience, and like I was, you know, there with everyone while they're getting everything ready for it. And so, you know, I do that and already in the back of my head I'm thinking, I've seen the newsletter, sort of my Total TypeScript stuff, creep up. I understand, okay, there's something really big here. And when I made a pre-release sale, that just went crazy. I was earning, let's say, X at Vercel, and that was like 30, 40x or something, you know, it was immediate.
And X at Vercel was already a really, really good compensation.
Absolutely. Very, very, very happy with that. But yeah, so it was obvious there was no other decision I could make. I loved working at Vercel. I would probably go back at some point, but I just couldn't stay. So I had to do this thing.
And then tell me about Total TypeScript. So you had this idea, you started to build it two days a week and on the weekends, and then you did this pre-release sale. What's the
I almost... I try never to work on weekends, basically. I'm extremely radical about this. I don't know, I mean, I think it's something I mostly fail at because I'm quite an obsessional person. I like trying to make something work, but I'm not one of these guys who's doing, what was it like? What's the SF thing where people go like 96 days?
996.
996. It turns my stomach, you know? I just hate that stuff. Like I am trying, with everything I do, to build a lifestyle and build a life where I can spend most of it with my family. That's my goal. And so just to prefix that, all of my decisions after that hopefully make more sense in that light.
So Total TypeScript, I was working with a guy called Joel Hooks. Joel Hooks is extremely funny, extremely influential on me. I've worked with him now for years, and he came up with Egghead. He's worked with Kent C. Dodds on his courses. Extremely successful course creator in the background, and I basically reached out to him and I said, would you like to make this course? And he said, hell yes, and we went from there. And so straight while I'm at Vercel, I'm also working with Joel, and we do this pre-release, and as I said, it just goes nuts, and I realize, okay, I've got to fully commit to this. And we get to, I think, about January 2023, February 2023, and we released the full course, and I don't know, I think I need to look at the charts from around that time, but it reaches seven figures extremely quickly. And that's a revenue split between me and Joel, of course there's expenses in that, but in terms of raw revenue it was extremely exciting.
Yeah, but the seven figures, that's $1 million, which is, I mean, an incredible milestone, right?
Which is nuts, you know, and life-changing. And I realized, okay, I can wake up in the morning and this money is still going to come in, you know. This is something that I dreamed about for a long time when I was a singing teacher as well: making material that I could sell online. You know, this is something I've been aiming for for a long time, sort of high-leverage work where I can do the work and then step back and go back to my family.
And for the next couple of years, I worked on Total TypeScript, sort of expanding the course, selling a couple of supplementary courses, and yeah, that's basically where Total TypeScript was. And so the main portion of my success in the last four years has been Total TypeScript and building that out.
Yeah. And Total TypeScript has been very inspirational, especially that you openly shared a big milestone when it hit $2.5 million of total revenue, which again, I think for many software engineers, you know, of course we know this is before revenue share and there's expenses involved as well, but it's something that is pretty clearly a higher earning potential than many great software engineering jobs. Not necessarily all of them, especially when we're looking at the US and some of the AI labs and whatnot, which was probably an exception. But the fact that there is a market and a business to be made of what I feel is a bit of an honest model, in the sense of, hey, I created this thing, people pay for it because they want to learn and they hopefully get value from it, because otherwise they would ask for a refund, right?
Exactly. We do a very extended refund policy. I don't tend to want to accept any other forms of money either. I don't like necessarily doing sponsored content. I'm not going to say never, you know, but I don't really... I've got a GitHub Sponsors page, but I'm really trying to take it down. For a long time, I've tried to take it down. I like the idea of just being someone who has, okay, these are the products you can buy from me. This is how you can support me. And hopefully this gives you, you know, 10x in terms of returns, because this is a very lucrative industry, right? And a lot of people have education budgets that they can spend. And if you want to spend some of your education budget on me, that's basically my model.
And a lot of that money, a lot of the people taking the course, this is from people's education budgets. You know, this is companies coming and spending big on education. And that was a market that Joel was very keen on pushing me towards and realizing this. You know, I didn't have much of a sense for how much money was sloshing around in the industry, especially in that age, and honestly still. But the fact that it just hit that milestone so quickly, within a couple of years, I mean, life-changing.
Well, I mean, this sounds like an amazing story, and it could be a fairy-tale ending where you keep creating educational content for the rest of your life and it's highly in demand, but then AI happened.
Yeah.
And as we know, it's changing a lot of how we work, how we find information. For example, I don't Google that much. I actually work with AI agents or deep research or some of those things. I heard stories about educators, online educators, who are saying that their revenue and market share and mind share is just falling down, because people might not want to sit through courses or sessions when you can just turn to the bot. How did you see AI impacting the industry, how people learn, and also your business, and also you as a teacher?
It's complicated because AI has changed the game, right? It's changed how important knowledge is, and specifically the types of knowledge that are important. So when I'm teaching my courses, I sort of think of there as being two layers. I'm teaching the syntax, obviously, I'm teaching the what, but there's also the why behind it, right? And it's very hard to teach the why without touching on the what, if that makes sense. So the sort of medium is, I'm going to teach you this syntax and maybe you might gather the sort of wisdom around it, right? I'm teaching you knowledge, but I'm also trying to teach you wisdom. Knowledge is now very cheap to acquire, right? Very, very cheap. You can just look it up, you know. I have a teach skill that can just take you and teach you the knowledge that you need. But the wisdom has gotten no easier to learn, right? It's still knocking about. Like, you are still going to run into the same issues that you ran into if you didn't have that wisdom before, even with AI. So in terms of my revenue from Total TypeScript, that's gone down, obviously, because I think people are not so interested in that material anymore, and I think people teaching that material are going to find it tricky because, again, that knowledge is really, really hard to come by. The only way that I've been able to, not necessarily survive, but it took me a long time to figure out where I wanted to be in the AI space, because I'm not a researcher from OpenAI. I don't have the credentials to talk about this stuff really, especially in late 2023, which was when I started looking at it. And I was initially making courses about how you put AI into applications. I thought, okay, I've been building frontend applications for a long time, it makes sense that AI is changing things a bit there. It started to be clear to me that was the wrong bet to make. I wasn't seeing the returns I was expecting, and the material was good and I feel proud of it, but I didn't think I wanted to make more of it. And around December last year, which is a date that many people cite. Oh yes, we know, or as we call it, the winter break where everyone came back AI-pilled.
Yeah, exactly. The Peter break, right?
Peter break.
The OpenClaw break, when Opus 4.5 is out. People have a lot of time off and they just start slamming it, and they realize, wow, okay, things are really happening. And that happened to me too. And I realized, okay, the AI is now good enough that you can delegate to it. You can actually make structures around the agents, and the agents can handle the knowledge, the syntax, the sort of, I call it, the tactical stuff. Yeah.
And you can handle the strategic stuff, the long-term thinking.
I use that a lot. I know you had John Ousterhout on this podcast. I've been really wanting to chat to him myself. He's a huge influence on me, and he talks about the difference between tactical programming and strategic programming. AI has largely eaten tactical programming, in my view, and it's up to us to handle the strategic. And I realized, okay, in the strategic layer there's a course I can create there. I was looking at Ralph loops at the time. Geoffrey Huntley was building this really cool stuff where you can loop the agent and get it to follow these goals, and I thought, okay, there's definitely material here. I just need to find a structure within which I can put it and organize how to fiddle around with it.
And I remember with the Ralph loops, you also made a video that became very popular on YouTube, X, everywhere, where you basically said, all right, here's how I created a Ralph loop. Here's a project that has a lot of to-dos. And you did a great job, we'll link that video in the show notes below, where you said, all right, here's how usually we would try to get agents to work: do a plan up front and then implement each step, the kind of traditional top-down planning. And you're saying the problem is that as you're implementing, or even as the agent's implementing, it realizes, hang on, I need to do more stuff, and then how do you modify the plan? And then enter the Ralph loop, where you gave the structures that you used at the time. There was an MD file, it adds tasks there, and it kind of eats through it but keeps adding. And it was actually a pretty eye-opener to me, as a way to think about how to do these agents.
Actually, the couple of years that I spent sort of trying to put agents into applications was really beneficial there, because when you try to build an app that contains an agent, you're always thinking about data flow. You're thinking about how the data is going to get in, what shape it's going to be, what priority, whether you're going to put it in the system prompt, the user prompt. You're working at a lower level than you usually get to with the harnesses. And so when I got to working with the harnesses, it felt like, oh, this just feels very familiar. I just need to, you know, where is the state going to live? How am I going to pass the state into the agents? What shape is that going to look like? How am I going to compact it or clear it? You know, and this was still pretty early days of Claude Code. Claude Code had been out, you know, five, six months at that point.
And it just felt very natural. And from there, I just got obsessed with these, I suppose we would call them loops now, but really they're just processes. They're sort of different ways of stringing agents together. This is kind of like the diagram, right? If you can draw arrows from one thing to the other, you can call that a loop, especially when there's a part that goes back, or you can call it a workflow or a flow or whatever, right?
I would call it a finite state machine. You know, it felt very similar to the stuff I've been working on in XState, which is process-based, which is state-based, event-based sometimes as well, where, you know, you have an agent at the bottom there that's calling an event back at the top. And I started just to see really good results from that. And I would do these experiments where I would try building out my process, and I would build out a feature of, you know, I have a few apps that I work on to extend what I do, like I have a custom video editor, a huge codebase. I have a few open source projects as well, and I was just building these little loops and little pipelines, and I would sometimes just drop it and go back to what the
default setup was and I just noticed a huge difference like I just felt wow okay the stuff that I'm doing here is really setting me up for success and I started thinking what's the best way that I can distribute that how can I share that with other people better and that's where I sort of started landing on skills as the distribution mechanism for this stuff.
Mhm. These are the skills for AI bots harnesses where you can typically define them and now you can once you install them you can invoke them with a slash command.
Exactly. Skills really they're just a folder of markdown files that can sit in your computer somewhere and the agents can either invoke them themselves. So model invoked skills or you can have skills that like the agent doesn't know about but you can invoke yourself. So user invoked skills and I started seeing these skill sets pop up everywhere like Superpowers and you know Claude Code plugins that you can install, I think gstack as well. I realized okay maybe I can distribute what I have, this process, as a set of skills and see what people think of it and initially I just put it up and I was doing other stuff. I was working on a course and I check back in and I realize oh it's got more stars than anything else I've ever done. I've not even really talked about it. You know, it's just sat there on its own. Word of mouth, I suppose. I've done a little bit of documentation, but really not much and it's just exploded already.
So, I thought, okay, maybe I should put a little bit more work into this. Maybe I should talk about them. And I did a talk at, I think, where we met last, which was AI Engineer London, about in April. That talk was entitled Software Fundamentals Still Matter. And that talk is now up to, I think, 1.2 million views or something. And I mentioned the skill set. The skill set is now at 230,000 stars. It is now the second most starred skills repo in the world. I think somewhere like 20th to 25th of the most starred repos of all time. Wow. You know what I mean? Like what's going on? So there's obviously a hunger for this. So this was the second time in my career, just like when I was putting out the little TypeScript videos, where I felt, wow, there's a momentum here. There's something happening. And so I felt I had to double down on that.
The skills, how did you write them? Is this trying to capture your workflow, your understanding of what works with agents, you know, not just right now, but of course, you're thinking about state, you're thinking about how you were integrating AI into applications, which again didn't take off all that much, but you learned. So, is this kind of like Matt's workflow, Matt's way of what works for me?
Yes, that's what it is. I try to think first of all people are going to use these skills and they're going to tinker with them. So how do I make the simplest set of skills that people can audit very easily? I'm trying to think how do I maximize people picking these up and using them at work. So for instance the grill me skill which is the most popular one, you know, I don't know if you've used it but
I use it as well. Yeah.
Okay. It's annoying how damn it grilled me. I just asked it. I was like, I'd like to expose an API endpoint that can tell whoever has the authenticated token, I have some basic authentication, is this email a subscriber to my email list or not? Because I want to connect it with one of the events that I'm doing to give priority to paid subscribers. And that's very simple, right? And then the grill me thing, it starts to just really grill me like, okay, so what about authentication? Do you want the bearer token or do you want it in JSON which is not as safe etc. Like okay, well that's a decision to make, and then we go through all of these decisions and it goes really low level, including like okay, how do we enforce rate limits? When it comes to rate limits, do you want to exactly do it when you get like a thousand per day and not allow a single more, which is more complexity? And I just realized it's been a long time since I've had such an involved design discussion with a team or an engineering team, and you typically have it when someone has deep domain knowledge. And I was both annoyed by it, like this is just simple, no, don't worry about that, but also impressed that this thing, this AI, this LLM, through a series of prompts is able to do all of this.
It's, I mean, everyone's got a grill me story. So I get so many of these in conferences. People say, you know, you've got this very very simple skill. It's really just telling the agent to interview you relentlessly about the topic. It's a very small skill. It just has this weird emergent behavior with it where the models just start thinking a little bit outside the box and they start throwing ideas at you. I think I got it originally from Tariq who works with Claude Code. He's saying basically get the agent to interview you and then you'll see better results. So I encoded that into a little skill and I realized wow okay it's just sort of 10 times better than anything I've ever used.
And that grill me skill was the first one that sort of reminded me of the discussions I would have at my first job, you know, with the guy with the sandals. It's this very senior engineer in the room really getting me to think about everything that I'd done. It was the most familiar thing to me to actually working with someone like Anderist Rake at Xate. You know, it just felt like a really high quality developer was asking me these good questions. And I thought, wow, okay. And then I started sort of taking that and going like how do I mine this agent for more software fundamental stuff? How do I make it feel more like a proper developer, a real senior? How do I tickle the right latent space in order to get its behavior to change and challenge me in interesting ways? Because if you can do that, if you can increase the quality of the conversation you're having with the agent, you're going to increase the quality of the outputs.
What I liked about the grill me skill is it forced me to make decisions that I know what decision to make when I think about it, but it is my decision. So unlike when you do the /goal command, like build this, and it goes off and does this and it makes all the decisions or most key decisions. What I like about grill me is I both make the decision, but also sometimes it reminds me about things that I didn't think too much about, or maybe it reminds me that I should do a bit of research. For example, it asked me which authentication would I want to do, a bearer token or over POST or even over GET, and then I'm like hang on, I'm going to look up what the differences are or ask a different session to educate me. So it makes me a better professional. And I do have this belief that when you're working with AI, as long as we're learning I think we're fine. As long as we stop learning and outsource the learning to this thing, trouble will be brewing maybe, you know, months or years down the road.
100%. There's two things there, right? I think what everybody underestimates about agents, everybody, is that there is a communication gap between you and the agent, right? There is a barrier there. You feel like because the agent is not a human and because you understand your hierarchy of values, you think that the agent will just pick up on them, right? There's this sort of feeling of, yeah, just trust the model. Especially with the top tier models, you know, just trust the model. But the agent, however good it is, however smart the model is, you know, even Mythos, it can't read your mind. It can't read your mind. So there has to be some process of communicating your values to the agent, because often when you do like a goal, when you just go okay just spam me out some code, give me some slop, the agent is going to produce something that's totally misaligned from you because it doesn't understand what you think is important. And so grill me is not only about implementation details, it's also about establishing okay, do this, this is in scope, this is not in scope, here's what I think is important. And so it's the agent getting to know you.
Matt just described the grill me skill. When I used this skill to design an API endpoint, the first questions it asked were about what the endpoint was and wasn't allowed to do and who it was allowed to do it for. Now, in my case, I had a decent idea of what I wanted, but it's generally a terrible idea to let an agent improvise authentication and authorization as they would often do.
This brings us to our season sponsor, WorkOS. How do you authorize AI agents? The problem you have is how you want to control the scope of the agent. The tricky part is how permissions are static, but the job of what the agent does is dynamic. So teams pick between two bad options. Either you read the prompt, then approve every tool invocation and call by hand, and then read the prompt again, then approve by hand again, until you eventually just stop reading the prompt, or you just run in YOLO mode, letting it rip and hoping that whatever the agent does is not irreversible. But there's a better way.
WorkOS just launched Airlock, intent-based access control for agents. You write the rules in plain English. For example, read repos and comment on PRs. Anything touching auth or billing needs sign-off. Never push to main. Every call that the agent makes is judged against its task and allowed, denied, or sent to a human. Every verdict is logged by Airlock. The neat thing about Airlock is how there are no pre-granted scopes and there's no rules to assign. The task itself is what defines what the agent can do. WorkOS Airlock is in early access. Request it at workos.com/airlock.
I'd also like to mention our presenting sponsor, turbopuffer. Matt and I are discussing a fundamental question. How do you get agents to remember what's important? Here's an idea. What if instead of building a complex memory system, you just let the agent search its entire history? This seems like it will be very, very expensive, but with turbopuffer, it isn't. Turbopuffer's object storage native architecture means that the marginal cost to store session transcripts is almost nothing, making it economical to index the entire chat history. And because turbopuffer namespaces scale virtually without limit, you can create a dedicated search index for every agent.
Here's a good example of this. Entire, another season sponsor of the podcast, indexes hundreds of millions of agent session transcripts for search and then lets the coding agent retrieve what it needs to recall how and why an engineering decision was made. Entire showed that their agent was more accurate, used fewer tokens, and took less time to find memories when it used turbopuffer instead of git history and a CLI. Agent memory is a complex and evolving use case. But perhaps there's a bitter lesson here. Maybe the best solution is the simple one. Just search every transcript. With turbopuffer, this is actually possible. If agent memory is something you're trying to solve, then please reach out to the turbopuffer team at turbopuffer.com/pragmatic.
And which other skills did you create?
So from there I thought okay, how do I take that conversation and turn it into code? And I was immediately scared because I'd been working with models, you know, just before they were good and before the December winter where, you know, things got really good, and so I felt the constraints from what I'd been working with before. I knew that for instance the more context you give to the agent, the worse it performs. I know you had Dex Horthy on this podcast. Yeah. And Dex is a really big influence on me, especially his idea of the smart zone and the dumb zone.
Smart zone and dumb zone. Yeah.
Yeah. So the idea of that, just so you don't have to go and listen to that podcast in full, although you should. You have essentially the more context you give to the agent, every token is shouting for attention. And the more voices you put into that room, the harder it is to hear the important ones. And so the model starts losing the connections between things and making mistakes because of that. And you can think of that as a slow decline. But there is a portion of the context window where it's better and where it's worse. And so you have the smart zone, which is currently I would say about the first 150,000 tokens of frontier models
of a 1 million token window.
Yeah. Of any size token window. Doesn't matter the context window size. It's all about raw amount of tokens, raw amount of attention relationships, and then the rest of it will slowly degrade more and more and more.
And so I started thinking, how do I take work that's bigger than 150k tokens, which is not very large, and portion it out over multiple context windows, multiple sessions? And this took me a lot of tries, a lot of different fiddling around with different approaches. The Ralph loops was one version of that. Ralph loops are designed to make the most of the smart zone because they essentially just give the Ralph loop a goal and they say do the smallest possible change that will get us further towards that goal and then clear your context
and then clear the context, start from fresh.
Exactly. And you're not technically starting from fresh because you've got the code base, right? There's a little bit of state saved in the file system and in the environment, but not in the model essentially. So that's the idea. So I started thinking, how do I take that Ralph loop idea but make it a little bit more stable and turn that into skills? And so what I realized I needed was two different types of documents. You need a document for where you're going, which is the destination document.
I used to call that a product requirements document, or a spec is what I call it now.
So that's the specification that declares when you've reached the end. And then you need to break that spec down into individual tickets, one ticket per session. And so I have a very simple skill just to spec and then to tickets. And so you take that grilling session that you've had and you turn it into a spec. Now that spec can work over, you know, 30, 40 tickets, let's say. You can have really massive great big chunks of work that are all tied into that spec. And so that's the main idea. You just grill. You turn that grilling into a spec. And then you just run some kind of implement loop over those tickets until you've got a huge chunk of work.
After grilling, do you get user input as well, or throughout this process, or it depends?
I was mostly designing this to be run for the user, like to be away from the keyboard totally
because there's this idea of like the day shift and the night shift. Have you heard of this?
No. No. No.
It's great. Basically, the optimal way to work with agents is to plan during the day shift and then get the agents to work during the night shift, right? And so hopefully you wake up in the morning and you've got beautiful clean code to look at. And that's what I was trying to optimize my process around, because I was really sick of what I, and what I still do to an extent, of just switching between terminals, context switching all the time, just going boom boom boom boom boom boom. What I wanted and what I'm trying to optimize for is to just get a good chunk of planning done and then let the agent work for a couple of hours, and then I can do other work, decent chunks of time, 15-minute chunks working on one thing, planning on stuff, and then I can review the code and do that. So that's what I was trying to optimize for all the time when I was doing Ralph loops, and that was the big thing that I found in December, is these guys are good enough to delegate to and so I can run them AFK.
And then you have a different skill as well which is a bit more ambitious called the wayfinder skill. Can we talk about that?
Absolutely. So in exactly the same way that implementation I noticed needed to be split out over multiple sessions. Sometimes you're grilling something and
you're hitting the limits. You're going to grill something, you know, build me a Stripe clone or something, right? You are going to hit the limits there. There's no way you can plan that in 150k tokens. And so I thought, how do I break that up so that I can run grilling sessions that can be infinite length, right? How do I split up grilling so that it can work like that?
And so I came up with this idea. Again, I'm thinking about the flow of information essentially, like what does it need to perform well in a grilling session? It probably needs to understand exactly what the purpose of that grilling session is, but it also needs to understand what's been decided so far. It needs to understand what other grilling sessions might be happening at that moment.
And I came up with this idea of a map. And the map would be the sort of center point of everything that was needed for all the decisions that you were coming up with. And once you've got a map, you realize, okay, as I'm trying to find my way to a destination, there are certain things I know I need to decide, certain points that are kind of like milestones on the map. And there's a fog of war. And that lovely metaphor just sort of carried me through designing the rest of the skill, right? Because you've got your map, you've got your fog of war, you vaguely know where you're going, and every time you have a grilling session, it opens out more points on the map. And so you sort of figure out where you're going. And so this is kind of like a directed cyclic graph where you're walking down until you reach your final destination.
And so you've got the map, and then each individual session in there are tickets on that map. And I realized, okay, grilling is good, but what if you need to prototype? What if you need to do research? What if you need to do an arbitrary task, like provision some infrastructure or something? Well, those are different types of tickets on the map. And Wayfinder basically just guides you through this process. I've had maps that have, you know, 50, 100 tickets or something until I finally reach my destination.
I've actually been using it for course planning as well, so non-technical stuff, which is really great. I've been using it to build a garden office in my garden, right? A lot of these skills, we say, okay, these are great for engineering, then you realize, okay, engineering is just a discipline. What are we doing here? We're just discussing something, we're doing things in real life like clicking around websites and stuff. You realize how easily that can map onto other domains. So maybe we can touch on that a bit later, which is I am thinking how transposable this stuff is into different disciplines and into different areas of life. So Wayfinder has been great.
Yeah. But one interesting thing about engineering and software engineering: when Hillel Wayne was on the podcast, he interviewed engineers who he thought are real engineers, chemical engineers, mechanical engineers, civil engineers, to try to find out: is software engineering real engineering? And in the end he found that it probably is. But he said that one interesting thing with software that is very different to every other engineering profession is the materials that we work with. In every single place, mechanical engineering, civil engineering, even chemical engineering, you have a material that has a threshold of things. You don't know exactly what it's like. You know that it can take about this much load, etc.
But in software, the material is software, which just works like a program. I mean, take out nondeterministic, which maybe brings us to more engineering, but software, code, you run it a thousand times and it does the same thing a thousand times, whereas in other fields it doesn't. And he said that he sees a big difference. But now I guess with LLMs maybe we have this thing where you run it a thousand times and it will have this variance that most of engineering has. So who knows if what works with LLMs will be useful in other engineering, where again they already had this variance, or we can take some approaches from other engineering professions that will maybe work nicely with working with this material called AI.
I totally agree. What I think is interesting about software engineering, and the reason agents are good with it, is all of the inputs and all of the outputs are text-based, everything. So the inputs: code, documentation, instructions for the agent on what to do, all text-based. And the output is more code, is test suites, is type checking results, linting. All that stuff is text-based.
The thing that agents really struggle with is anything that's non-text-based. You see these amazing demos of people one-shotting a perfect UI first time. Well, what about if you have an interaction problem in that UI? What if you're hovering over something and the animation doesn't look right? How are you going to get that to the agent? I mean, you can record it a video, I suppose, and it sort of pauses on certain frames, let's say. But it's actually not that good in terms of vision just yet. And so anything that's non-text-based is just garbage from the agent. It just can't handle it.
And so I think in those sorts of professions, I assume they're doing simulations, right? I assume they're doing some kind of, you know, I don't know if you have a linter that can work on an architectural diagram. I'm sure you have some variety of that, right? Some simulation. If you can make that text-based, if you can take the interactions that you have in your day-to-day life and turn them into text, which mostly they are anyway, then the agents are going to do a pretty good job. That's something I'm trying to do currently: take all of the services that I use and plug them into agents, right? Make them available to the agents. But yeah, the more we can make our work agent-friendly, the better results we're going to get.
Circling back to AI as a whole and what has changed. It has changed so many things, but one thing that comes up with AI, often especially from researchers and people working in AI companies, is "no prior": with AI you should let go of everything that we knew before, because this thing is different. Start from scratch. The approaches might not work; in fact, let's assume they don't work and come up with new approaches. Having been a developer before AI, and you were really interested in building quality, great software, how much do you think AI has changed everything, including the fundamentals?
This is something that I thought too. I thought, right, AI has changed everything. I'm going to throw the baby out with the bathwater, right? I think we just need to look at everything in a new way. I started doing that. I was especially looking at spec-driven development, you know, which I have sort of mixed feelings towards. I think it's a strange term; it encompasses too much. And I thought, okay, right, maybe English is the hot new programming language, right?
Which went viral at some point when Andrej posted it.
Exactly. Like, maybe I can just write a spec, and that specification is going to be persistent, it's going to be something I can edit, and just get the agent to change it as it goes. And as I experimented with it, I tried it a lot, and I was just getting worse results than if I'd coded it by hand. And it wasn't getting better as well. And I noticed that every time I would run this loop of change the spec, see the code change, the code would get worse. You're not supposed to look at the code, of course, but I looked at the code and it was garbage. And I thought, how is the agent going to perform well in here? How is it going to work? Because the feedback loops are so important to the agent. If you have a bad test suite, the agent is going to get bad signal from it, just like a human would. And I thought, how do I improve the test suite? How do I get this setup not churning out garbage every time?
And I opened a book that I had on my shelf that I think was still wrapped in plastic the first time I took it out, which was The Pragmatic Programmer. Everyone told me to read it. Everyone said, you know, this is the best book ever, you just got to. And I bought it and I didn't read it for some reason. And I opened it and it had a whole section on software entropy. And entropy is the idea that things go towards a more disordered state; that is more likely than them going into an ordered state. And I realized, okay, software entropy is inevitable. What I'm seeing here is that agents are producing software entropy at higher rates than ever.
And I started looking more into that book. And almost every line I read, I thought, wow, this feels like it was written for today. You should go back to that book. These ideas of don't outrun your headlights, always work within your feedback loops, programming by coincidence, tracer bullets, so many smart ideas. And I realized this book has been out for 25 years, right? This is probably in the agent's prior. Maybe if I just mention some of these concepts, especially the ones that are really pithy, like tracer bullets for instance, which is the idea that you should always get feedback really quickly on the work that you're doing.
I guess the idea of the tracer bullet is, right, like a tracer bullet that leaves a mark, you implement a path that works, an important piece of the software, instead of building a database layer and the application layer and, I don't know, whatever layer, building all three and then putting them together. Just build one part of each, but they should work together.
That was the problem I was seeing with agents. You would get it to build a piece of software, even with Ralph loops, and it would build the entire database and then it would build the entire application layer on top of that. Then it would build the entire React component library. Only at the end would it start actually plugging things together and getting feedback on what it was doing. And it was maddening, because things in the database will affect what you show on the front end. You only really know whether things are actually making sense when you see it crossing those integration layers. And so another concept is vertical slices, right? Instead of these horizontal slices across these different deployable units, you have a vertical slice where it gets feedback on what it's doing straight away and builds out from there.
And so I just started using these phrases in my prompts when I was talking to the agent, and I started noticing that it was saying those phrases back to me. It was repeating them back to me. It was saying, "Okay, I'll turn this into a tracer bullet because this is a tracer bullet. I'll do this." It was using the words that I was using in its own reasoning traces. And so this is what I call a leading word, a Leitwort, let's say, which is a sort of fancy literary term, where you lead the agent just with a simple phrase that you repeat a couple of times in the skill or the prompt to change its behavior. And so tracer bullets was a fantastic one. And I just started diving into different books, all the books I could find, to try to mine them for leading words. And another one was John Ousterhout's book A Philosophy of Software Design, where I picked up tons of great stuff like deep modules, which is a massive one for me.
It's interesting to consider: these agents have obviously been trained on those books, which are still available in print, and of course there are arguments about what they're doing with those books and whatnot. But if it's in their training data, and the agents, as they're trained, connect all these different concepts, then these leading words could invoke those concepts. And I wonder if this is much different to when, on a topic, you talk with a professional and you're trying to describe as an amateur what you want, and a professional says a word that does that, and a fellow professional gets it. This is jargon, right? And jargon, on one end, it's not very inviting when you join a company and there's jargon, but we use it because it makes things faster, easier, fewer misunderstandings.
Definitely. That was something that idea led me to, because obviously you've got these leading words that are in the agent's prior, right, like tracer bullets, all that stuff. What about describing my application? What about describing my code? Because the agents are just awfully verbose, right? Especially Opus 5; for some reason people really go after that model for being verbose, and it really is. And I thought, how do I get it to be less verbose? How do we start talking a common language between me and the agent, this communication barrier again? And it led me to DDD, domain-driven...
Domain-driven design. Eric Evans' incredible book, where he talks about ubiquitous language. Again, really deep in the agent's priors. It understands it really well.
I started toying with the idea of maybe changing grill me a little bit, because grill me is a very simple skill. But what if, while we were ideating, while we were thinking about the application we were going to build, what if we were also building a domain language? What if we were also deciding on the right terms to use? And this turned into a skill called grill with docs, which is a terribly named skill, but it essentially creates this domain language as you go. And if you get the agent to use the domain language, the difference is night and day, because suddenly you're speaking the same language. You're able to describe the things you want to change in so many fewer words.
Like, I had this app that I sort of work on. There's this complicated interaction where there are ghost lessons and real lessons. And what happens when you turn a ghost lesson that's inside a ghost section inside a ghost course into a real lesson? That means the ghost section needs to become real. The ghost course needs to become real. How do you explain that? Well, that's the materialization cascade, right?
And you came up with these terms?
With the agent, right? The agent is actually really good at coming up with these terms. And so I have a domain modeling skill. And we talk about jargon, but really it's domain language. And if you can integrate that, not only with the way you talk about the app but with the app code itself, then you've got a stew going, right? It's very, very exciting, and it means that the agent can navigate your codebase a lot easier. It can find the functions that mention that specific domain terminology just with a simple grep. It's just gorgeous. So that's something I've been really integrating with every part of my setup, is DDD.
But this is so interesting, because in an effort to make these agents work more efficiently, or do workflows that mean that you can produce better software with fewer mistakes, you start to go back in time. You found this book that is now, what is it, like 20, 30, 40 years old. And you're even still going back and finding gems from those. I'm sure at some point you'll get to The Mythical Man-Month.
Yeah, I've got it already. Absolutely.
Which is now more than 50 years old. And you're trying to find the right words to describe things, which is very curious, because when I talked with Kent Beck about how they used to program with Ward Cunningham, as they were coming up with the concept of design patterns, they had a thesaurus with them, and they would look through it trying to find the right word that has the right meaning, and they had it on their desk. Right now, this feels like we're going back to the fundamentals, the questions that people have been asking themselves, and
Every now and then people write it in books and it kind of spreads as wisdom, and now we're back to where we started, which is what you're trying to teach is the wisdom part. It's wild, right? Because AI is so different to humans, you need to optimize it. Imagine you essentially had a human who wakes up every morning and cannot remember who they are, right? The guy from Memento. This is Memento-driven development, right? We are trying to optimize our code bases for new starters.
So, we're trying to have the most healthy code base that we've ever had, because a human can work around a bad codebase, they just develop memory. They just slam their head against the wall again and again and again until they've got there. But an agent can't do that. It starts fresh every single session. And so, you need to optimize your codebase for that person. That leads you down into really interesting paths. And it turns out that software fundamentals have been saying we've been trying to do that for the entire time, right? I am fully, I don't know, Eric Evans-pilled. I'm fully software fundamentals-pilled. We are sort of changing the rules a little bit, but maybe we're just emphasizing rules that we knew we were supposed to do but maybe we didn't. And I find that really fascinating and it's definitely a lot of fun.
Now, okay, I think it's easy enough to follow with this train of thought why fundamentals matter, but which fundamentals? And if I'm an engineer, especially maybe someone who has been just kind of heads down coding, more technical coding, how do I go about and find those fundamentals that matter and go back to? What have you found works?
This is a really tough question, right? It's really tough because strategic programming has always been really hard to learn. The reason for that is that the feedback loop on it is really long. You would often find people who quit their jobs after 6 months, their strategic mistakes never catch up with them, right? Maybe that strategic mistake takes nine months to come back at you.
I think of strategic programming as kind of like you've got a huge mixing desk in front of you with loads of these different sliders. Maybe one of those sliders is the amount of deployable units that you have. You turn it up, you've got more microservices, right? You turn it down, you've got a monolith. How do you make that decision? Where do you put that slider? Because it's kind of like you're mastering something. You're mixing some music, but you can't hear what's wrong until 9 months later, right? Until the mistakes come and get you.
So, I think the only thing that can make that feedback loop faster is moving faster. AI now lets you move faster, right? And so, your strategic mistakes will come back at you quicker. They will come back at you quicker because AI is just able to produce so much code. And so, what you need to be thinking about is that your code is the environment the agent operates in. And you should always be thinking about improving that environment, thinking about how to do it better. And obviously that requires a bit of tactical knowledge, right? You need to understand what code is and how it fits together and what the memory constraints are and all that stuff.
But in order to get better at strategic programming, you just need to be thinking on that level all the time. And I would say reading these books as well, because just having the language to explain that and understanding the difference between applying strategic techniques and not is the whole game.
I mean, up to, you know, pre-AI, senior engineers, staff engineers, they were the people who often... You didn't see a senior engineer under five years of experience, because you typically needed, even in a fast-paced environment, that much time to get the feedback loops, to make your own mistakes. And by the time people got to staff engineer, oftentimes around 10 plus years of experience, some people did it earlier, but they often just had battle scars all over them. Someone would start a new project, they would go in and they would just make a tweak and it wouldn't be clear why, and they were like, "Trust me on this. We're avoiding disaster in production or on call" or whatnot. But all of this came through lived experience.
Now AI speeds things up. It also makes it easier to fix mistakes. So I'm wondering how this might change. On one end, I can see how it could just speed up experience. In a year, some teams will ship more projects than they have in four years, or about the same as in, let's say, three or four years before, so you get a lot more experience. But I wonder if sometimes the mistakes that you make are just not as serious because you can fix them quickly, and I wonder if the learning is not as strong. Because again, some of these battle scars, these war stories, are: it was just a really bad outage, we lost a lot of money because we didn't have idempotency. Now of course you know what idempotency is. It's not an easy concept, but it's important if you've been hurt by it, and so on.
If you're a company right now and you want to train the next junior developer, because this strategic programming knowledge is so valuable now, because you can use it at such higher leverage, are you really going to employ someone without it? Why would you? I did an interview with Uncle Bob the other day, and his recommendation was: okay, you just hire someone and you treat them as an agent for a while. You just delegate to them. You keep them in that tactical mindset for a while until their mistakes start coming up at you. But that's such an enormous waste of money for people, right? When the tactical stuff in software engineering has gone below minimum wage in a lot of countries.
So I don't know what the answer is. I only know that the strategic stuff, the understanding of the code, the understanding of the long view, has gotten more valuable than it's ever been, right? Because you can just get so much leverage out of it.
I asked people about interesting things they'd like to know from you, and this is very related to this. This person asked: how do you convince non-engineering stakeholders that investing in software fundamentals is important, even if it might reduce the speed and productivity on paper? I think the question here is if some people advocate, look, we do want to get the fundamentals right, which means we want to take it a bit slower, think about our decisions, maybe educate ourselves as well, as opposed to just churning it out.
I mean, you could have asked the same question 10 years ago, right? And it would have still been relevant, you know what I mean?
Except we sort of thought of elements, we would have asked about paying off tech debt.
Exactly, and it's the same thing, right? We have been having the same conversation, which is quite satisfying to me, because you need some sort of metric for figuring this out. And it's a little easier to figure this out because agents allow you to move faster. And the first step to this is getting observability in your organization over every single agent, on what it's doing and what its success and failure rate is.
Yeah.
We've never been able to have that with developers before. You know, that's kind of invasive for developers.
Yeah. But for agents, it's okay.
It's okay, right? We are paying for this service, right? We need to understand how well we're optimizing for it. The first step there is actually getting a harness or observability around your agents, the entire organization, to work out what's working and what's not. And you probably need someone whose job it is, or part of their job is, to look at that data and figure out what we're doing. Maybe some repos in your organization have better success rates than others. And so you take the lessons that are in there and you pass them out.
I also think that most organizations need to gather around a common set of skills. You need a common software workflow process so that everyone can contribute back to it, so that you can experiment with things. You can A/B test things. You can have one team doing one set of stuff and one team doing another set of stuff and then you ask them afterwards. And so everyone working with agents in any kind of organization needs this experimental mindset. You need to be thinking, how do we get more juice out of these tokens that we're spending? And observability is the first step there.
Yeah. And I also wonder if there's a human feedback loop, in the sense of just talk to your colleagues. We do have rituals, team meetings, companywide meetings for a reason, to share: here's what's working for me, here's where it didn't work, here's what I'm learning. In the end, we are in charge of setting up the rules, deciding how we use them, where we use them, where we don't use them, and where we say, no, humans need to take 100% of this, we're not even getting AI involved. Which again will be different everywhere.
And it's not only that. A lot of this stuff now you don't need to be human in the loop for, right? You don't actually need to delegate that much time in order to build up a better codebase. I have loops where essentially every morning it will run my improve codebase architecture skill and give me a proposal for something that I could improve in the codebase. And then I can just press a button. I can say, okay, turn that into tickets and then let's ship that. That is pretty easy to do and it's pretty easy to stream that in with other work.
And so I think, I don't know whether you need like 20% of your time focusing on the factory that builds your software as well as the software, because I feel like that's a massive, incredible investment into your future leverage, and not only your leverage with your work but also your team's leverage and understanding and getting better at those skills. But of course you need results, and you might need to hide that work for a bit before you actually reveal it: this is what we've been doing all along.
Well, and this is down to your environment, but yeah.
And no one's going to be mad at you if you come back saying, oh, by the way guys, I also did this.
Yeah, exactly.
I wanted to ask you about specifically how you use tools. First one is coding agents: local or in the cloud? You recently posted a pretty provocative tweet, which I'll quote: "I'm moving away from my local dev setup. Makes zero sense to me now."
A lot of people ask me, how do you make your skills collaborative? How do you have a collaborative grilling session? And the answer to that is that you need more than just your terminal and you, right? We're in a phase now where every dev has like 100 terminals available to them. And that seems crazy. It feels like you need those 100 terminals available to your entire organization. You need to be able to collaborate in a shared space. You need to be able to tag someone into your grilling session and say, "Okay, do this."
And so it makes a lot of sense to me to have a lot of those interactions in the place where you already work, in Slack or in Discord or in Teams, whatever, or Linear. And that is really the thing that's driving me to explore this. I don't work with a team particularly, but I understand the value of that, and I've been trying to build that into my flows.
So on the train over here, I'm in Discord chatting to my Hetzner box, you know, building stuff for my course or fixing bugs that students are coming across. So I can see less value now in just doing things locally when I have this setup that I can port forward into, let's say, and see the dev server as it's making changes. And I don't know, it just feels like it makes way more sense to me than having a very, very expensive laptop that can do this stuff. It feels like wasted compute. And especially because on that remote box, I can set up schedules. I know the box is always going to be on. I have a morning standup with my agent where it schedules my day for me, and it understands all of my Discord chats and all that. Yeah, having that remote feels like it makes just so much more sense for me. And the only thing I do locally now is debugging issues with the remote bot.
Yeah, I wonder if there's a question of how easy it is to replicate some pretty complicated local setups in the cloud. But once that becomes possible, it's probably a matter of when, not if.
Yeah. And if anything, people are having a similar issue with local setups, right? With just a thousand Git worktrees spamming their hard drive, and with how do I have a worktree where I've got to run like five Docker containers in order to get my local dev setup? Well, that's often a little bit easier in the cloud, because you can just provision the resources that you need on demand.
And by the way, we're seeing that companies like Ramp, Stripe, Uber that have platform teams that manage to take the local dev's full setup and put it into the cloud, on a cloud machine that you can now invoke with an @ in Slack or a website, are seeing people use these agents far more. Except for front-end work, where you still want to have that feedback loop. There are a few exceptions where you really want to have a local dev setup for latency or whatnot, but they're also seeing like 70, 80% of devs just voluntarily going for the cloud.
Yeah. I mean, I think you can just tunnel through. If it's running a dev server and you just have that appearing on your local machine, how is that different from having it locally, right?
Okay.
I don't know. I've not experimented with that, but when I talked about that and said, "Oh, maybe front end is a good exception," that was the immediate response that I got, and it makes sense to me.
I want to ask you about planning and requirements. You're a big believer in Grill Me and planning up front, or getting the plan and then having the agent work. But there's a devil's advocate here. Agents are so fast at implementing. You could actually even have a few agents implement different architectures. What about the approach of, well, they're fast at implementing, so I might not need to do as much upfront planning. I can just course correct as I go.
It depends what type of work you're doing, right? Because I believe that you shouldn't be using Grill Me for everything. Essentially, you need Grill Me for pieces of work where the actual thing being done is going to be quite large and hard to row back from. If you feel like, okay, this feature, maybe it's a whole new page, maybe it's a big feature, you think if the agent gets it wrong, then the wrong code is going to be in its context window influencing everything that comes afterwards. And actually going back and editing the stuff afterwards and doing the alignment after the fact is going to be expensive. Whereas for those cases, it makes sense to align first, to answer all of the tricky questions, like your JSON cookie or whatever, your authentication token, first, and then do it. But for some cases, like simple bug fixes or just move this button three pixels to the left, it's obvious that you don't need to align before that. You can see the thing if it's just a five-line change or something. You can align afterwards.
And so that's how I think of it: where you can, you should shift right as much as possible. And actually there are certain features where, in my video editor, I have a button that I can send feedback with. And I often use this for very simple tasks, where I send the feedback, it goes into a GitHub issue, this immediately gets picked up by an implement agent, gets worked on immediately. Then a code review agent comes in and reviews
the code and then at the end I get to see this actual thing being fixed and I can do my alignment then. And that's worked really well for things that are very easy to specify, things that I don't need to grill on. So those are the choices you've got. Is it a small enough thing that I can align afterwards? Then don't use grill me. Does it fit into a single session? Then use grill me. Does it span multiple sessions? I need to align over the entire thing, then use wayfinder.
Interesting, because this is not all that different to where some tech companies landed years before, which is on the PRD, the product reference document. If it's something trivial, just build it. If it requires the team, like it's a team-level scope, I mean, write a PRD, send it out to the team, maybe CC some other teams, but it's not a blocker. And if it's something bigger, then it's a blocker, like we need to wait for feedback. Basically, the way we would say it is like, look, if it's like a one-month project, like spend two days, like it's not a bad thing to spend like one or two days planning it because we're going to save time on it. But if it's a one-day project, like forget about it. If it's a one-year project, I mean, what are we doing? Like it should be a smaller one.
Totally. And I want to, like, there's a bit of sort of criticism I hear just from outside the room when you say that, which is: doesn't this sound like waterfall, what we're doing? When I'm talking about wayfinder and when I'm doing any kind of, like, building up any kind of spec, I do a lot of upfront aggressive prototyping before we get there. That's something that comes up again and again and again. It's like, this is just waterfall. What are we doing going back to the 70s? But agents give you this ability of just churning out slop, right? And sometimes you can use that to your advantage because a prototype, right, just getting a sense for what it should look like. You can build out three or four different versions and just choose your favorite and iterate on it and just keep churning, churning, churning. That can be a really powerful setup that we've not really had before, right? It was always expensive to produce prototypes. Now it's the cheapest that it's ever been. And that's an essential part of writing specs to me, is actually producing these prototypes.
Yeah. But also, like with the waterfall criticism, I think Grady Booch might have told me this as well, is like, don't forget, we should not criticize waterfall because, for example, a lot of big tech, the largest tech companies from like Amazon, Microsoft, Google, Meta, you name it, they are kind of doing mini waterfall. Like pre-AI, they've been doing pretty mini waterfall, which is: let's do a plan, let's agree on it, let's build it, let's ship it, and this is all done in like 2 weeks, a month, two months, 3 months. 3 months is kind of the extreme.
But Grady Booch was saying the problem was never this with waterfall. The problem with waterfall was the planning was literally taking like a year, like one year, and then the implementation taking 3 years, and by the time it was ready 4 years later, it's not what we wanted. And that was the problem. He was like, the problem is not like having like a one or two month project or one week project with a waterfall. The problem was always that we're talking years. And he said that the industry has not seen waterfalls for decades now. And so here we're using this term, which is a bit like we're criticizing, or mini waterfalls were criticized in that one. It's actually not a bad thing necessarily. You see what I mean?
A scarecrow that we're punching or something.
Yeah. It's a piñata which stopped existing. It might exist in some crazy enterprise projects that none of us know about in regulated industries. But I feel even there it's probably gone out of style.
Yeah. I think if we're hitting the piñata, I think it's actually a useful thing to have up there. It's like a useful ghost or useful cautionary tale, right? Because which one fits the agentic setup more closely? It's going to be agile, right? Because the cost of labor has gone down so much, we can just make changes very, very quickly. I don't know, that feels like the right metaphor to me. So I don't mind hating on waterfall even though no one really does it anymore.
Well, one other thing that just went out of style. We didn't hate it, but test-driven development, TDD. What is your take on using it for agentic stuff? Well, when I talked with Kent Beck, we talked about how this could be a great fit for many reasons, but I still don't see people really using it. I see people writing tests. The agents also write tests after the fact, which is how most people work. But I think you've been an advocate for TDD, right?
Yeah. So I have a TDD skill which I recommend using, and this is quite timely because I have been thinking about it, but I haven't really posted about it yet. TDD optimizes for having a very small working memory, right? You write one test and that test is supposed to fail. And it means that even if you get distracted, you go for a coffee or something, you go for a long walk, when you come back, the test is still failing, reminding you of where you are in the implementation and guiding you to the next thing. Agents don't need that. The thing that's great about agents is that they have a much larger working memory than humans, right? They can actually hold a lot more in their heads than humans can currently, which is very useful. But they don't have an infinite working memory. And TDD, it's sort of aiming at the wrong problem, I think.
But the thing that agents really do need is that they need to have feedback loops. So they need to see what they're doing and how it's interacting with the environment of the code. They need to probe it all the time. And having an agent that builds the failure first, it's also very hard for an agent to cheat that. So not only are you forcing the agent to build its own feedback loops, the agent is providing proof to you that the thing is actually working as it goes. And even if I'm not using TDD directly, where it, you know, writes the failing test first, then fixes it, then refactors, I will often say, provide proof that your change does the thing it's purported to do. Give me TDD evidence, right? That it would fail without this change. And that's been really good for just improving the feedback loops essentially.
Because another thing with TDD that agents get wrong is they will often just write crap tests. They'll often just write, especially, tautological tests where the test is just asserting the implementation itself. It's just like a duplicate of it. You know, it writes a constant and then it says expect this constant to be this value. I mean, what's the point in that test? You know, it's just asserting the implementation. So yeah, I have a mixed relationship with TDD. I do still recommend it just because it gives you so much more confidence in what you're building from a human perspective. But yeah, I'm starting to see the counterarguments.
Let's talk about tech debt. Jared Friedman at Y Combinator wrote a tweet that I'll quote from him: "Technical debt used to be something you just had to live with with a sufficiently large codebase. No longer." And to which you replied, "Yes, now you can live with it even in a tiny codebase."
That's good. I read it out loud actually. You really gave the sense of that one. Yeah, it's just so easy for agents to produce rubbish, right? Even really smart, powerful agents, because they're unable to think strategically, they're just focused on what they're doing right now. It's very easy for them to produce tech debt. What is tech debt? Right, tech debt is anything that makes the codebase harder to make modifications to over time. A good codebase is one that's easy to change, easy to make a change in that doesn't result in cascading failures, right? So a codebase with solid test coverage and a good test suite is a codebase that's easy to change.
Yeah.
But it's so easy for agents to just make a codebase worse over time.
Yeah. And it's a really hard problem and it's one that you need a strategic mindset to think about, because one thing that I found works really well is automated review. So you have one implement agent to do the thing and then you have another automated review agent that sort of imposes your coding standards, that looks for these tautological tests, that improves the quality of the test suite over time. But then how do you know if the automated review agent is doing a good job, you know? And so even in tiny codebases, even in one-line changes, the agent can produce crap, you know. And so I think it's just something we need to live with and something we need to be in a constant battle against.
It's also not a bad thing. We bring a bunch of value when you understand what good code looks like, when you can recognize what tech debt is.
And it's also a problem that we've always had. You know what I mean? Like
It hasn't gone away.
Hasn't gone away. You know, this is just what I feel like. We're just having the same conversations we've had for 20 years. It's just there's this new elephant in the room.
I want to ask you about living in the UK and AI. This is a question that also came from one of the readers. Now that you're based in the UK, and outside of London, but you're now educating about AI, is being further away from Silicon Valley and the HQ of the labs making things easier or harder for you?
I'm really just trying to plow my own furrow, really. Like, what I realized quite early on is that I have no power to predict the future, right? Because I'm so far away from things. I'm just a person in the field working with this stuff. I have no way of knowing what's coming, right? I don't know whether the model's going to improve. I don't have privileged access to stuff. And so I'm just trying to focus on what's working right now. And because of that, I think that's narrowed my scope a little bit. That means I can just try to get my stuff working. And it's sort of quite surprising to me that it's working as well as it is, you know, because I don't have this privileged access. I'm just trying to make this one approach work. So I think, yeah, you're probably right. I probably would be able to do this stuff if I lived in San Francisco, but then I'd have to live in San Francisco. You know, I don't want to do that. That's miserable. You know, I've got a great setup here. My parents are just down the road. You know, I've got my son growing up in the countryside. So it is what it is.
Now, you're an educator at heart. How have you seen the business of teaching or educating software engineers change, and also how people want to learn? If you've observed any trends from before, like already when you started, I feel you were on at the time where online courses and learning over video became a lot more popular, as opposed to, let's say, a decade ago where it was maybe tutorials, and then before that it was books. Obviously they still exist, but there were just different preferences.
Yeah, it was around COVID time that video tutorials really took off. I think people wanted a much richer learning experience and I was kind of just after that wave, I suppose. I think that the way people have learned hasn't changed that much, right? And their desire for certain types of materials hasn't changed. I think it's very sexy, the idea that, you know, an agent can just come in and teach you everything. And that sort of works in some contexts, but really what you want is curation, right? You want a human to have come in, understand the flow of the information. I always think of information as kind of like a graph, right? You have a piece of information that's dependent on another piece of information dependent on another piece of information. And turning that graph into a linear path is how I think of my job, right?
I'm just trying to teach you, like, find Dijkstra's algorithm through the graph so that you can learn it in the most sensible way. And that level of curation is just not something that, again, that's strategic, right? That's not something that AI is particularly good at. So I mean, I've obviously made this huge pivot from TypeScript, from tactical stuff really, to this strategic layer, and it's working okay for me. I really can't speak for other folks doing this work, and I know that lots of people are not having this level of success, I suppose. So I think what it shows is that agents have just changed the game in terms of what people value and what people prioritize, and the industry has shifted in 7 months faster than I think it's ever done. You know, this is a huge shift. Doesn't mean we need to throw away our working practices, but it does mean that what we need to focus on is different. And I feel like I've been able to move with that quite well, whereas I think others just haven't because they're focused on different things.
And I wonder if in your case it's also, with Total TypeScript and even before with TypeScript and some other things you shared, you were helping people use the very popular tool at the time. TypeScript was gaining market share. There were migrations happening from Java to TypeScript, from Python to TypeScript and so on, and so developers wanted to get really good, a lot of them, or the top 10% or top 20%, you name it, wanted to get really, really good with TypeScript and they were looking for efficient ways of doing it. Now AI is here, is changing how we work as software engineers, and I think it's particular that building software is valuable, but there's a question of how do I use these tools more efficiently, which is more pressing right now than how do I write TypeScript efficiently, especially with the agent. So I wonder if you've kind of just, a little bit, how you pivoted from voice acting, which you couldn't do from outside of London, to a thing that you could do outside of London, which was still teaching. You've just pivoted to teaching a different area, which right now is, again, on so many people's minds.
I think I've just been lucky, basically, of choosing the right thing at the right time. It would have been very easy for me to, and it actually took quite a fair bit of convincing to move into AI. Like back a couple of years ago, it was Joel, my business partner, who was pushing me to actually go, you've really got to try this. It's actually pretty good and you can use it for all sorts of stuff. And it took about 3 months of me actually trying it and failing and trying it and failing before I realized, okay, this is great. I just feel quite fortunate that I've landed in the right place at the right time. And I try not to narrativize it. I try not to think, well, well done, Matt, you've been so smart, you know, making the right play at the right time, because I've made several mistakes as well, and I could have easily found myself in a different zone. And I mean, that's no bad thing. I would just get back to being an engineer. That's what I love, too.
Putting yourself back into the shoes of when you were someone just starting out in the industry: today, for people starting out in the industry, early career, junior folks, what would you recommend them for tactical things to do? Like they will know, like, look, I want to get that experience. I want to get that judgment, that taste, those fundamentals. You'll need to get repetitions in. If you found yourself in those shoes, how would you approach, like, I want to be a builder, a software engineer with all these AI tools and whatnot, which is now confusing because now there's a mix of: do I use these AI tools just to do stuff for me? Do I get in the fundamentals, which is slower, and so on?
Yeah, I mean, I would love to be a junior right now. I would love to be in the exact position I was in like 2014
where I was building these tools for my students, right? I actually got really nostalgic for it the other day. I thought I'd love to go back and do some singing teaching because just the ability to, like, I could finish a lesson and then just prompt the agent, okay, this tool didn't quite work in that way. I could maybe modify it a little bit and, you know, see it working. I just think the right thing to do is to use these agents as much as possible because that's how people are going to be working now.
And I think the thing that I find valuable about my skill set is you're constantly in touch with the changes that are happening. Grill me, not only you're having a discussion with a senior developer, right? That's beneficial for the developer, but it's also beneficial for you, keeps you thinking about these deeper ideas.
And the absolute rubbish that I was churning out, you know, with my spectrogram analysis tool, that would have been so much better if I had an agent to work with. It ran like a pig, you know, like its performance was absolutely terrible. If I'd have been able to say, okay, this frame rate has dropped to 10 frames per second, how do I fix that? It would have seen the six nested for loops and gone, okay, maybe you should do something different there.
So, I think that there's never been a more empowering time to work on this stuff, as long as you're interested in not only the code you're producing, but also the process of creating the code. There's never been a better time to be a kind of navel-gazing programmer, just constantly thinking about your own processes and being introspective.
So, it sounds like if you're motivated, you should be able to learn really fast compared to even before.
Absolutely. It's just about being curious, about being adaptable. And the people that I see who are thriving in this new environment are the same people who were thriving 10 years ago, because they're just interested in this work, interested in making better software and interested in their own process.
And interested in making better software. I want to ask you about gardening. A software engineer on X, Lauren, posted, I'll quote her: "Every team needs a gardener. Someone quietly watching the stream of PRs flowing into your codebase, noticing the smells, the lint suppressions creeping like ivy across your careful garden. A steady hand in tending the weeds that would otherwise engulf the garden." And to which you replied, "I'd argue the only thing your team needs are gardeners."
You probably do need a couple of other people as well.
Yeah. Yeah. But more specifically, I want to ask you about this concept of gardening.
I actually really love how Lauren described the weeds taking over the garden and getting them out. I think I made a tweet a while ago that we are, this was when I was sort of thinking about Ralph and sort of the agent looping over stuff, we are essentially just Ralph's platform team, right? That's what we are now. We are our agents' platform team. We are trying to build the environment for them to succeed. That's exactly how you should be thinking about it. Again, it's strategic.
And that gardener metaphor is nice because, you know, it's very easy for the garden itself just to suffer entropy, right? To gather weeds and to do all that stuff. So understanding and diagnosing that stuff before it becomes a problem in your own codebase is an essential skill and might be the essential skill, right? As long as you can queue up work for agents, as long as you can build these loops, now that we're starting to see these processes where agents improve the codebase based on bug reports and feedback, that feels to me like really cool work and noble, interesting work as well.
We talked about some great standout software engineering that you learned from, that you got inspiration from, today. What skill sets, experience, approach do you think makes a great software engineer?
I'll use an example, which is Lars Grammel, who works at Vercel on the AI SDK, who I had a chat with the other day, and he is building an entire software factory for his extremely popular open source library that gets a ton of issues. We're talking about plumbing again. We're talking about gardening. We're thinking about the processes of software development.
And I suppose if I had to put it in a word, it would be introspection. It would be looking at yourself and the ability to take what you do and put that into something the AI can work with. You're essentially trying to put your process into words. And that's what I've been doing with the skills. That's what I've been trying to do with the automations I've been creating as well. I just look at what I'm doing and think, how could I do this better? And also, how could I encode this into this strange animal that I have in front of me? How can I make it work like I want to? And that attitude has been really, really helpful for me, and it's something that I value in Lars and I value in all the people that I work with when they approach agents.
And then as closing, what is a book that you would recommend, or multiple books?
I'll go with The Pragmatic Programmer, A Philosophy of Software Design by John Ousterhout, and I'd say the first three chapters of DDD, the Eric Evans book, the ubiquitous language one. That one in particular, it's really great for the ubiquitous language concepts, the domain modeling. The actual sort of encoding it into code I'm not such a huge fan of, but those three are the big three.
Awesome. Matt, well, thank you. This was really interesting and really fun.
Great to finally be on the podcast. Yeah. Meet the famous guy himself. It's great. We've met before, obviously, but it's great to be here.
It was so nice to sit down with Matt, and I have to say, knowing that he was a voice coach and actor makes me understand how he talks so smoothly and how he's so pleasant to listen to. Probably the most amusing part from this conversation was how, as Matt was searching for how to work better with AI, it wasn't modern approaches that he found really useful. Instead, he went back to classic software engineering books: The Pragmatic Programmer, A Philosophy of Software Design, and Domain-Driven Design. There's some irony as to how the best practices documented 20-plus years ago, like tactical versus strategic programming in this book, not only do they still work, but they become more important when writing code with AI agents.
A related point I want to emphasize is the importance of leading words with AI. When Matt started to use terms like tracer bullet or vertical slices, the model started to follow his ideas better in planning. And if you think about it, this makes sense, because software engineering literature is part of LLM training, so these terms are also part of the model's priors. Just as interestingly, using the right words for describing your problem is not a new concept. For example, when I had Kent Beck on the podcast, he talked about how 30 or 35 years back, when he and Ward Cunningham had a thesaurus on their desks, they used it to try to find the best words for the specific thing they were describing. This was just another full circle moment on how words do matter.
Finally, I appreciated Matt's push on how you should want a clean codebase. Not just because it's easier for humans to navigate, although I think you really want to do it for that as well, but also, conveniently, agents do not have a long-term memory, and they will look at your codebase for the first time on every new run. And it's much easier to get around inside a well-structured codebase than one that is really messy.
Check out the show notes below for an interview with John Ousterhout, the author of A Philosophy of Software Design, a book I really love, and related deep dives for AI engineering and context engineering. If you like this episode, please make sure to be subscribed on your podcast player, and submitting a rating is always appreciated. Thanks and see you in the next
Article published · Updated
