Matt Pocock on Grilling Your Agent: Why Old Software Fundamentals Matter More with AI

Open on YouTube ↗
Overview

Matt Pocock built Total TypeScript into a multi-million-dollar course business. More recently he created a set of AI agent skills, including "grill me" and "wayfinder," that spread quickly among developers. In this conversation on The Pragmatic Engineer podcast, he explains how he got into tech without a computer science background, and why his search for better ways to work with AI coding agents led him back to decades-old books. His argument is that AI has taken over much of what he calls tactical programming. That leaves humans responsible for strategic thinking, and the classic literature on software design turns out to be a practical way of steering agents.

31 min read
4:09

From voice coach to self-taught developer

Before he became a developer, Pocock spent six years as a voice coach. He started as a singing teacher while at university, set up his own small company, and later did a master's degree in voice. After that he taught Shakespeare at drama schools, taught accents, and coached public speaking, including a couple of large engagements teaching consultants to deliver speeches. He says he left because doing the work at a decent level meant living in London. He tried it for about two years and hated it. He wanted to go back to the countryside where he grew up, so he taught himself to code in order to have work he could do remotely.

He had already been writing software for his students. There were small flashcard apps, and his first project was, in his words, the most ambitious thing he has ever attempted: a web audio analyzer that displayed a spectrogram of a student's voice so he could see which resonant frequencies were present. He admits it "ran terribly," but it made his lessons a little better. He looked at job postings, decided his JavaScript and CSS were enough to get started, quit his job, took a couple of months off, and was hired around 2017. He notes that getting a job in the UK was easier then than it is now.

He calls his communication background an "unbelievable advantage." In interviews he could come across as a reasonable person, unlike some candidates straight out of CS degrees. He had very little technical knowledge at first but could explain technical ideas to people, and once his knowledge grew he moved up quickly across several companies. His first employer was a tiny company with a couple of senior developers he found inspiring. One of them, who wore sandals and had lived on a canal boat, had Pocock install CentOS 6 on his Windows PC on his first day, because that was what their application ran on. That company ran into financial trouble, and Pocock then moved through several agencies.

TypeScript as a coordination tool

He was already arguing hard for TypeScript by his second job. The agency was building a learning management system for a car manufacturer. The front-end team was small and full of bugs. The back-end team, based in Portugal, kept changing API contracts without warning. The front end adopted TypeScript to keep the two sides in sync. Pocock says their velocity rose so much that they overtook the back-end team, and eventually people were moved off the front-end team because it was working so quickly. He describes this as his origin story with TypeScript.

Open source, XState, and the Stately job

In his fourth job he worked on a complex application where two people on a video call could move around a house together in real time. It involved a lot of state that had to cross the network boundary. He used XState version 4, the state machine library, and considered the project a success. He then began building tooling to make XState more type-safe, including a CLI. That caught the attention of David Khourshid, XState's creator, and Pocock joined the core team.

He says the work put him in contact with developers at a level he had never seen, naming Khourshid and Mateusz Burzyński (Andarist). When Khourshid raised funding for Stately, a company betting on state charts and visual programming, Pocock was hired. It was his first job paid at American rates, which he says changed his life and how he thought about money and flexibility. At Stately he also did advocacy work. He still thinks state charts are incredibly useful for certain kinds of work, although he has backed off his belief in them somewhat, especially now that AI is here.

Three days a week at Vercel, and the Total TypeScript launch

His advocacy caught the attention of Vercel's developer education group, which Lee Robinson led at the time. Pocock joined under Jared Palmer. While trying to force XState to be type-safe, which he calls "a mostly impossible job," he had learned many advanced typing tricks. He also missed teaching. He began posting two-minute TypeScript tips on Twitter and found they spread in a way nothing he had done before had. One Sunday he recorded 13 to 15 tips and scheduled them over the next few weeks, and his follower count went from about 4,000 to about 10,000.

That made him confident a course could work. He negotiated an unusual contract with Vercel: three days a week, initially for three months. Vercel had wanted him full-time. He admits it is slightly embarrassing to say, given that it is many people's dream job, but he treated Vercel as a stable 9-to-5 while he tested the course idea. During his time there he wrote some of the initial documentation for Turbopack and flew to San Francisco for Next.js Conf, where it was announced.

Around two months in, it became clear he couldn't stay. A pre-release sale of Total TypeScript earned roughly 30 to 40 times his Vercel pay, and he describes that pay as already very good. He says he loved working at Vercel and would probably go back at some point, but he had no real choice.

18:42

Building Total TypeScript and its business model

Pocock says he tries very hard not to work weekends and that the "996" culture "turns my stomach." He wants a life where he spends most of his time with his family, and he asks listeners to read his later decisions in that light.

He built Total TypeScript with Joel Hooks, who is known for Egghead and has worked on Kent C. Dodds's courses. The full course came out around January or February 2023 and, by his account, reached seven figures in revenue very quickly. That revenue was split between him and Hooks, and expenses came out of it. The host notes that he later publicly shared a milestone of $2.5 million in total revenue. For Pocock, this was the high-leverage work he had wanted since his singing-teacher days: do the work once, step back, and keep earning while spending time with his family.

He wants his income to come only from products people buy. He avoids sponsored content and has been trying to take down his GitHub Sponsors page. Many buyers pay from company education budgets, a market Hooks pushed him toward. Pocock says he had no idea how much money was "sloshing around" in the industry. They also offer an extended refund policy.

23:22

What AI did to technical education

Asked how AI has affected education, Pocock separates teaching into two layers. There is the "what," meaning syntax and knowledge, and the "why," which he calls wisdom. The first is usually the vehicle for the second. He argues that knowledge is now very cheap to acquire. He even has a "teach" skill that teaches you whatever you need. Wisdom, he says, has gotten no easier to learn, and you will hit the same problems without it whether or not you use AI.

He says Total TypeScript revenue has gone down, and he thinks people teaching that kind of material will struggle. It took him a long time to find his place in AI. He felt he lacked the credentials of an OpenAI researcher. Starting in late 2023, he made courses about building AI into applications, which seemed like a natural move for a front-end developer. He is proud of that material, but he did not see the returns he expected and decided it was the wrong bet.

26:11

The December shift: tactical versus strategic programming

The turning point was what the host calls the winter break when everyone came back "AI-pilled." Pocock associates it with Opus 4.5 and OpenClaw: people had time off, used the models heavily, and saw that something had changed. His conclusion was that agents had become good enough to delegate to. You could build structures around them so they handle the tactical work (syntax and knowledge) while you handle the strategic, long-term thinking.

He borrows the tactical/strategic distinction from John Ousterhout and says AI has "largely eaten tactical programming." That left room for a course aimed at the strategic layer. At the time he was also experimenting with Geoffrey Huntley's "Ralph loops," which repeatedly run an agent toward a goal. The host recalls a popular video of Pocock's that contrasted traditional top-down planning, where plans break once implementation reveals new work, with a Ralph loop driven by a markdown file of tasks that grows as the agent works through it.

Pocock says his two years building agents into applications were good preparation. Building those apps forced him to think at a lower level than harnesses usually expose: how data gets in, what shape it has, whether it goes in the system prompt or the user prompt, and how to compact or clear it. When he moved to Claude Code, which had been out for five or six months at that point, those questions felt familiar. The processes he built looked to him like finite state machines, much like his XState work, with agents firing events that send control back to earlier states. He ran experiments on his own projects, including a custom video editor he describes as a large codebase and several open-source projects. Sometimes he dropped his process and went back to the default setup, and he says he noticed "a huge difference."

30:47

Skills as a distribution mechanism

He settled on skills as a way to share these processes. He describes a skill as "just a folder of markdown files." Some are model-invoked, meaning the agent decides to use them. Others are user-invoked, meaning only the user can call them. He had seen skill sets like Superpowers, Claude Code plugins, and GStack, so he published his own and moved on to other work. When he checked back, the repo had more stars than anything he had ever made, with almost no promotion. At the AI Engineer conference in London in April he gave a talk called "Software Fundamentals Still Matter," which he says is now around 1.2 million views. By his count the repo has about 230,000 stars, making it the second most-starred skills repo and somewhere around the 20th to 25th most-starred repo of all time. He says this was the second time in his career he felt that kind of momentum, after the TypeScript tips.

His design goal is the simplest possible set of skills that people can easily audit and tinker with, which makes them more likely to be adopted at work.

33:18

"Grill me": getting the agent to interview you

The host describes being grilled by the skill while building a simple API endpoint. The endpoint needed to tell an authenticated caller whether an email address belonged to a newsletter subscriber. The skill asked 35 questions: whether to use a bearer token or put credentials in a JSON body, how strictly to enforce rate limits, and so on. The host found it annoying and impressive at once. It had been a long time since they'd had such an involved design discussion, and it forced them to make decisions and sometimes go research the options. The host adds that as long as people keep learning while using AI they will be fine, and that trouble comes when the learning itself gets outsourced.

Pocock says "everyone's got a grill me story." The skill is tiny. It just tells the agent to interview you relentlessly about the topic. He got the idea from someone named Tariq who works on Claude Code and suggested that having the agent interview you produces better results. He says the skill has "weird emergent behavior": models start thinking outside the box and suggesting ideas. It reminded him of design discussions with the senior developer at his first job and with Andarist at Stately. That led him to ask how to "tickle the right latent space" so the agent behaves more like a real senior developer, since better conversations lead to better output.

He argues that people underestimate the communication gap between themselves and the agent. The belief that you can "just trust the model" ignores the fact that no model, "even Mythos," can read your mind. If you just ask for code, you get something misaligned with your values, because the agent doesn't know what you consider important. So grilling is about more than implementation details. It establishes what is in scope, what is out of scope, and what matters. In his words, it is the agent getting to know you.

40:56

Smart zone, specs, tickets, and the night shift

Turning that conversation into code raised the problem of context size. Pocock credits Dex Horthy's idea of the "smart zone" and "dumb zone." Every token competes for attention, and as context grows the model loses connections and makes more mistakes. Pocock estimates the smart zone at roughly the first 150,000 tokens of current frontier models. He says the window size doesn't matter, because what counts is the raw number of tokens and attention relationships. Beyond that, performance slowly degrades.

The question became how to spread work larger than 150k tokens across several sessions. Ralph loops were one answer: tell the agent to make the smallest change that moves toward the goal, then clear context. State survives in the codebase and file system, not in the model. To make this more stable, he defined two kinds of documents. The first is a destination document, which he used to call a PRD and now calls a spec, and which says when the work is finished. The second is a set of tickets, one per session. His skills turn a grilling session into a spec, then turn the spec into tickets. A spec might cover 30 or 40 tickets, and an implementation loop works through them.

The loop is meant to run with the user away from the keyboard. Pocock describes "day shift and night shift": plan during the day and let agents work overnight, so you wake up to clean code to review. He was tired of constantly switching between terminals, which he admits he still does to some extent. He wants long planning sessions, a couple of hours of autonomous agent work, and focused blocks of his own time for other tasks. He says the December models were the first good enough to delegate to that way.

45:08

Wayfinder: grilling across many sessions

Planning can also outgrow a single context window. Pocock's example is "build me a Stripe clone," which he says cannot be planned in 150k tokens. He asked what each grilling session needs to work well: its own purpose, what has been decided so far, and what other sessions are in progress. His answer was a map. The map is a central record of decisions, with milestones and a "fog of war." Each grilling session reveals more of the map, and he describes the overall structure as a directed graph that you walk until you reach the destination.

Each session is a ticket on the map. Tickets aren't limited to grilling. Some are prototyping, research, or arbitrary tasks like provisioning infrastructure. He says he has had maps with 50 to 100 tickets. He also uses Wayfinder for non-technical work such as course planning and building a garden office, and says he is interested in how well these skills transfer to other areas of life.

Why agents are good at software, and bad at what isn't text

The host brings up Hillel Wayne's research comparing software engineers with traditional engineers. One difference Wayne noted is that software, unlike physical materials, behaves the same way every time you run it. The host wonders whether LLMs bring back the kind of variance other engineering fields have always dealt with.

Pocock's explanation for why agents do well in software is that everything is text. The inputs are code, docs, and instructions. The outputs are code, test results, type-check results, and lint output. Anything that isn't text is, in his words, "just garbage from the agent." A one-shot UI demo can look impressive, but if a hover animation looks wrong, it is hard to get that to the agent. You could record a video, but he says vision isn't good enough yet. He guesses that other professions will get better results if they can express their work as text, for example through simulations or something like a linter for architectural diagrams. He is currently connecting all the services he uses to agents.

51:03

Rediscovering the fundamentals

The host asks about the claim that AI requires setting aside prior knowledge. Pocock says he believed that at first. He tried spec-driven development, a term he has mixed feelings about because it covers too much, on the theory that "English is the hot new programming language": keep editing a persistent spec and let the agent update the code. He says the results were worse than coding by hand and weren't improving. Each cycle of changing the spec made the code worse. "You're not supposed to look at the code," he says, but he did, and it was garbage. He reasoned that a bad test suite gives the agent bad signals, just as it would a human.

He then opened The Pragmatic Programmer, which had been sitting on his shelf still in plastic. Its section on software entropy helped him frame the problem: agents produce software entropy faster than ever. He says almost every line read as if it had been written for today, citing "don't outrun your headlights," working within feedback loops, programming by coincidence, and tracer bullets. Because the book has been around for about 25 years, he figured its concepts were probably in the models' training data.

Tracer bullets addressed a specific failure. Even with Ralph loops, agents would build the whole database layer, then the whole application layer, then the whole React component library, and only connect them at the end. By then, decisions in the database had already affected the front end without any feedback. A tracer bullet, or vertical slice, builds a thin path through every layer first so feedback arrives immediately.

When he used these phrases in prompts, the agent began repeating them in its reasoning, saying things like "Okay, I'll turn this into a tracer bullet." He calls these "leading words": short phrases repeated a few times in a skill or prompt that change the agent's behavior. He began going through other books for more. From Ousterhout's A Philosophy of Software Design he took "deep modules," which he calls "a massive one" for him. The host compares leading words to professional jargon, which makes communication faster and reduces misunderstandings between experts.

56:38

Ubiquitous language and "Memento-driven development"

Pocock took the idea further. If established terms work, what about terms for his own application? Agents are very verbose, and he singles out Opus 5. So he turned to Eric Evans's Domain-Driven Design and its "ubiquitous language." He built a variant of grill me, which he admits is "terribly named" grill with docs, that develops a domain language during the grilling session. He says the difference is "night and day."

His example comes from an app with "ghost" and "real" lessons. When a ghost lesson inside a ghost section inside a ghost course becomes real, the section and course also have to become real. With the agent's help, that became the "materialization cascade." He says agents are good at coming up with such terms, and he has a separate domain modeling skill for it. When the vocabulary also appears in the code, the agent can find relevant functions with a simple grep. He has worked DDD into every part of his setup and says he already has The Mythical Man-Month too.

His overall framing is to imagine a coworker who wakes up every morning with no memory, like the protagonist of Memento. A human can work around a bad codebase by building up memory over time. An agent starts fresh every session, so the codebase has to be optimized for a permanent new starter. He argues that software fundamentals have always aimed at this, and that AI is mostly emphasizing rules developers already knew they should follow but often didn't.

1:01:14

How to learn strategic programming, and what juniors face

Asked how an engineer can find the fundamentals that matter, Pocock says strategic programming has always been hard to learn because the feedback loop is so long. Someone who leaves a job after six months may never see their strategic mistakes, which can take nine months to surface. He compares it to mixing music on a desk full of sliders, such as the number of deployable units (turn it up for microservices, down for a monolith), except that "you can't hear what's wrong until 9 months later." He thinks AI shortens that loop because the extra code brings mistakes back sooner. His advice is to treat code as the environment the agent works in, keep improving it, think at that level constantly, and read the books to gain the vocabulary.

The host wonders whether faster fixes also make lessons less memorable, since painful incidents, like being burned by a lack of idempotency, are what teach people. Pocock raises a harder question: if strategic knowledge is now so valuable, why would a company hire someone without it? He mentions that Uncle Bob, in a recent interview, suggested hiring juniors and treating them like agents, delegating tactical work until their mistakes come up. Pocock calls that a huge waste of money now that tactical work costs less than minimum wage in many countries. He says he doesn't know the answer. He only knows that strategic understanding is worth more than ever.

On a listener's question about convincing non-engineering stakeholders to invest in fundamentals, he says the same question could have been asked ten years ago about tech debt. He suggests measurement. Start with observability across every agent in the organization, tracking success and failure rates, which he says would have felt invasive for human developers but is fine for a paid service. Have someone analyze the data, find which repos do better, and spread those lessons. Organizations should share a common set of skills so people can contribute and A/B test different workflows. His own setup includes a loop that runs an "improve codebase architecture" skill every morning and proposes an improvement he can turn into tickets with one click. He floats spending something like 20% of time on "the factory that builds your software," and adds, with a laugh, that you might keep that work quiet until it has shown results.

1:09:14

Moving from local to cloud agents

The host quotes Pocock's tweet: "I'm moving away from my local dev setup. Makes zero sense to me now." He explains that people keep asking how to make skills collaborative, such as tagging a colleague into a grilling session. Every developer now has something like a hundred terminals, and he thinks those should be available to the whole organization, inside tools like Slack, Discord, Teams, or Linear. He doesn't work with a team himself, but he chats with his remote box over Discord, including on the train to this recording, to build course material and fix bugs students report. He can port-forward to see the dev server, and he views an expensive laptop running this work as wasted compute. The remote box is always on, so it can run schedules, including a morning stand-up where the agent plans his day using his Discord chats. The only thing he now does locally is debug the remote bot.

He adds that local setups have their own problems, like piles of Git worktrees and several Docker containers per worktree, and the cloud lets you provision resources on demand. The host mentions that some companies with platform teams have moved full dev setups into the cloud and see most developers choosing cloud agents voluntarily, though front-end work is often cited as an exception. Pocock suggests tunneling to the dev server might solve that, but says he hasn't tried it.

1:12:46

When to plan and when to course-correct

The host asks whether fast agents make upfront planning less necessary. Pocock says it depends on the size of the work and how hard it is to undo. If a large feature goes wrong, the wrong code stays in the agent's context and shapes everything after it, and realigning afterward is expensive. So align first. For a five-line change or moving a button a few pixels, align afterward. His rule is to "shift right" wherever possible. His video editor has a feedback button that creates a GitHub issue. An implementation agent picks it up, a review agent checks it, and he reviews the result at the end. His decision tree: if it is small enough to align afterward, skip grilling. If it fits in one session, use grill me. If it spans many sessions, use Wayfinder. The host compares this to how tech companies used PRDs: skip them for trivial work, share them for team-sized work, and require sign-off for large projects.

On the criticism that this is just waterfall again, Pocock points out that he does a lot of aggressive prototyping before writing specs. Agents make it cheap to produce three or four versions of something, pick the best, and iterate, and he calls prototyping essential to spec writing. The host cites Grady Booch's view that waterfall was only ever a problem when planning took a year and implementation took three. Pocock says he doesn't mind criticizing waterfall as a cautionary tale anyway, because cheap labor makes agile a better fit for agentic work.

1:18:14

TDD: wrong problem, useful evidence

Pocock has a TDD skill and recommends it, but says he has been rethinking it. As he describes it, TDD is designed for small working memory: a failing test tells you where you were after a coffee break. Agents have much larger working memory than humans, though not infinite, so he thinks TDD is "aiming at the wrong problem" for them. What agents do need is feedback loops, and writing the failure first is hard for an agent to fake. Even when he doesn't use strict red-green-refactor, he asks the agent to "provide proof" that a change works and would fail without it, which he calls TDD evidence. He also notes that agents often write tautological tests, such as asserting that a constant equals its own value. He still recommends TDD because of the confidence it gives humans, but says he is "starting to see the counter arguments."

Tech debt at any size

Responding to Jared Friedman's claim that large codebases no longer have to live with tech debt, Pocock replied that now you can have it "even in a tiny code base." He defines tech debt as anything that makes a codebase harder to change over time, and a good codebase as one that is easy to change without cascading failures, for example because it has a solid test suite. Agents can't think strategically, so they easily make codebases worse. Automated review, where one agent implements and another enforces standards and removes tautological tests, helps. But he asks how you would know whether the review agent itself is doing a good job. His conclusion is that this is a constant battle and the same conversation the industry has been having for 20 years, now with "this new elephant in the room."

1:23:13

Distance from Silicon Valley, curation, and advice for juniors

On living in the UK, Pocock says he has no ability to predict the future and no privileged access. He is "just a person in the field" and focuses on what works today, which he thinks has kept his scope narrow. He probably could do more from San Francisco, but he would have to live there, and he prefers being near his parents and raising his son in the countryside.

On education, he doesn't think the way people learn has changed much. An agent teaching you everything sounds attractive, but what learners want is curation. He sees knowledge as a dependency graph and his job as turning that graph into a sensible linear path, which he considers strategic work AI isn't good at. He says his pivot from tactical TypeScript content to the strategic layer is "working okay" for him, but he can't speak for others, many of whom he knows are not doing as well. He says the industry shifted in about seven months, faster than ever. He calls his own success mostly luck: Joel Hooks had to push him into AI, it took about three months of trying and failing before it clicked, and he has made mistakes too. If it hadn't worked out, he says he would simply go back to engineering, which he also loves.

For people starting out, he says he would love to be a junior now, and imagines going back to singing teaching just to build tools for students between lessons. He would use agents as much as possible, since that is how work is done now. He thinks his skills help because grilling gives you a senior developer to talk to and keeps you thinking about deeper ideas. His old spectrogram tool, with its "six nested for loops," would have been far better with an agent pointing them out. The key, he says, is caring about how you produce code as well as the code itself. It is a great time to be a "navel-gazing programmer," and the people doing well now are the same curious, adaptable people who did well ten years ago.

Gardeners and introspection

The host quotes Lauren, a software engineer, saying every team needs a gardener who watches incoming PRs and pulls weeds like creeping lint suppressions. Pocock had replied that teams need only gardeners, and adds with a laugh that you probably need a few other people too. He describes developers as "Ralph's platform team," people building the environment agents succeed in. Diagnosing entropy before it becomes a problem, and building loops where agents improve code based on bug reports and feedback, may be the essential skill now.

Asked what makes a great engineer, he points to Lars Grammel at Vercel, who is building a "software factory" to handle the flood of issues on the AI SDK. If he had to choose one word, it would be introspection: looking at how you work and putting it into words an AI can use. That is what his skills and automations are, he says. He keeps asking how he could do something better and "how could I encode this into this strange animal that I have in front of me."

He closes by recommending three books: The Pragmatic Programmer; John Ousterhout's A Philosophy of Software Design; and roughly the first three chapters of Eric Evans's Domain-Driven Design, for ubiquitous language, domain modeling, and encoding that model in code. He says he is less of a fan of the rest of that book.