Dex Horthy on Context Engineering, and Why HumanLayer Shut Down Its Own Lights-Off Software Factory
The Pragmatic EngineerDex Horthy, CEO and cofounder of HumanLayer, is credited with coining the term "context engineering." His team also tried running a software factory in which agents shipped code that no human read. They started it in July 2025 and shut it down by November. In this conversation on The Pragmatic Engineer podcast, Horthy explains what context engineering means and why context windows degrade as they fill. He describes how his team uses automated loops, why they believe the "dark factory" fails, and the workflow HumanLayer is building instead. His position is that he is heavily invested in agentic coding and still keeps humans reading the code. He argues that current models are not trained to write code that stays maintainable over months.
From moon-crater pathfinding to "the thing that builds the thing"
Horthy studied physics as an undergraduate and decided against academia. He saw three exits from physics: a PhD, finance, or programming. His first programming experience came at 17, during a high-school internship with NASA researchers at the Jet Propulsion Laboratory. The team had a new high-resolution topographic map of the Moon's south pole. Craters there are deep enough that some have never seen sunlight and may hold frozen water. His task was to find a route between two points that stayed within a rover's maximum incline and decline. He had never opened a CS textbook, so he wrote what he calls a "really naive bad version" of Dijkstra's algorithm.
After half a CS minor, he joined an API platform team at Sprout Social in Chicago. Within a few months he concluded that the most valuable people in the company were the ones building the developer platform: CI/CD, sandbox environments, preview infrastructure. He says he has been "obsessed with software factories" ever since. He is surprised when developers say they hate CI/CD. In his view, lazy engineers should want to build "the thing that builds the thing," because that is the highest-leverage work.
A stint at the consumer fintech Aspiration followed. The VP of engineering who hired him left within three months, and Horthy served as acting CTO for a while. He came away convinced he is "a B2B guy." He then spent about four years at Replicated. The company built its own container orchestrator before Kubernetes and Docker Swarm matured. The goal was to let SaaS vendors ship installable versions of their products that customers could run in their own data centers or VPCs, much like GitHub Enterprise. After two years in core engineering, he says he kept arguing with the CTO about the build pipeline. His manager told him to stop fixing CI and do his tickets.
He moved into a customer-facing engineering role. Replicated's customers included HashiCorp, DataStax, Puppet, TravisCI and CircleCI. In his first three months he met every stalled prospect in the pipeline, and he recounts closing about 12 deals. He grew that organization to around 25 people. The team later shrank after zero interest rates ended, and the company shifted toward a more product-led, self-service model. Horthy then became a product manager, carrying a "laundry list" of usability fixes from his customer work.
He also gives a personal reason for the customer-facing years. He felt introverted and socially awkward through his twenties. His uncle, a music producer who had worked with Randy Newman, once told him that greatness comes only from making something the only thing you do. Horthy decided that instead of reading self-help books, he would "make it my freaking job to just talk to people." He says the move also forced him to become organized, which did not come naturally. He recommends that every engineer spend a year or two in a customer-facing role.
Metalytics, the pivot, and 12-factor agents
Around 2020, Horthy and a friend in Chicago started Metalytics, a data engineering company. HumanLayer is technically the same company after a pivot. He describes the data-tooling market of that era as a funding boom. By 2021–2022, people realized the total addressable market was smaller than assumed, and both fundraising and customer acquisition became hard.
When his cofounder burned out and left on good terms, Horthy started experimenting with AI agents. Frameworks such as LangChain and CrewAI were popular then, with large Discord communities and plugin ecosystems. He saw a missing piece: controlling which tools an agent calls. In a chatbot, you can show approve and deny buttons. Horthy cared more about "outer loop" or proactive agents that run in the background and are triggered by events. He cites OpenClaw's heartbeat model as the biggest current example. He would not trust such an agent unless it could reach him through Slack or iMessage, and unless he could deterministically approve, deny, or deny with feedback. HumanLayer entered Y Combinator in fall 2024 with an API platform for this. He compares it to PagerDuty, but routing approvals rather than on-call incidents.
He then talked to about a hundred engineers who were shipping AI into enterprises on six-figure contracts. Most had tried the popular frameworks for a month or two and dropped them. They wrote LLM calls by hand and built systems that looked more like pipelines and workflows than open-ended "tools in a loop." A major influence was his friend Vaibhav at Boundary, which built what Horthy calls "protobufs for AI." Vaibhav framed every step in an AI workflow as tokens in and tokens out. The engineer's job is to choose the input tokens that maximize the chance of good output.
Horthy distilled these conversations into twelve principles and published them as a GitHub repository. It stayed on the Hacker News front page for about two days. The factors include: natural language to tool calls, own your prompts, own your context window, tools are just structured outputs, unify execution and business state, launch/pause/resume with simple APIs, contact humans with tool calls, own your control flow, compact errors into the context window, small focused agents, trigger from anywhere, and make your agent a stateless reducer. He notes someone corrected him that the last one is technically a "transducer." Of all of them, factor three, "own your context window," is the one he considers most durable. Whether the system is agentic or a single pipeline step, the only lever on output quality is careful crafting of the input.
Naming context engineering
Horthy wrote the 12-factor document in March, published it in April, and gave a talk on it at AI Engineer in early June. About one to two weeks later, Shopify's Tobi Lütke posted that he liked the idea of "context engineering." A week after that, Andrej Karpathy argued for thinking in terms of context engineering rather than prompt engineering. Horthy jokes that depending on the day, Gemini will credit him, Lütke, or Karpathy. He adds a caveat: he did not invent the practice. He saw what the hundred builders he interviewed had in common and put a name on it. He considers clean vocabulary important in a field full of hype and jargon.
His definition is that context engineering is "de-abstracting" the layers built on top of LLMs. RAG, memory, agentic history and structured output are all, in the end, different ways of passing tokens into a model and usually getting structured output back. Off-the-shelf agent and memory frameworks can get you to about 80% and a good demo. To reach 95% or 99%, you have to go down a level. You decide what goes into the context window and in what order, which can vary by model.
Asked why the idea caught on when it did, Horthy says context always mattered. It took time for enough builders trying to sell reliable products to discover that token-level thinking breaks through the quality ceiling. That means chains of LLM calls and workflows rather than a single model with tools and a system prompt. The approach needs more work and deeper intuition about LLMs, but it offers many more levers.
On cost, he advises "make it run, make it right, make it fast." First check whether the best available model solves the problem, and whether people want the result. Engineering time is the real bottleneck. Building evals and tuning prompts costs more than simply paying for a smarter model, until you reach something like millions of requests. At that point, context engineering lets you split a task into several calls. Some calls can run on far cheaper models; he mentions an open-weights model at roughly a thousandth of the cost of Opus. The expensive frontier model is then reserved for the steps that need it. He connects this to Eli Goldratt's The Goal: find the actual bottleneck. Latency and cost may become the bottleneck someday, but usually not at the start.
Harness engineering: inner and outer harness
Around October or November 2025, Horthy posted about what he called harness engineering. He later learned that Viv, now at LangChain, had written about it a few weeks earlier. His framing was this: you use context engineering when you build an agent, and harness engineering when you use one. The question is how to engineer against the integration points of a harness like Claude Code or Codex — commands, MCPs, skills, codebase organization — so every turn produces the best possible result. It "raises the floor."
The term quickly blurred. Some people took it to mean building a harness, and others took it to mean building around one. Horthy prefers Martin Fowler's framing. The inner harness is what a product like Claude Code, Codex or Amp exposes: tool definitions and integration points. The outer harness is what the human adds to adapt it to their codebase and languages.
He is surprised that "context engineering" still means roughly the same thing a year later, given how little AI writing from 15 months ago still holds up. His explanation is that it rests on how transformer attention works. It will stay relevant until post-transformer or linear-attention architectures arrive, and he says no one knows when that will be.
The physics of context: instruction budgets and the "dumb zone"
Horthy acknowledges that longer context windows let you talk to a model longer. But a million-token version of a model, such as the long-context variants of Opus, is not a smarter model. The model's intelligence determines how well it can attend to everything in context and pick out what matters for the next tool call. He cites a 2025 study finding that frontier models could follow about 150 to 250 instructions before adherence dropped off quickly. He adds that the models were older and a later study reportedly showed improvement, though he has not looked at that data himself.
He divides context engineering into two budgets. The information budget is the familiar one: use RAG to fetch the relevant pages instead of loading the whole book. The instruction budget is less discussed. Too many instructions, especially conflicting ones, cost the model computation. That includes changes of direction mid-conversation ("actually, I don't want any of that"), since the model has to work out what to ignore. When those instructions sit far back in the window and get only partial attention, the chance of the model honoring exactly what you said 100,000 tokens ago drops significantly. He does not claim a mathematical proof. He says he is "not a PhD in machine learning," but attention is quadratic, and more content spreads it thinner.
This leads to his "smart zone" and "dumb zone." In November his team spoke of the first 40% of the window. Once million-token windows arrived, they switched to absolute numbers. As training wheels, he suggests about 100K tokens for smaller models and about 200K for heavier models like Codex and Opus 4.8. Past that, results may degrade. He personally sometimes works up to 250K–300K tokens, occasionally 400K, when his intuition says the task can tolerate it. He attributes the core insight to Geoffrey Huntley's Ralph Wiggum work: the less context you use, the better the outcome.
His clearest warning sign is a model deep into a session trying to make a test pass. It tries increasingly extreme things, up to something like deleting your env file. At that point he writes everything done so far to a file, or uses built-in compaction. He then starts a fresh session at 30K–50K tokens with an explicit instruction to fix the test without being "stupid about it."
He treats this as a skill learned through suffering, not from textbooks. He quotes his friend Jake from Netflix, who said at AI Engineer that you learn bad patterns by debugging them at three in the morning.
Trajectory, and why "You're absolutely right" means start over
The host raises Horthy's advice that when a model says "You're absolutely right," you should start over. Horthy says models produce this phrase when a user angrily points out a mistake, and then they often keep making the mistake. (He jokes that Opus's newer version is "You're right to push back on that.")
He lists four properties of a context window: its size; whether it contains incorrect information, such as a reasoning trace that concluded something false; whether information is missing; and its trajectory. Trajectory is the history of what the agent has actually done. Models are autoregressive and predict the next message in the conversation. If the agent previously made a change, ran the tests, fixed them and reported back, it will likely follow that path again. If it skipped the tests, it is on a different trajectory. In his "No Vibes Allowed" talk he gave this example: the model makes a mistake, the user yells, it makes another mistake, the user yells again. The most plausible next message is another mistake. That is when to reset.
Loop engineering: back pressure and slow loops
Horthy traces current interest in loops to Ralph Wiggum. About a year before the recording, Geoffrey Huntley visited San Francisco and showed that he had run Sonnet around the clock. He spent about $6,000 over six weeks and built an entire Gen Z–themed programming language, including a stage-two compiler written in the language itself. Horthy's main lesson from it is back pressure: letting the model check its own work by automating feedback through linters, unit tests and similar tools. A programming language worked well for this because it is endlessly verifiable. If you can make a problem verifiable, you can treat it largely as a black box and let it loop.
His simple example is speeding up slow CI during a release. He gives the model a goal and a fixed set of steps: change code, commit, push, launch a subagent to watch the job, then decide what to do next. He compares this to the goal commands in Claude Code and Codex and to auto-research, where a prompt keeps the model trying until results improve.
He warns against getting carried away and rebuilding everything as a grand agent-first "dark factory" meant to be your infrastructure for five years. Few teams can accept a 60,000-line PR from a three-day Ralph loop that fixed every lint error, because no one will review or sign off on it. What excites him are "iterated" or "slow" loops. A GitHub Actions cron job runs every night with a simple instruction: run this linter, fix one thing, commit and push. The team wakes up each morning to one PR that makes the codebase a little better.
This can grow in two ways. You can add checks: React Doctor for the frontend, or a pattern with no deterministic tool at all. For one such pattern, his colleague Kyle wrote a description of good and bad examples. It is "prop narrowing," which makes unnecessarily optional props required so code is easier to reason about. The team now wakes up to about four PRs, one per check. As confidence grows, you can also widen the scope from one fix to four. Kyle has shipped a skill so others can build these loops. More broadly, Horthy says a good loop is one that no human has to start. It is triggered by a Sentry alert, a support ticket, a PM's ticket, a failing test or a cron schedule, and it follows a defined workflow.
The lights-off factory that HumanLayer shut down
The host quotes a Horthy tweet that did "a lot of numbers." It predicts a one-to-three-year period in which things break at 3 a.m., loops are expected to fix them, no one understands what's under the hood, and companies face existential risk. Horthy grants that with today's tools you might get away with not reading the code. The problem is that loops eventually produce more code than anyone can read. He places the strong "dark factory" position, and what he describes as Ryan Lopopolo's version of harness engineering — spend as many tokens as possible — in this category.
HumanLayer tried it. The team built a lights-off factory in July 2025 and shut it down by November. Horthy estimates it takes three to six months of shipping unread code before you realize things are getting much worse and starting over is easier than fixing. At one point they hit a bug that Opus 4.1 could not root-cause, no matter how carefully they prompted. The team spent about three weeks re-onboarding into a codebase they had stopped reading three months earlier. After several days of digging, they found that a primary key was being routed through the whole system and needed to become a different type of object. At the time, he concluded it was still worth it: skip reading code most of the time and occasionally pay two weeks of manual repair. He no longer believes that. He thinks the volume of code teams can now generate has grown something like 10x or 100x, so the problem keeps getting worse.
That is why HumanLayer's loops improve code quality rather than ship user-facing features, and why the team reads all the code. Horthy distinguishes system architecture from what he calls program design: where the interfaces and seams are, how dependency injection works, and whatever else keeps a change in one place from breaking another. Software engineering as a discipline, he says, came about in the 1970s to avoid the "giant ball of spaghetti."
The host compares this to how engineers become senior: over years, you learn how small mistakes snowball into disasters, and that tech debt can let competitors overtake you. Horthy allows that "GPT-7" might solve it. But he describes a company that sends every user complaint, crash, PM ticket, and "obnoxious essay" from the CEO in Slack to agents. It replaces human review with agentic testing and agentic review because PR review became the bottleneck. In that setup, none of the remaining components have any feel for architecture, because that intuition has not been trained into them.
Why models don't learn maintainability
Horthy says he always tries to understand one layer below where he works. For software factories, that meant spending recent weeks studying reinforcement learning with verifiable rewards (RLVR) and coding benchmarks. In the typical setup, a model gets a small problem. Its test changes are reverted, a test patch is applied, and the run is scored on whether the tests pass. SWE-bench-style benchmarks take a real commit and issue from Django or another repository and check whether the model reproduces the fix.
In his view, reinforcement learning is what made Claude Code good. The model and harness were trained together, so the model became very good at that harness's tools for reading, writing and searching files. But that is one dimension. The cost of bad architecture cannot be measured by running unit tests, because it shows up three to six months later when the software has become hard to change. No lab has shown a benchmark that tells him whether a model makes a codebase better or worse over time, and benchmarks tend to reflect lab priorities.
The best he has seen is Cognition's Frontier Code. It checks whether tests pass, then uses one judge model to assess whether the patch is functionally equivalent to the reference answer and another to review code quality. He also mentions a "marathon"-style benchmark. He calls these better but insufficient. He is also skeptical of agentic code review. It raises the floor, but models are sycophantic. Ask a model whether code is good and it will praise it. Tell it a coworker wrote the code and it will find problems. So he finds it hard to trust models to judge code quality. His own benchmark idea is to have a model build 20 features in sequence without knowing what comes next, as a real product team would. The benchmark should be hard enough that most frontier models fail by feature six or seven.
Software factories, from 1968 to agents
Horthy says "software factory" first appeared at a NATO conference in 1968. The idea was a system of steps like a factory floor — coding, testing, validation, integration — at a time with no CI/CD and barely any version control. Companies including Toshiba adopted it. DevOps was the next wave: Chef, Ansible and Puppet replaced people resizing disks by hand or clicking through the AWS console. A disk hits 90%, Nagios fires an alert, Chef enlarges the disk. Around 2018, Nicolas Chaillan, then chief software officer of the U.S. Air Force, wrote a roughly 100-page document calling for a "DevSecOps factory" at the Department of Defense. It covered Jenkins, code-quality and security scanning, and CI/CD. The goal was to ship daily instead of every three months or once a year, with automation catching about 90% of issues.
He then describes the pre-AI factory. A work tracker such as Linear or Jira holds items in stages. People plan, pick up tickets, open PRs, get reviews, pass CI and ship. Users complain to support, crashes reach monitoring, and both feed back into the tracker for prioritization. The host notes the long latency in each stage: a bug report can take months to reach a developer and much longer to be fixed. Even at a company like Netflix or Meta, which can deploy many times a day, the work itself still takes hours or days.
In the agentic version, an agent replaces the human builder. It needs orchestration, a sandbox, an LLM, an inner harness and an outer harness. He cites Cursor's background agents as an example, since they add browsers and video recording. Building then takes about ten minutes, so code review becomes the bottleneck, and teams add AI review and agentic testing. Next, support tickets and Sentry or Datadog crashes flow straight to agents that open PRs. He calls this "the Ramp Inspect thing." Finally, faced with too much code to review, some teams turn the lights off. If users don't complain, the system is treated as working.
Horthy agrees with the host that every team is changing its factory at a different pace. His advice is to build one loop at a time and keep each small and contained. "Everything except stop reading the code is really good advice." Turning support tickets into tracker tickets, and possibly into PRs, is fine.
Three options, and seeking leverage
Horthy states his three commitments: cut through hype by trying things and talking to people who use them; protect useful words from "semantic diffusion," another Martin Fowler term, as has happened to "agents"; and understand one layer down. From there he lays out three options for teams going all in on agentic factories:
- Turn the lights off. Let everything flow, and hope you don't create too much slop before better models arrive and you're left with "a giant pile of ash."
- Slow down and read every PR and every line. He says this yields modest gains, and from his work with teams he would expect roughly a 30–50% productivity lift.
- Find leverage points. Identify where an hour of human planning saves four hours of rework in implementation. With the right leverage points, he says, teams can move two to three times faster while staying close to 99% faithful to what careful humans would have written.
He explains that reviewing PRs of hundreds or thousands of lines is costly. It is especially costly when the direction is wrong, because once an agent has committed to a direction, restarting beats steering. His goal is to spend an hour of human-agent discussion up front so the resulting PR takes 20 minutes to read, rather than six hours of back-and-forth.
Research, Plan, Implement — and what it got wrong
Horthy first presented RPI (Research, Plan, Implement) in August 2025. In the research step, agents read large amounts of code with parallel subagents, without being told what task was coming. They produced a markdown document explaining how the systems work and connect. Perhaps 100,000 tokens of reading became a 10K-token document. A fresh context window then produced a plan, and another implemented it.
Looking back, he thinks planning became popular in mid-2025 because it made agents work longer. Ask for "a B2B SaaS for burrito delivery" and you got a homepage. Ask for a plan first, then hand it over, and the agent kept going until the plan was done. But he now calls those plans "actually terrible." They contained every line of change as diff blocks. His team told people to read them, and read their own. Eventually he was only skimming them, so they no longer served to steer. People who reviewed both the plan and the code spent 20 minutes on each, and the two differed. That doubled the reading instead of reducing it — "anti-leverage."
The host compares RPI with spec-driven tools such as Amazon Kiro and GitHub's workflow, which he says looked good but were mostly abandoned. Horthy notes that some people call RPI spec-driven development, since for some the term just means markdown files used while coding. One OpenAI researcher proposed never reading code and treating specs as source that "compiles" into code. Horthy says that never materialized, "maybe with GPT-7." He is on a GitHub issue in Spec Kit that has been open for a year, where people keep complaining that specs and code drift apart. Two sources of truth stop being useful. His team now treats RPI documents as tactical and disposable. They are created for a task, then thrown away, and research is regenerated from scratch next time. Tokens are cheap, his time is expensive, and stale research is risky. He has seen people keep docs and code in sync but has never met anyone who found it worth the effort. He treats the code as the source of truth.
He describes the current workflow as "frequent intentional compaction," a direct application of context engineering meant to keep work in the smart zone. Research is compacted into a document; he says he doesn't read research documents, since models are good at describing how code fits together, though not at opinions or bug-hunting. The ticket and intent become a design document: current state, desired end state, and the model's design questions, which he calls a thorough, perhaps over-engineered plan mode. Planning then happens in a fresh session with both documents. Each step exists because models have a weakness at that stage. Models are weak at designing end-state architecture and program design, so a human reviews the design.
Models also favor what he calls "horizontal" plans: database, then services, then API, then frontend. In an existing codebase, that means 2,000 lines of changes across the system with nothing testable until the end. He would build vertically instead. He starts with a mock API endpoint with fake data, shapes the frontend, wires a services layer, adds the migration, then business logic, then error handling. He would rather read five small, manually verifiable diffs than 2,000 lines that don't work for reasons no one can locate.
Token harder vs. token smarter
Horthy belongs to a group chat called "hyperengineering," where members try to max out their Claude subscriptions. Some run six Claude Code accounts timed so every five-hour window is fully used. That is "token harder." In Goldratt's terms, it optimizes the utilization of one node in the factory instead of the end-to-end goal of shipping stable, lasting value. The dark factory is the same idea. It is named after fully robotic factories that need no lights because no people work there: raw materials in, cars out. Removing humans from review pushes more tokens through the system. He grants that small internal loops can safely be "dark." A review agent can send problems back to a builder agent with no human involved. The full dark factory, where no one reads code, maximizes token use. It suits people who see their job as extracting as much intelligence from the "machine god" as possible.
"Token smarter" means getting as much from AI as possible while keeping control, taste, judgment, understanding of the architecture, and years of hard-won opinions about program design. He compares it to Google's SRE book, where a small team had to manage growth from a handful of data centers to many more without growing headcount linearly. The host adds that Google never removed SREs and the team did grow. Horthy agrees and restates the point: headcount should scale like a square root or logarithm while output scales linearly. Good architecture and program design make that possible, and today that requires humans in the loop. Whatever the aim, he adds, a company with zero engineers is not a place either of them could work.
Slop in, slop out
On AI slop, Horthy repeats his line that AI can write your specs and PRDs as well as your code, but slop in means slop out if you outsource your thinking. He describes three speeds. Working closely with an agent and reading every change is faster than hand-coding, but not by much. Above that is the maximum speed at which you can still care about the code. The fastest is lights-off. HumanLayer targets the middle through a staged expansion. A two-sentence request or a voice note becomes a one-pager, which is checked. That becomes a three-pager, which is checked, then a ten-page outline, and finally around 100 pages of code. The documents need not be perfect. Each check narrows the range of possible outcomes; he calls this his physics habit of "collapsing" probabilities. He also suspects that people who like real-time strategy games will do well with AI. He cites Matt Pocock's discussion of "fog of war": estimating probabilities from partial information and seeking the information that best reveals the likely path to success.
HumanLayer and the idea of killing the pull request
Horthy describes HumanLayer, recently out of stealth, as an AI IDE, a collaboration platform, and building blocks for a software factory. It is aimed at the production end of the spectrum, where failures can cost millions, not at vibe coders building side projects. The goal is to help those engineers solve hard problems in complex codebases two to three times faster while staying near human-level quality.
Its core ideas extend RPI: start at a high level, zoom in layer by layer, and re-steer where it gives leverage. The day before the recording he had posted, "should we kill the pull request?", and he says he can't say much about it yet. His broader argument is that the IDE must be rethought for agents. Many editors started as text fields with an agent tab added on. He says that in Cursor 3 he can't even find the text editor, though he's told it exists. HumanLayer started with an IDE for directing and managing agent work, then made it collaborative using a sync engine and durable streams. The aim is for humans to give feedback on agent work in real time rather than at PR time.
He points out that strong teams have long held design reviews and sprint planning. AI can help with all of that, and using it only to write code misses much of the benefit. People who say those meetings are unnecessary because the dark factory runs everything are, in his view, giving up quality. The platform includes a Google Docs–style commenting layer where agents can show mockups, Mermaid diagrams and HTML. Everything is in the cloud and shared, in the style of Figma, so everyone sees everyone else's sessions. He compares this to Slack's advantage over email: you can see channels light up and join only the conversations you care about. He wants to move beyond discrete units like the PR and even "agile" processes that are still waterfall in practice. The open question he names is the data model for a world of agent traces, documents, tasks, projects and continuously streamed git diffs.
The host compares this to how GitHub made teams' work visible across a company. Horthy agrees that the goal is something like what GitHub did, but more continuous, real-time and collaborative than the PR. The PR, he notes, was a GitHub invention, and probably better than emailing patches, which the Linux kernel still does and which the host says works only for them. Horthy recalls learning Subversion as an undergraduate at the University of Chicago, reportedly because Subversion's creator came from there. The school switched to Git the year after he graduated.
Silicon Valley, hiring, and reading the classics
On location, Horthy points listeners to Paul Graham's talk on why San Francisco matters rather than repeating it. He says he has never felt more "seen" or connected than there. People are competent, care deeply, and share the same problems. Friends co-work in the office until 11 p.m. on side projects, which he believes needs a critical mass not found elsewhere.
HumanLayer hires for strong fundamentals: distributed systems, operating systems, core CS. His reasoning is that he can teach someone to be a good AI developer in a few months but cannot teach a CS degree in three. The problems he finds most interesting are real-time systems, cloud sandboxes and sync. He mentions admiring the ElectricSQL team's durable streams. Coding agents need to run anywhere, briefly or for a long time, on demand or on schedule, while staying part of one shared "brain." Some of the stack is boring, such as data in Postgres. He calls collaboration platforms hard distributed-systems problems, even with newer building blocks. The host compares it to cloud primitives taking about a decade to mature, from AWS to Kubernetes.
For a book, Horthy recommends Martin Fowler's Refactoring. His team spends much of its time improving the design of existing code and trying to get models to write maintainable, readable code. He adds that classics such as Clean Code and The Pragmatic Programmer are "more relevant now than [they have] ever been." That fits the unresolved problem running through the conversation: models are trained to make tests pass, and the question is how to get them to write code that is still easy to change three months later.
So what is context engineering?
It's kind of like deabstracting the abstractions that have been layered on top of RAG, memory, agentic history. At the end of the day, they're all different ways to pass tokens into a model.
What is a smart zone and what is a dumb zone?
The less context window you use, the better outcomes you'll get, always.
A new paradigm that is spreading up is loop engineering. What do you think is bad about it?
Problem with loops is, like, at a certain point, you're going to generate so much code that you can't read it anymore. We built a lights-off software factory in July of 2025 and by November we had shut it down.
Can we talk about what you mean by token harder and token smarter?
I'm in a group chat called hyperengineering and it's all people trying to max out their Claude subs. That's my idea of token harder, and the goal is
What happens when you let AI agents ship code for months and no developer reads a single line? Today's guest tried exactly that. He built a lights-off software factory and four months later he had no choice but to shut it down as things just stopped working. Dex Horthy is the founder of HumanLayer and the person who coined the term context engineering days before Andrej Karpathy and Tobi Lütke made it famous.
He spent the last two years talking to hundreds of AI engineers about what actually works when you build with LLMs and is testing the most extreme ideas with his own team.
In today's conversation we discuss context engineering, what it is and the physics of context windows, including what the dumb zone is; loop engineering, from the Ralph Wiggum technique to the slow loops that Dex's team runs every night to wake up to code cleanup PRs; the rise of software factories, from a NATO conference in 1968 through DevOps to today's agentic factories; spec-driven development and why specs always drift from the code itself; and many more. If you want to understand increasingly important concepts like context engineering and harness engineering, or want to know how far you can push the let-agents-build-everything idea from someone who pushed it further than almost anyone, then this episode is for you.
This episode is presented by Antithesis. If you work with agents, your job is no longer just writing code. It's specifying and testing it, and Antithesis is the most effective method of verifying agentic code today.
Today's episode is brought to you by Buildkite, the CI/CD platform trusted by OpenAI, Anthropic, Cursor, Nvidia, Uber, Canva, and more. Today, we're talking about pushing the right context into models so that they write better code. Right after that starts working, your agents will write more code, a lot more. Trusting that code avalanche is where many teams face a challenge today. Every change that an agent makes still has to be built, tested, and proven safe before it ships. Works on my machine is not enough. So, you obviously need CI. But when agents are pushing five, 10, or 50 times the commit volume into your pipelines, faster CI runners won't save you. Shaving 30 seconds off a single build is meaningless when the queue is 100-plus jobs deep. What you really want is a CI system that gets faster as the volume grows, and CI that offers instant parallelization to give you unlimited concurrency and to intelligently route changes at runtime.
This is what Buildkite does and why global software leaders continue to rely on it. The same architecture that absorbed the scale of Shopify and Uber a decade ago now runs about 1.4 billion job minutes a week across Cursor, Meta, and Snowflake. While the rest of the CI world are cracking under the weight of rearchitecting your platform, Buildkite continues to reliably grow. Agents running on your infrastructure or Buildkite's. Any cloud, any chip, your secrets, your scale. Every artifact and log is captured. So when something fails, either you or your agents have immediate insight for why. As you're ensuring the context you'll give to your agents, think about how you'll verify what they hand back. If your system is buckling under the increased volume, head to buildkite.com/pragmatic. 30-day all-access trial, no credit card, and an actual human engineer on standby. His name's Ola, and he's very helpful.
So, Dex, welcome to the podcast.
Super stoked to be here, dude.
Before we get into some of the context engineering and some of the more spicy stuff as well, how did you get into tech? How did you fall in love with computers?
Oh, man. So, I was doing undergrad as a physics major, and I realized that I didn't like academia. And there's like basically two or three paths out of physics: basically you go get a PhD, or you go into finance, or you go do programming. At that time, this was, you know, 2012, 2011, when it was in the middle of undergrad and deciding what to do. And I had done an internship when I was in high school. I was working with NASA researchers at the Jet Propulsion Lab in California. They had just gotten this really high-fidelity, like the most fine-grained data set of altitudes, like the heights, a very topographical map of the south pole of the moon.
And the south pole of the moon is really interesting because some of the craters there are so deep because of the angle it has. It got hit by meteor storms like no other part of the moon. So there's very deep craters that have never seen sunlight, and so there's frozen liquid water in there from the formation of the moon.
And so scientists were really interested in getting down there and exploring. And so we had this really fine-grained map and it's like, okay, cool. Let's build software so that I have point A to point B. I know the limitations of my rover: max incline up is this, max incline down is that. Find a path from point A to point B that doesn't break those rules of the incline. So I was, you know, 17. I had never cracked a CS textbook. So I basically wrote a really naive, bad version of Dijkstra's algorithm for pathfinding. So I was in college. I was like, I don't know if I want to do the academics thing, but I really enjoyed programming back in the day. And so I decided to go, I got like half of a CS minor and then started working on an API platform team at a software company in Chicago and
Sprout Social, right?
Yes. And basically never went back.
Yeah. And then where did you go from there? Where did you pick up the parts of the trade? Because very early on, your first job, that's not really common. You were doing platform engineering back, you know, more than a decade ago.
From that point, it took me about two or three months to notice that the most valuable work that was being done in the company was being done by, of course it's obvious, like the first couple engineers who know everything and understand where everything was. And you spend a day on a support ticket from a customer and they solve it in 5 minutes, but you have to solve it so you learn and whatever. And I realized the most valuable people in the company were the people that were building the developer platform: CI/CD, sandbox environments, preview stuff. And so that was my first step into the journey, and I've basically been obsessed with software factories since that, like three or six months into my first job.
We talk about software factories now, but you're talking about software factories back then. So you were starting to already think that this is how we can produce better software. This is pre-AI world, right?
Well, and I'm always surprised. There's a huge class of developers that say, I don't want to work on CI/CD. I hate CI/CD. I'm like, really? Because building the thing that builds the thing, and building the thing that builds the thing that builds the thing, is like, as software engineers we're lazy. We want to do the most high-leverage thing that makes our job easier. So if we can build a thing that helps us build a thing that helps us move faster, then that's the best use of my time as a lazy engineer.
And then you went to another startup, Aspiration.
Aspiration. Yeah.
Aspiration, also platform engineering.
Yeah. I was brought in and then three months into the job, the VP of engineering who hired me quit or got fired. I don't know. There was some drama about it. I probably shouldn't talk about it. And then I was there for about a year and was kind of acting CTO for a while, hired a couple people, helped hire the new VP of engineering, but I was out of there. I don't think I'll ever do consumer again. I think I'm actually a B2B guy.
Good to know. And then you went to Replicated, where you spent a good, solid four years and went from engineer to forward deployed engineer to product manager.
Yeah, I did core engineering for like two years. We were building a container orchestrator before Kubernetes, before Docker Swarm was really a thing. We built our own orchestrator. The founders had this vision that, oh, Docker is going to make it much easier to ship on-prem software. And when I say on-prem, I don't mean literally a rack in a colo. It's more like, hey, look, bring the app to where the data is rather than sending the data up to some cloud vendor. And Docker makes it much easier to package up apps and move them around. And so they had this thesis that basically you could build a platform, the experience that you get when you use GitHub Enterprise, which is like you install it and it has this admin panel, but then you just get GitHub running in your data center and your code never has to leave your data center. Suddenly you could build a generic SaaS where everybody could have that.
So I did two years as an engineer there and then our head of sales, we parted ways with our head of sales. And honestly I was having a lot of arguments about the software factory with our CTO, and it's kind of almost like a too-many-cooks-in-the-kitchen kind of thing. I'm sure many listeners have had this experience of like, well, yeah, I know I have these tickets to build, but CI sucks. I got to fix CI because it's too slow, or there's too many different builds and it's always breaking. I'm going to fix that and then I'm going to do the tickets. And it's just like, Dex, I need you to stop fixing the build pipeline and do the tickets I gave you. I'm sure you've had this experience, perhaps.
Yeah. And was this what led you to forward deployed engineering?
Yeah. So I really loved our customers. Our customers are HashiCorp, DataStax, Puppet, all these really cool engineering brands. TravisCI, CircleCI. I was like, yeah, I actually love working with our customers. Our customers are awesome. And it was a great way to get in the trenches. A lot of really good engineers who were solving the hardest problem at the company, which is: how do we take this 3-to-5-year-old SaaS platform and package it all up so that someone who knows nothing about our architecture can run it reliably in their own AWS VPC, in their own on-prem data center, whatever it was. And so I was our first kind of customer-facing engineer, and in about three months, I met with every customer that was kind of in the pipeline but wasn't moving sales-wise, and we closed like 12 deals in 3 months. And the CEO was like, "Holy crap, Dex. The investors are taking my calls again. I know you want to get back to coding, but I need you to go hire three people and build this team out, because I think you might have been born for this."
Wow.
Yeah. So I did that for about four years, built that org to like 25 people, and then Zer happened and it got a lot smaller. And we kind of realized, hey, we have a product that's pretty good, and we've been solving what lots of early startups do, which is like, okay, there's some usability issues. We'll get a bunch of smart people, throw them in the trenches with our customers. Great for sales, great for retention, all this stuff. And it was like, oh, actually the margins on that aren't good enough. And so we basically were like, cool, we actually just need to make the product way more usable, do a more PLG-shaped thing, make it product-led growth. Make it a little more self-service so you don't need an expert to teach you how to use it. And I was like, cool. If that's the most important thing, then I want to go be a product manager, because I have tons of opinions. I've now spent four years in the trenches with our customers. I have a laundry list of roadmap things that I think would make the product way easier to use and adopt and implement and deploy.
And now you went the full, you went towards the dark side.
Exactly. Yeah, I did. I was like, this is going to kill my street cred, isn't it? But I was really glad. You know, I think a lot of engineers are afraid that if they go do a customer-facing thing they lose all their credibility. And yes, I wasn't coding for 10 hours a day. I was coding for like three or four hours on a Saturday for fun. But I mean, we were helping people build YAML, we were building CLIs. We owned a lot of the tooling that customers use, but it was the last-mile delivery side of it, not the core platform.
And on a more personal note, I had spent most of my 20s feeling like, okay, a little bit introverted, a little bit socially awkward, what a lot of engineers I'm sure experience. And I had talked to my uncle, who's a music producer. So he used to work with Randy Newman and a bunch of really famous musicians.
Oh wow.
Yeah. This guy Mitchell F. And I was sitting at dinner with him at some point, I think it was when I was still in undergrad, but he gave me this lecture. He was basically like, if you want to be really good at something, you have to make it the only thing you do. The guy playing guitar nights and weekends trying to get his band off the ground will probably never achieve greatness. The people who become great are the people who basically make it like, if I don't play guitar, I don't eat. And you go and you sit on the street all day and you play for 14 hours a day or whatever it is. That's the only way to become great. So, I said, okay, instead of trying to read self-help books about how to be less introverted and less socially awkward, what if I just made it my freaking job to just talk to people and make friends and help people and solve their problems? And I think it worked out. I recommend it. I think everyone should spend a year or two at least doing something really customer-facing.
Did you do this because you felt that it was holding you back, being introverted? I know you got the motivation from the whole musician thing. I get it on one part, but what was it that made you say, a customer-facing thing is what I'm going to be doing? Because clearly you were pretty great at writing code by that point. You could argue you were doing it night and day. So where did you find that, I actually think customer-facing, or getting this introvert off of me? Did you feel it was holding you back, or you just wanted to be good at it?
It was just kind of a thing that was interfering with my general life satisfaction. And it was also like, I'm not a very type A person. I'm very disorganized. I don't know if people call it, like, okay, I'm ADHD, now that's why I can run 30 Claudes in parallel or whatever it is. But I was really bad at email and calendars and spreadsheets. I just didn't care about these, didn't understand them. And so another side effect of this was it just forced me to be organized and keep a lot of things going. And so, I don't know, there's weird benefits you get from stepping outside your comfort zone and learning industrial disciplines
that are separate from what you've been doing, and so the opportunity presented itself and I was like, oh, I like working, I'll try this for a little bit. Started going really well. I'm like, cool, let's see how far this thread goes.
And then afterwards, you're now in your second startup. You became a founder and you also got involved in AI pretty early, even before it was so obvious that it would change how we develop software, right?
Well, I would say I was later than I could have been, because we started the company, me and a buddy in Chicago started a company in the data engineering space in about 2020, November 20. We decided in like August of 2015,
This is Metalytics.
Metalytics. Technically still the same company as HumanLayer, we just pivoted the mission. But yeah, the advice I got from every angel investor, you know, people who just knew CTOs I'd worked for before and stuff, they were just like, look, hitting a lot of heads wins. I don't know if you know the whole dbt data engineering fervor, that whole arc where it was this huge party and tons of investor money going into all these different companies, and then by like 2021, 2022 there was kind of the ZIRP thing and just this general realization that the TAM for those sorts of tools is not as big as everyone
Yes, the total addressable market for those sort of tools was not quite as big as we all thought it was. So it was a hard place to raise money. It was a hard place to get customers.
Yeah. And then I met you while you were at HumanLayer in SF at an event. You actually talked and we chatted afterwards. But by that, this was about a year ago, you started to have some really strong opinions on using AI. And one of them was this now famous 12-factor agents manifesto.
Are we calling it a manifesto now?
I'm calling it a manifesto. It's a manifesto. I'm calling it. Let's talk about this. This was 12 engineering principles to build reliable, production-ready apps. How did you come up with this, and maybe we can also talk about some of them.
Yeah. So I'll kind of go to around August. The co-founder I was working with kind of burned out and left, and we were on good terms. It was very mutual. And I decided to start messing with AI stuff, and I was building AI agents, and what was really in vogue right then was like the LangChain, the CrewAI, these agent frameworks. And it seemed like there was a ton of, you go in the CrewAI Discord, there's 10,000 people. It's like, okay, this feels like the right shape, and there's clearly this eco... You go in every single one of those projects, they have a Chroma DB plugin. They have like a Composio plugin. There's clearly, this is the shared interface that everybody is building for. I say, okay, what's missing from all of this?
The agents can call tools, but it's really hard to control which tools they call. And if it's a chatbot, obviously you can show approve/deny in the UI of your application. But I kind of was obsessed with what I would call outer loop agents, or proactive agents. Agents that would run in the background, get triggered by events. I mean, OpenClaw is basically the biggest manifestation of this, of like you have a heartbeat, it wakes up, it sees if there's any work to do, it tries to do stuff. And my thought was like, I'm not going to trust that agent to do anything meaningful if I can't get a Slack message or an iMessage or something when it wants to do something, and kind of guarantee deterministically that I can approve or deny that, or deny it with feedback and say, actually, no, do it like this.
So we played in that space for a while and talked to a lot of founders and founding engineers and builders. We came into YC in the fall of 2024 with this idea. We're building out this API platform, and it was sort of like PagerDuty, but it wasn't who's on call to fix the servers. It was like this routing mechanism for who needs to approve this agent, and can they escalate it or delegate it or defer it, all this stuff. And we built it for this ecosystem: CrewAI, LangChain, there's so many, Griptape, there was so many in that time. And then I talked to tons of AI engineers who were actually building really interesting things and actually making money, doing six-figure contracts shipping AI to the enterprise, and all of them had tried that stuff for like a month or two and then they had thrown it out, and they were just writing all API calls by hand, and they were building things that look more like pipelines and workflows than these sort of hands-off call-tools-in-a-loop kind of thing.
And so I talked to a hundred people, and I spent a lot of time hanging out with one of my best friends, Vib from Boundary. So they built like this protobufs-for-AI thing, and I think they're about to launch their full-fat, Turing-complete programming language thing. But he had this way of thinking about agents and building with models and building with inference where it was a lot more about understanding what structured output really is under the hood. And every single step in your AI workflow is just tokens in, tokens out. And your job as an engineer is to figure out, okay, what tokens do I need to put in to maximize the chance that the tokens out are going to be good. And kind of distilled all these ideas into about 12 principles and wrote about it on GitHub, posted just like this 12-page GitHub repo, threw it on Hacker News, it was on the front page for like two days, and I think it really resonated with a lot of people.
Yeah. So I'll just quickly read the 12 principles and then let's talk about one or two that resonate. The 12 are: natural language to tool calls, own your prompts, own your context window, tools are just structured outputs, unify execution state and business state, launch/pause/resume with simple APIs, contact humans with tool calls, own your control flow, compact errors into context window, small focused agents, trigger from anywhere, meet users where they are, make your agent a stateless reducer.
The stateless... yeah, the stateless reducer one, actually someone hit me up on Twitter and corrected me. It's actually a transducer, because there's technically multiple steps in the workflow, but there we go.
But of this one, this was a year ago, which is like forever in how the tooling is evolving. Which ones still stick with you, where you're like, "All right, these were good, these still seem to hold up"?
I think I spent most of March writing it, published this in April. And then swyx hit me up from AI.engineer and he said, "Hey, do you want to come talk about this?" So I gave this talk, 12-factor agents, on like June 6th, I think. And small room, it was packed, but it was maybe a hundred people. That was the year at AI Engineer where physically, on the second basement floor, was all the super corporate stuff, and you go up a level, it's a little bit more, and then on the top floor is all the weird, cutting-edge startup stuff that you probably shouldn't care about yet, kind of thing. So we were up there on the top with this weird way of thinking about agents.
And then about a week later or two weeks later, Tobi Lütke from Shopify says, I really like this idea of context engineering. And I'm like, I wrote about this two months ago. This is great. Tobi gets it. And then a week later Andrej Karpathy is like, well, I think what we should think about is not prompt engineering but context engineering. And I was like, "Yes, that's my..." Anyways, I don't know. If you ask Gemini, depending what day it is, they will tell you either me or Tobi or Andrej came up with context engineering. You can't really own a word. No one remembers who invented the word prompt engineering.
But of all the factors, factor three, own your context window. And basically, whether it's agentic or a single step in a pipeline, the only way you can impact the quality of your output from AI is by caring a lot about what the inputs are and crafting them.
Let's talk about context engineering, which I am going to credit you that you coined. I did some research and I think you were earlier by a few days. So there we go, you coined it. We're adding to the SEO juice. We'll have it in a transcript: Dex coined context engineering.
Well, an asterisk on that is basically, I learned about context engineering from talking to these hundred engineers and founders. I just kind of looked at what was the same about what they were all doing, and I put a name on it. So I didn't invent doing it. I was just like, I think there's this thing, and vocabulary and names are really important, and having clean ways to talk about the problem, especially when a lot of the content about AI right now is so much hype and jargon that is meaningless. I was like, okay, I think there's a word here that is useful to builders that explains how they should be thinking about building their software.
So what is context engineering?
It's kind of de-abstracting a lot of the abstractions that have been layered on top. So you have RAG, you have memory, you have agentic history, you have structured output, you have all these things that are different ideas in the frame of agentic programming. And at the end of the day, they're all different ways to pass tokens into a model and ask it to produce, usually, some structured output. And understanding that is a lot more powerful than trying to learn memory and trying to pick some agent framework off the shelf and some memory framework off the shelf. I mean, these things are all really good if you want to get to like 80%, you want to get a really good demo. But when you have to go from 80% to 95% or 99%, you need to go down a level and think about what's everything we're putting into the context window. What order is it going in, depending on which model we're using? And all of this stuff matters. You have all of these levers that you can pull. And it just felt like the right abstraction for thinking about how do I get AI to do the thing I want as accurately as possible.
Why has context engineering started to become more talked about? It was about a year ago. Did it have to do with the context window that we could pass on to LLMs? Did it start to expand, or did we just start to realize that we can do a lot more by passing on... the easiest one is of course system prompts, but of course whenever you build an LLM behind the scenes you will pass additional context as well, not just the prompt of the user, you will add a bunch of stuff, that's I guess a dirty secret of any LLM. But why do you think the focus is moving on to, all right, context is important?
I think it always was important. I think what had to happen is a ton of smart people, again, all these builders I talked to, had to focus really hard on producing, like, I want to make software that I can sell. I want to make something that is accurate enough that I'm proud of and I can sell to an enterprise and they're going to be happy with it. And the easiest way to get to really high-quality AI applications is by thinking at that token level. Thinking about a string of different LLM calls, rather than just tools in a loop, where it's kind of open-ended and very flexible but not that reliable. Thinking of agents as workflows, as pipelines, as some mix between maybe a couple tools in a loop, versus just, hey, I have my tools and I have my model and I have my system prompt and these are the only levers I have. And it's actually, no, you have way more levers. It's going to take more work and you're going to have to understand the LLM with a deeper intuition. But it was a thing that we always needed, and it just took time for people to build with this technology to figure out that this is the layer of abstraction that allows you to break through the quality ceiling.
And how are cost and context engineering connected?
Yeah. I was talking about this with someone this morning, about when you're working with LLMs, one of the things I like to say is kind of like make it run, make it right, make it fast. See if the world's best LLM at the time, I think we did a podcast episode where at the time it was like o3, see if o3 can solve your problem, and then give it to people and see if they want that. And then if people want it and you use it a lot, then go do a bunch of context engineering, because your engineering time is always the bottleneck. Humans trying to figure out and solve problems and build evals and improve and try different dimensions or set up jeep or whatever it is, is always going to be more expensive than just using a smarter model
until you have millions of requests a day, and then it's like, okay, we're going to do a bunch of context engineering, break this up into three calls, and get it to work on GPT-4o. And then we're going to take two of those and make those two work on GPT-4o, using old model names. But the point is, for a certain task in your workflow, can you get GPT OSS 12B, which is like 1/1,000th of the cost of Opus, can you get it to solve parts of the problem, so that the things you're using the smartest frontier models for are just the things that really need that level of intelligence? But you shouldn't go build all of that and overengineer it until you've proved that you need it, that it's valuable, that it's like, okay, this is now...
I mean, we get to Eli Goldratt, he had this book, The Goal, right? It was about how to model your factory, and I'm sure we'll get to that when we talk about software factories. It was like, what is the bottleneck in your system? And one day it will be latency and cost, but it's probably not that when you first start out. And context engineering is how you add human effort to the equation to improve the efficiency, the speed, the price, the cost efficiency of your system.
Interesting. And then one thing that came up more recently is harness engineering. What is harness engineering?
So I made a post in like October, or maybe November, of like, hey, there's this new thing that I see, I'm calling it harness engineering. My definition that I had at the time is not... actually, this guy Viv, who's at LangChain now, does a lot of really good writing on agents and how to think about harnesses. He had written something called harness engineering a couple weeks before me, but I hadn't read it at that point.
And my take was basically, okay, when you build an agent, you use context engineering. When you use an agent, because we gave this talk in August of 2025 about how to apply context engineering to how you use coding agents, and that kind of evolved into this idea of how do you take a harness like Claude Code, like Codex, how do you engineer against the integration points of that harness? So commands, MCPs, skills, how you organize your codebase. How do you optimize the environment that the coding agent runs in to get the best results? The same way with context engineering, how do you optimize the inputs to every single prompt? Well, harness engineering just is, how do I raise the floor so that every single turn of this thing, the results are as good as possible. And the term got super blurry, and some people think harness engineering means building a harness, and some people think harness engineering means building around a harness. I actually like what Martin Fowler came up with, as usual, he's very good at naming things
and he kind of defined the, you have the LLM and then you have the inner harness, which is like the thing, the tool definitions and the integration points that, like, say, a Claude Code or a Codex or an Amp actually exposes. That's your inner harness. And then you have the outer harness, which is the stuff that you, the human, do to customize that for your specific needs, your codebase, your languages, etc. That's the best definition I think we have for harness engineering.
It's interesting how naming is still so, so important, isn't it?
Well, it's like as soon as you name anything, people are, most people are... I'm actually surprised that context engineering still means the same thing to most people that it did a year ago and that it's even still relevant. Like, that's honestly the craziest thing to me is, like, how many things that were written about AI 15 months ago still matter or are still interesting or still have good advice baked into them. Stuff changes. I think context engineering has been so long-lived because it's grounded in the fundamentals of how transformer attention works, and until we have post-transformer models or linear attention or whatever it is, which, who knows when that's going to happen, context engineering will be interesting and important to anyone building on AI.
And can we talk about the physics of context? You had a tweet, this one, the context reality check. This is a graph of, as you get to 1 million context, the quality just drops. It goes down. What do we need to know about the context? Again, we now have models that do have a 1 million context window. Maybe we'll have even longer ones, but when you start to just put more stuff into the context, it starts to become less efficient. What do we know so far from a practical perspective, as someone who is using the context window to add on a bunch of stuff? May that be MCP, may that be tools, may that be skills, may that be all of these things?
Yeah. I mean, so the longer context windows are good. You can talk to it for longer. They're doing a good job. But at the end of the day, especially when you had, like, Opus, it was like Opus 4.5 and then Opus 4.5 1 mil, or 4.6 and 4.6 1 mil, you're not actually getting a smarter model. The intelligence of the model is what drives its ability to attend to all of the tokens in the context window, to figure out on the next turn which parts of this 100k or 200k context window are the most relevant to making the decision of, like, what is the next tool we call, and doing that over and over again in a loop.
So I don't know, there was some study that came out in 2025 which found that, and again, these are old models, so inflate your numbers, but it was like frontier LLMs can follow about 150 to 250 instructions before it starts to drop off. Their ability to follow all the instructions just drops off pretty quickly. And I think Lori Vos, I haven't actually looked at the data, but they did a study with the next generation models a year later, and it looks like it's much better, the number of instructions you can get in.
In any case, I split context engineering into two categories. Most people think about the information budget of, like, okay, I can do RAG and I can pull out chunks of this document rather than putting the entire book into my context window. I can just go grab the pages that matter. But it's also your instruction budget. If you give the model too many instructions, and especially too many conflicting instructions, and that's in your initial prompt, and also if you have a conversation, you start going down a path and then you change your mind and you start going down a different one, "actually, I don't want to do any of that, I want to do this," it's a lot of computation the model has to do to notice that it has to ignore that whole thing. And when both of those things are far back enough in the context window that they're only half getting attended to, your likelihood that it's actually going to remember the exact instructions you gave it 100,000 tokens ago goes down quite significantly.
This is all very interesting because as engineers, we are expected, when we're AI engineers, which now is a lot of software engineers, meaning you just use LLMs to build software, like underneath there's an LLM layer somewhere, you're an AI engineer, congratulations. But it sounds like the expectation is, to be a good software engineer pre-AI, you need to understand how to write good code, and it helps when you understand a little bit of the underlying. We didn't need to do that that much over time, but it never hurts. But it sounds like right now we're in this phase that to be an engineer who can write an efficient AI system that uses LLMs, you need to understand the dynamics of the context. You need to understand why stuffing your context one way or the other can introduce latency and all of these. It sounds like it's kind of more of an intuition, and of course there's some understanding, but from talking to you, you're like, well, it does this computation, like, I know, because you tried it out, right?
Yeah, I'm not a PhD in machine learning. I couldn't actually go draw a mathematical proof of how this works, but we know attention is quadratic, and the more stuff you put in, the more it has to spread this attention out over everything.
This just feels like an absolute new area and a little bit very different to what we're used to in software engineering, which is pretty kind of black and white, right? It compiles or doesn't compile.
That's true. I mean, there's a different kind of intuition. I was talking about this earlier as well: there's a different kind of intuition that you develop over years as a software engineer, and there's many categories of it, but the one I'll call attention to is a thing that you cannot teach, you cannot learn in a textbook. The only way to learn it is, like, I know bad patterns in software because I have debugged them at three in the morning. My buddy Jake from Netflix said this in his talk at AI Engineer Code. There's no better way to learn what is good and what is bad and what works and what doesn't than suffering through the thing that doesn't work.
Well, speaking of suffering through the things that don't work, a new paradigm that is spreading is loops. Loop engineering. The idea that instead of writing prompts, just write loops. Set up your loops. And this all started with the Ralph Wiggum technique, where it will just, well, I guess that's an early version of loops that just looped around, and now we're hearing some of the biggest labs talking about how they're actually just doing looping. What is your take? Have you done some looping yourself? Have you set up some loops? And what do you think is good about it and what do you think is bad about it?
Yeah. So I think of loops as, I mean, I could ramble on this for 10 minutes. This is an entire talk, but I'll try to lay out some high-level stuff and then we can dig in wherever you think is most interesting.
We had Ralph Wiggum. It was actually a year and four days ago was the first time I saw the Ralph Wiggum demo, and Jeff Hunley was just visiting SF and he just came through and dropped everybody's jaws with his, like, "Yeah, I just ran Sonnet around the clock and spent six grand in six weeks and I built an entire Gen Z programming language. Look, it compiles, and it has a stage two compiler where the compiler for the language is written in the language itself," and all the insane stuff.
And the core lesson from all of that, I think, was the idea of back pressure, which is basically, and I think a lot of people have been doing this for a long time, how do I let the model check its own work? How do I automate the process of getting feedback into the model? And there's lots and lots of different flavors of this. You can have deterministic linters. You can have unit tests. Part of what made the programming language easy to build with Ralph is a programming language can be infinitely verified. You write the code in the language, you compile it. If the compiler fails, you go fix the compiler. You run the program. If the program fails, you go fix the compiler. It's very, very verifiable. And I think the lesson in loops engineering is, if you can make a problem very verifiable, you can kind of treat it like a black box
and then have it loop, because it will keep improving itself because the verification loop is already there.
Exactly. And so you can do this with CI/CD. I do this every time I'm doing a release. I'm like, I'm tired, the CI/CD is slow. Cool. Go research the codebase, make a change, make a pull request, run the test, see if it's faster, try again. Run the test, push to the branch, check again, see if it's faster. And so if it can verify its own work in a loop, instead of saying let's try this approach or let's try that approach, or suggesting and being really back and forth, you just say, my goal is to make CI faster. And you tell the model, here's the five steps: you're going to write some code, you're going to commit it, you're going to push it, you're going to launch a sub-agent to watch the job until it's finished, it's going to tell you what happened, then you're going to decide what to do next. And so that's the very simplest example I have of designing loops.
And you just set the goal, which Claude Code and, I think, Codex both ship: /goal, which is you just set the goal and it iterates until it reaches it, or as long as it makes progress towards it.
Exactly. And so if it's verifiable, if you can measure. This is auto research too. Auto research is like, hey, go make this model twice as fast, and it's just a prompt that tells the model to go at it over and over again and try things until it actually has good results. So that's what I think of loops engineering.
I don't know, we do a very interesting kind of loops engineering where the challenge is, I think it's very easy to get very excited about building the thing that builds the thing, or building the thing that builds the thing that builds the thing we talked about. And so people say, "Oh, we need to redo everything as this big agentic-first factory, maybe even a dark factory." And they're redesigning their entire thing to be their infrastructure for the next 5 years. And one thing we know of in engineering, and especially pragmatic engineering, is how can you make this more incremental? How can you make it more continuous? And a lot of people don't have the option to just, "Hey, I ran a Ralph loop for 3 days and it fixed every lint error in our codebase. Here's a 60,000-line PR. Who wants to review it, and who wants to sign off on merging and deploying it and that there's not going to be any bugs?" Nobody.
So I think the thing I'm most excited about is actually what we call iterated loops or slow loops, where we basically have a cron job. The structure of the loop is really easy. It's like: run this linter, fix one thing, commit and push. And then we run that every night in our GitHub Actions, and we wake up every morning to one PR that makes the codebase a little bit better.
I like the slow loops.
Yeah. And it has two dimensions. Now we have a blueprint for it, and actually Kyle just shipped a skill so that you can build these yourself. You can add more feedback mechanisms. So we have React Doctor for the front end. We have another anti-pattern that has no deterministic tooling, but Kyle's just like, "Here's what good looks like. Here's what bad looks like. Go fix one thing and bring it back." It's like prop narrowing, basically. We have a bunch of optional props and most of them don't need to be optional. It's like, here's how to make the prop not optional, so that the code is cleaner and easier to reason about. And so you can add more conditions, more things of like, fix one thing, I want to wake up to a PR. So now we wake up to four PRs because there's four separate things. And then the other dimension you can do here is, as you gain confidence, you can increase the scope. Instead of fixing one thing, fix four things.
And so these are other ways to think about loops, where something that's not a human triggers it to start. Whether it's an alert from Sentry, whether it's user feedback like a support ticket, whether it's a PM writing a ticket, whether it's a test failing, or it's a cron that runs on a schedule. The trigger should be something that you don't have to press a button on, and there's a defined workflow, and it makes everything a little bit better.
Dex just described letting agents fix things without a human pressing a button. But what if a bug is too difficult not just for an agent but also for a human to reproduce, let alone fix? This is where presenting sponsor Antithesis comes in. I was recently pairing with the Antithesis team, where we did a walkthrough of how they helped fix a nasty bug in etcd, the open-source key-value store used by Kubernetes. This is a bug that actually happened in etcd. The team noticed that the linearization validation assertion failed during the regular Antithesis runs. This is not good, because linearization guarantees strong consistency. So this needs to be fixed.
So what the etcd team did was run a causality analysis inside Antithesis. This generates this graph, which is a bug probability graph. Here the x-axis is virtual time and the y-axis is probability. Now we see that something happened just before virtual time 24 that caused a huge jump in the probability that the bug would occur. Going deeper, we can look at the entire set of timelines. Vertical lines going down represent events branching off from the same state, and the purple dots are where the bug happens. If we look closely enough, we see that all of the failures come from one parent branch.
Gotcha. This is such a useful debugging tool. In the end, the team was able to figure out that process pauses were causing the bug using all these Antithesis debugging tools. This non-deterministic bug was diagnosed in a deterministic way. How cool is that? Oh, and this is an actual bug that then got fixed in etcd. You can see the bug and the fix in etcd's GitHub repo. Honestly, the tools that Antithesis has built for debugging feel pretty darn futuristic, but they are also really powerful. Head over to antithesis.com/pragmatic to learn more.
I'd also like to talk about our season sponsor, Sentry. Sentry is a tool I use for application monitoring on all of my projects, including the Pragmatic Engineer back end. I've used it for 10 years now, starting with when I worked at Uber.
A neat Sentry feature I'm liking is their Seer AI agent, which helps investigate production errors. For example, here's an actual error I had in my application. I can just ask Seer what might be the root cause, and it brings context. And it can also make a plan to fix it right from the web interface. And a nice thing is how Seer also works great from Slack as well, not just from the web.
One place I find even more handy to use Sentry is from Codex or Claude Code using Sentry MCP. Also, you can set up neat automations, like when a resolved Sentry issue resurfaces, you can kick off a Cursor agent or GitHub Copilot agent to investigate the regression, read the relevant code, and open a PR with a suggested fix. I'm not a fan of using AI tools just for the sake of it, but I really like the practical integrations where I can fix errors faster and with more context. Check out Sentry at sentry.io/pragmatic and start monitoring and fixing regressions today. And with this, let's
get back to Dex and to agentic loops that trigger themselves. Now, you said we can get more ambitious and we can add more things to it, but I'm going to quote you with one of your tweets which says, "This may surprise you that this is coming from me, but I think we're in for a 1 to 3 year period where stuff might break at 3:00 a.m. and you're relying on loops to fix it and nobody understands what's under the hood, and you're looking at an existential threat to your company."
Yes. Yeah, that one was great. That one did a lot of numbers. [laughter]
It resonated.
Here's the other side of it: I think that today, with today's models, today's programming languages, today's infrastructure, you might get away with not reading the code. Problem with loops is at a certain point you're going to generate so much code that you can't read it anymore. This is the StrongDM dark factory. This is like Ryan Lopopolo's harness engineering. Just spend as many tokens as possible. We tried this. We built a lights-off software factory in July of 2025 and by November we had shut it down. I think it takes about three to six months of you shipping all the time with nobody reading the code before you realize, wow, this is getting way worse and it's easier to start over than it is to fix it. The models have made the codebase so bad that it is actually going to be easier to just rethink this from scratch. And maybe that's okay because we have AI and it's easier to rebuild things from nothing.
And usually when engineers say, "Oh, we can't fix this. We have to rebuild it," the feedback is, "No, just refactor in place. Just constantly keep the codebase getting better." You mentioned what I said. You'll notice what I said was not use loops to ship the features that users want. We use loops to actually improve the codebase quality and we read all the code because we care about how it's architected, and we care not just about the system architecture but what I would call the program design, which I think is something people are going to... where are the interfaces, where are the seams, how are we doing dependency injection, all of these things that make your codebase more maintainable over time and keep you from falling into this trap of, okay, well now if I change something over here I broke something over here.
This is the classic problem of software engineering: software engineering was invented in the 1970s because we realized we needed techniques for avoiding that problem of this giant ball of spaghetti. And I don't think the models are smart enough, and I don't think we actually have the training and the benchmarking and the eval techniques to get models to write code that is more maintainable over time, versus they're all trained on SWE-bench-looking things, right? All of the benchmarks are basically: here's a commit in Django, here's an issue that was filed around that time, see if you can create the fix that the human created. And it's Django and it's Apache and there's a hundred repos in Go and C++ and TypeScript and Java and all these different languages. The problem with training models on maintainability is the cost function of bad architecture and bad program design can't be evaluated by running the unit tests, because it hits you 3 to 6 months later when you're like, "Holy crap, this software has become so hard to change."
Is this not similar to senior software engineers, why it took years for someone to become a senior? Because typically, and in some environments you can become a senior faster, typically fast-moving ones where there's a bunch of issues and you have to keep fixing it. Sometimes some people are working in the same place for 10 years and they're still not at that level. The point was it just takes time for you to understand the small mistake that you make right now that snowballs into something disastrous later, and you get hit by it and you realize, okay, testing matters, architecture matters, tech debt can actually be a killer. We don't talk about it anymore, but we used to talk about how tech debt kills or slows down companies so badly pre-AI that their competitors can overtake them, or they're just stuck with a 2-year refactor not shipping any new features, and the competition ships a bunch of other stuff and now they're ahead.
And I will say it is possible that GPT-7 will fix this, but if you are turning the lights off in your software factory and you're saying, "Hey, you know what? We're not going to read the code. It's fine. The models are smart enough. If we give it the right feedback and just throw enough tokens at the problem, it will keep getting better." This is what led to this tweet. That might work, but if nobody read the code in three months and you replace all of your code review with loops of, hey, if a user complains, we give it to an agent. If something crashes, we give it to an agent. If a PM writes a ticket, we give it to an agent. If a CEO writes an obnoxious essay about what we should be building in Slack, we give it to an agent.
Yeah. [laughter]
And then you stop reading the code because that's going to produce way too much. No one can read it. And the PR reviews become the bottleneck. So you replace that with agentic testing and agentic code review. But none of these things have intuition for software architecture because we haven't trained it in yet. And so you're going to wake up one day and you're going to have an issue. This happened to us, and we got through it, and at the time it was still worth it. We spent 3 weeks onboarding back into the codebase that we had stopped reading 3 months ago because no matter how much sophisticated expert prompting, we could not get Opus, I think it was Opus 4.1 at the time, we could not get Opus 4.1 to actually find the root cause. We had to go spend several days digging through the code and figuring out, oh, there's actually a primary key that's being routed through this whole thing that needs to be changed to a different type of object and it needs it.
This actually happened to you.
This happened to us. Yeah.
And when it happened, I was like, you know what? That sucked. That was terrible. But we did it. We solved it. And it's still worth not reading the code most of the time, at the cost of every once in a while I'm going to have to spend two weeks fixing an issue by hand. And I don't believe that anymore, because I think the amount of code we're able to write now is actually 10xed or 100xed and I think the problem's just getting worse.
So let's talk about software factories.
Yeah.
In your mind, because I feel it's an overloaded word, what do you think of a software factory before AI and now post-AI?
Do you know the first definition of software factory, the first time it was used?
No.
It was a NATO conference in 1968.
Oh, Grady Booch would know about this.
Yeah, exactly. You should ask Grady about it. They talked about the idea of, okay, you actually need to build a system of steps, just like a factory floor. You have the coding part and the testing part and the validation part and the integration part. We had no CI/CD, we barely had version control, but you needed a factory. And then it was adopted by Toshiba and a bunch of companies. And then the next moment was DevOps, and you have this idea of, okay, we're going to do CI/CD, we're going to automate, Chef and Ansible, Puppet, whatever, all these technologies. Instead of having dudes running around data centers resizing disks and stuff, or clicking around the AWS console.
Yeah, exactly. It was like, cool, we build loops. The server hits 90% disk space, that sends an alert to Nagios, Nagios triggers a Chef run, Chef makes the disk bigger. Feedback loops, right? This has been around for a while. And in 2018, I want to say, this guy Nicolas Chaillan, who was the chief software officer of the Air Force, wrote this 100-page essay of, hey, the DoD needs a software factory.
The Department of Defense.
Yeah, the Department of Defense and the Air Force. And he called it the DevSecOps factory. And he said we need all the things that all of the good enterprises are using. We need Jenkins. We need code quality scanning. We need security scanning. We need CI/CD. We need to be able to ship. We're shipping once every three months or once a year. We need to be able to ship every day like all these other companies. And the way we do that is we actually embrace all these automations and technologies so that 90% of the issues are caught by automations instead of people manually checking it or manually reading the code or manually integrating modules together.
Wow. Talk about forward thinking in the government.
I know. I was surprised, like, oh nice. And part of it is, hey look, we're falling behind. I don't know exactly all the reasons, but I imagine it's also about attracting really good talent: hey look, if we have the modern software stack and we're building things fast and we care about efficiency and we care about using people's time well, we care about them spending time on the hard parts of the job, not manually looking for SQL injections. You could automate that.
So this was software factories pre-AI.
Pre-AI.
Now I've heard the term a lot more because of AI.
Yeah.
Is it the same? Is it different?
So this is really hard to say without a drawing, but I'll try to draw it out. At the core of a software factory, you have a source of work. You can imagine a Linear, a Jira. The source of truth, your object, whether it's a spreadsheet or whatever, is what stages is the work in.
Yep. And pre-AI you would maybe do some architecture review planning. You would maybe do some sprint planning, and then people would take tickets off the queue and they would go build them. And then you would make a pull request and people would review it and you would run CI checks, and then you would send it to prod, and then it would make contact with your users, and your users would complain about stuff and that would go to your support team and back into your work tracker. And it would crash and you would have issues, and that would go into your monitoring stack and that would go into your tracker, and that was your loop. And then people would take stuff off the tracker based on priorities: product managers, engineering managers, engineers prioritizing work, and then we go and do that.
And the first change is this long-winded, lots of phases. And this is also why when a developer ships a bug, by the time it comes back to you it might be two or 3 months or even longer, and by the time it gets fixed it might be a year or two. And this is why when you're using a piece of software, it's like that annoying bug and you talk with customer support, but it's just very long latencies at each part of the factory, if you will.
Yeah. And the step where someone pulls a work item off a queue and starts working on it is a couple hours to a couple days before it actually gets integrated into everything else and touches users. And that's in a great world, right? Sometimes you go build it and then you merge it and then it actually gets released 3 months later. But we're going to assume we're somewhere fairly modern, like a Netflix or a Meta, where engineers are capable of shipping 100 times a day or a thousand times a day, but it still takes two, three hours to do the work.
And now with an agentic factory, what you do is you take out that person building the thing and you replace it with an agent building the thing. And so you have orchestration to trigger things. You have a sandbox, you have an LLM, you have an inner harness, you have an outer harness, which is the dev environment you build for the agent. And maybe you give it a browser, you give it a video recorder. If you use things like Cursor background agents, they've kind of built this outer harness around the inner harness that is the coding agent. And then you make PRs with that. Problem there is that, okay, now it takes 10 minutes to do a build instead of two hours or two days, and so now the bottleneck is code review. So okay, let's throw a bunch of AI agents at code review and let's do agentic testing so that we can basically catch a lot of the easy stuff, and humans are only focused on the most important, critical core parts of the codebase.
And then the next level up of your agentic factory is you do the top. It's like, okay, then it gets deployed, it goes to prod, and a user complains. You just hook your support queue right up to the agent. Someone complains about something, agent tries to fix it, and instead of looking at a ticket and then saying, okay, go send, you just close that loop, and instead every time something goes wrong you just get a PR. And then every time something crashes in Sentry or Datadog or whatever, it goes into the tracker, it gets picked up by an agent and you get a PR. This is the Ramp Inspect thing. The only difference is then you have so much code to review, and people say, well, let's try turning the lights off. Let's just take all the human testing and review steps out and we'll say, okay, cool, if users complain then it's broken, and if users don't complain then it's working, and we're not going to read the code. We're going to treat the whole system as a black box.
So you said you tried this out when it was Opus 4.1, and you built the software factory, it was running beautifully until it just blew up in your faces. How do you think of this model? Because I can see an ideal world where it works, but clearly we're not in an ideal world. Where do you think we are, and could some of this actually work at some point? What progress are you seeing right now, and what is the situation today? How much of this do you believe we can automate or should we automate?
Yep. So if you know me, you follow my stuff, you know I stand for three things. Number one is cutting through the hype and the jargon and going trying things and talking to people who are using things and figuring out which parts of this actually work and are valuable. Number two, we talked about words. I try to find and protect useful bits of language because I think it helps us all move forward. And when you take a useful word like agents or a useful word like software factory and then you semantically diffuse it, this is another Martin Fowler word, everybody likes the word and it all becomes hype. Agents means nothing anymore. Agents could be a chatbot, it could be a Slackbot, it could be a coding agent, it could be tools in a loop, whatever it is. So I like to protect important useful words and help us all elevate the conversation out of that hype and jargon.
And then I care a lot about going one level down beneath where I'm generally working. This is the same thing with context engineering: I was rarely actually going and building LLMs or training LLMs, but knowing how they're trained, how transformers work, informs how you build at one layer up. And for the software factory, my version of that is I spent the last couple weeks going really deep on reinforcement learning with verifiable rewards, RLVR, which is very productionized. RLHF is still fairly academic, and RLVR is this machine in these labs of how we train these models, and I'm studying the benchmarks for coding agents and the
techniques for training them and how we give it a small problem, have it solve it, delete the test changes it made, revert them, apply a test patch, see if it passed. And then even the frontier this year, we have — we can get into this later — but Frontier Code and Marathon, these new benchmarks that are supposed to be better at evaluating models' ability to maintain a codebase over time and write maintainable code. And they are better, but I don't think they're sufficient.
But it's basically this idea that the only thing that made Claude Code good was reinforcement learning. And the dimension along which it got good was: we made a model, we trained the model and the harness together. And so the model got really good at calling the specific tools in that harness. Really good at reading files, writing files, searching for files, all this stuff through doing these problems. And that was what made it feel so much better than all the other CLI coding agents that came before it. And so people were like, "Okay, that was so much better." And they're just going to keep getting better. But it's like it got really good in one dimension.
And the dimension that they're not getting better in, because it's hard, expensive — maybe we need to get a lot more creative with how we design these verifiers and benchmarks — is in how do I make code that in three months is going to improve the productivity of humans and agents, mostly agents, but humans and agents in the codebase, instead of making it worse over time.
And so you think that part is just missing? We haven't seen too much improvement.
I haven't seen — obviously no one knows what the labs are doing internally, because it's all very secret. But I think if we're looking at where the benchmarks — the benchmarks tend to reflect where the labs are, right? There is no benchmark that can convey to me: did this model write code that is going to make my codebase better or worse? The best we have is — I think Frontier Code from the Cognition team is really interesting. They have: did the tests pass, and then they have two layers of model review. So they have a judge model that checks, okay, is the patch the model made similar to the patch that is the golden answer set? So even if the model didn't write the exact code that the benchmark was expecting, was it functionally equivalent? And the next one is a code quality review from another judge model. And that's better, but it's not sufficient.
And this is why I also think agentic code review is — yes, it will catch things and it will raise your floor, but I don't believe — the model writing the code is the same model reading the code. And if you ask a model, "Hey, is this code good?" it's going to be like, "Oh yeah, it's great, comprehensive, it's got unit tests." You've tried this, I'm sure. And you say, "Okay, review this PR that my coworker wrote and tell me everything that's wrong with it," and it's like, "Oh, it has this problem and this problem." They're sycophantic and they want to tell you what you want to hear. And so it's really hard for me to trust a model to evaluate the quality of code that's written.
And so I have some ideas on, okay, can you build a benchmark where the model builds 20 features in a row and maintains the codebase the whole time, and it doesn't know what features are coming? You treat it like a real product team, where you don't know what you're going to build next week until you get there and you find out what's most important. And then can we try to evaluate — can we build a problem like that that's hard enough that most frontier models fail by issue six or seven.
Is it fair to say that, you know, we've had the software factory before AI? It was just lots of loops. It was the PM giving the ticket to the dev, the dev building it, deploying to production, customers using it, customer support getting tickets, and then the PM triaging, and it kind of goes around in this loop. Is it fair to say that the software factory — of how a company, a team builds and maintains software — that is changing, because now everyone's replacing some parts of it? You know, maybe the least advanced teams will just be: devs are starting to use Claude Code or Codex to write faster. They're not spending as much time on there. Some others are also having the deployment, the feedback. Some actually have the agents already one-shotting bugs. So is it fair to say that the software factory is just changing everywhere, maybe at different speeds? But I think every team who is building production software, they're frantically experimenting, trying, and everyone's at a different pace. You'll have the AI-native startups where most of this will have agents in them, and you'll have the laggards, or more cautious ones. They have agents in a few places but not in the others.
Well, and I think that's the key: if you want to do loops engineering, you should build one loop at a time and you should keep them small and contained. Basically, I think everything except "stop reading the code" is really good advice. Take support tickets and turn them into tickets in your system, and then maybe turn those into PRs. Great.
The advice that I have, and what we are kind of chasing at HumanLayer, is: how can I add another checkpoint in that factory? So instead of having one human review point where you're reviewing PRs — and sometimes they're 100 lines and sometimes they're a thousand lines, but it's quite a lot of effort, especially if it's bad, especially if it needs rework. It's quite a lot of effort for a human to be like, okay, this is wrong, go change it in this way, and then you loop back to the agent and then you come with another one, and doing a lot of loops on there. Once the direction has been committed to, it's really hard to steer off. You're better off just kind of restarting from scratch. How do you build controls and mechanisms around that?
And then my take is, if you do a little bit of human-agent planning and discussion before you hand it to the implementer — whether it's, I mean, planning and specs, whatever you want to call it. Again, spec-driven development is another word that has become very muddled as far as what it means. But basically, how can we spend an hour before we start building so that the PR, when we read it, only takes 20 minutes because the code is perfect, instead of not touching it, just literally saying every user-reported issue becomes a PR through the loop, and then we read that PR and it takes six hours because there's back and forth and we have to make changes and things. I'm all about: let's find leverage.
And so you basically have three options in the software factory world, if you're going to go all in on agentic software factories. You can turn the lights off and just let everything flow and pray that you don't create too much slop, and pray that the next generation of models comes fast enough before you create a giant pile of ash. You can slow way down and read every PR and read every line of code, and then you're only going to really get modest benefits from AI, because that becomes — I think you should expect maybe 30 to 50% lift in productivity is kind of what I see when we go into teams.
Or you can find the right leverage points where humans can actually — an hour spent over here in planning can save you four hours in implementation, in terms of fixing and going back and getting the design right. And that's what I call seeking leverage. If you can find the right leverage points for the agents to guide the work, then you can actually move two to three times faster while maintaining a 99% accuracy to: if the humans were carefully writing this code by hand, how would it come out?
Now, jumping a little bit back to ideas — I will come back to this. This was earlier, maybe it was last year, but you had the research-plan-implement. Can we talk about the original research-plan-implement framework, and then also what you've learned about this approach, what you got wrong about it?
Yeah, sure. So the first time we talked about RPI was in August of 2025. And it was basically: the research was this thing of, hey, before you go build anything, go read lots and lots of code. Use a bunch of subagents in parallel, understand all the code. It was this technique that worked really well for hard problems in complex codebases. If you just ask Claude to do a thing, it would read three files and make a change. It would have no context. So you start the research. You don't even tell it what you're working on. You just tell it, "Hey, can you tell me how this system works and this system, and how they connect together?" And then you get a markdown doc out. And this is the context engineering part: that would take 100,000 tokens of context, but you would get a 10k-token doc out of it that summarized it.
Then you would start a new context window and you would do planning. And the planning would be — and actually I realize the plans that we were building last summer were actually terrible. But it would basically be this long — you would say, "Okay, now here's what we're building. Here's the research doc. Build a plan to implement it."
And in retrospect, now that we see everyone is obsessed with how do I get agents to work for longer, I think the reason why in May, June, July, August of 2025 a lot of people became really interested in planning was it was a very powerful lever to get agents to work for longer. If you said, "Build me a B2B SaaS for burrito delivery," you'd get a homepage and that's it. But if you said, "Build me a plan," it would build out this big plan. And then in the next context window, you'd say, "Hey, here's the plan. Here's all the changes we're going to make. Go implement," and it would actually keep going until the plan was done. So the plan was a really good way to anchor an agent and remind it that, hey, you're not done until this is all finished. So that was the original RPI.
And the plan doc — what was bad about it is it didn't give you leverage. The plan was every single line of code that was going to change, in diff blocks, and all the new stuff to write. And so people would review these plans. We recommended this. We told people to read the plans. We read all our plans. And then eventually I found myself — I just kind of skimmed the plans. And so you're not really using it as a way to resteer the agent. It's just kind of there. And then you go write the code and it's crap. Some people would review the plans and the code, and it's like, okay, well, the plan took you 20 minutes to read, and then the pull request takes you 20 minutes to read, and they're different. And so you actually doubled the amount of time you're spending reading code instead of doing less of it. You've got anti-leverage.
And hang on, was spec-driven development not related to this? The one that Amazon Kiro, for example, and GitHub workflows, again a year ago, did — which was, it also first generated a plan and it had the human review it, and you could edit it as well, and then it went off and implemented this part. And it looked beautiful on the surface. It should have worked great, but it's tossed into the garbage outside of some maintenance projects. I think it just didn't work. All the feedback I got, people just stopped using it because it just didn't really work that well. It just rhymes with the RPI framework a little bit, the original one, right?
Well, so our thing too — the biggest difference between RPI and spec-driven development — and some people refer to RPI as spec-driven dev, because for some people SDD, all it means is "I use a bunch of markdown files while I'm coding," and forget what's in them. I'm spec-driven — those are my specs and I'm using them to drive development. There was this OpenAI researcher who talked about spec-driven dev, and like, hey, stop reading the code, just write the specs and treat the coding part as compiling specs into code. That part never really materialized. Maybe with GPT-7, you know.
But the challenge — I'm on a GitHub issue in Spec Kit that has been open for a year, and every couple weeks there's a new email on the thread of people complaining about this problem of, okay, I edit my specs and then I edit the code, and then the code drifts from the specs. How do I keep the specs up to date as the code is changing? And it's basically like you now have two sources of truth, and it stops being useful.
And so that's why with RPI, the idea of the docs is — for a while we kept them around, but after two or three months we're like, oh, these are actually tactical execution docs. I do the research, I do the plan, I do the implementation, I throw the docs out. And the next time I need research, I just do it from scratch, because tokens are cheap and my time is expensive, and the amount of time I might waste if I reuse a research that is no longer in sync with the real state of the codebase. So we just create it live every time.
This is why context engineering still matters. Creating artifacts that compress the state of the codebase and compress the intent of the builder into small things that can be reused in the future for the scope of a task is a very powerful tactical approach. But it's not a thing — I have very few opinions on what sorts of docs you should leave lying around your codebase that are evergreen. I've seen people try to maintain parity between documentation or specs and the code itself, and I don't think anyone actually found it very useful. You can do it and it works, but it's the ratio of the effort it takes to keep them up to date — and trivially you could do this with AI, probably — but I've never known anyone who was like, "Yeah, this is great and we're glad we have it." You could do it and it might help, but I don't think anyone found it useful enough to maintain a system to keep the specs and the code in sync, versus just using the code as the source of truth always.
Now, you mentioned something interesting, which is with context engineering you need to sometimes compact. And you've previously talked about intentional compaction: that when context is noisy, deliberately compress the useful part into a clear markdown artifact, verify it, and then start a fresh conversation. Can we talk about this kind of compaction and why it's important? And it sounds like it's going to be a building block, or it already is, for context engineering, right?
Yeah. No, frequent intentional compaction is the building block. It completely comes from context engineering. Context engineering is: how do we get the most out of today's models? How do we change what we're putting into the model, into the context window, into the agentic chat? How do we control that in such a way that we get the best results possible, which means doing as much work as possible in the smart zone, the first 100,000 tokens of the context window?
And this frequent intentional compaction is basically: okay, the research step, we're going to go read a bunch of code and turn it into a doc. That's our compaction. We take that forward in the next session. We're going to read the ticket and the intent and turn that into a design document, which is like, okay, here's the high-level spec of what we want to do. Here's a high-level current state, desired end state, and then a bunch of design questions the model has — kind of a very thorough, maybe even overengineered, plan mode. And then you take the research and the design and you do a new session, new context window. You're like, cool, you've compressed the intent and you've compressed the state of the codebase, so that you can then do your planning of, okay, we know what the end state looks like. We know where we're going. Now let's break down how we're going to get there. All of these different steps of the process exist because models have
shortcomings in each of these phases. So, the research is pretty hands-off. I don't read the research docs. It's just like go read a bunch of code and then make a doc out of it. Models are pretty damn good at that. If you ask it to find a bug and have opinions about the codebase, that's different. But if you just ask it what is the intent and how does this stuff fit together, that's usually pretty straightforward.
But designing the end state of the software, the architecture and the program design, models are not great at. They make decisions and sometimes they're right and sometimes they're wrong. So we want to have a human in the loop there.
And then the steps to get there, we talked about this before, but models love making what I call horizontal plans. If you ask a model, build a plan of steps to go build this app, it's like, cool. We're going to do the database and then we're going to do the services layer, then we're going to do the API and then we're going to do the front end. It's like, well, that actually kind of sucks because we're going to be on the other side of 2,000 lines of code, and let's imagine this is an existing codebase, right? We're going to make changes to all these different parts of the system. I can't test it till the end.
And so what I would do is like, okay, how would I have built this if I were building by hand? Well, okay, I would probably create a mock API endpoint with fake data. And then I would go kind of get the front end how I want it to look. And then I would actually go build a services layer and actually wire the data through. And then I would make a database migration and make my new table. And then I would actually add a lot of business logic. And then I would add a bunch of error handling.
And it's completely orthogonal to how models would write the database layer and all the error handling without anyone's ever touched or seen the code or whatever it is. And so this is another place where we like to have humans involved, because humans have really good taste and judgment. Like, I would rather read five separate little mini diffs of things that I can manually verify and explore than read 2,000 lines of code and be like, well, it's not working. I don't know where. You don't know where cuz you wrote the code. You were supposed to get it right.
We talk about compaction, context engineering. It's like, how can you stay in the smart zone of the context window, which is again the dumb zone. I will say, disclaimer, it's really good training wheels if you don't have intuition about this.
So let's just define these things. What is a smart zone and what is a dumb zone?
So, it's a little bit blurrier than I would like it to be. I think in November we talked about the first 40% of the context window, but then we had million
Smart zone.
Yeah. Then we had million token context windows. So then I changed it to like the first 100,000 tokens. If it's a really, like 4.8, I usually will go up to like 200k. But basically the thing Geoffrey Huntley had in Ralph Wiggum was like, the less context window you use, the better outcomes you'll get. And basically the smart zone meaning if you have context in that first part, it should work a lot better. And then the dumb zone is like, once you have stuff there, it's kind of forget about it. It'll be confused, it's not going to do much, it'll degrade.
Yeah. And there are times, and this is an intuition thing, I will often go up to 300, 400k tokens. Four is rare, but I will go up to 250, 300k tokens for certain types of work where my intuition tells me that I can keep working without degrading the performance. But if you don't have good LLM intuition, like 100K for smaller models, 200K for these really beefy Codex and Opus 4.8 models is usually a good training wheel guideline of, if you pass there, your quality of results may be degrading.
The biggest tell I see for this is often the model's trying to get the test to pass and you're at 200k tokens. Well, let me try this. Okay, let me try that. And it's trying a bunch of stuff and it's getting more and more extreme, and it's like, oh, let me delete your .env file and try again. This is where things get really, really weird. And so if you start to see certain types of, if I'm like, oh, we're at 300k tokens and I need to fix the unit test, I'm like, cool, write everything we did to a file, or even I'll just do a built-in compaction depending on the model, and then I'm starting a new session at 30k or 50k tokens and I'm like, cool, we're going to do a hard thing, which is you're going to get this freaking test to pass and you're not going to be stupid about it.
By the way, one thing that you said about the model being dumb is you said that if the model ever tells you you are absolutely right, you should start over. And we've all had that when it tells me like, oh, you know, you didn't, you're absolutely right, and we just get annoyed. But why should we start over? What's happening there in your observations?
Yeah, that's great. And the new you're absolutely right, I think, is you're right to push back on that, right? Yes.
That's Opus, right?
Yeah, Opus is like, you didn't run the test, did you? Right to push back on that. I totally did it. But no, for me, you're absolutely right was always what the model would respond. If you were like, "That's totally wrong. You did it." If you said something where you were angry or frustrated or just wanted to point out that it's done something wrong, it would respond with, "You're absolutely right." And most of us have had the experience of it says that and then it continues to do the wrong thing.
So, it's like once it starts doing dumb things, because there's four things in your context window that matter. There's the size of it, how many tokens? There's the quality of the information, like is there any incorrect information? Like if the model had some thinking trace where it decided the wrong thing was true. Is there missing information? Does this have context missing that it should have? And then there's the trajectory. And the trajectory is very subtle, but you may have had sessions.
The trajectory meaning you're prompting
The actual history of everything. I call it trajectory, it's the actual history of what the agent has done in the past.
And so if I say, "Hey, make this change," and the agent makes the change and then it runs the test and then they're broken and then it fixes the test, I have very high confidence the next change I ask it to make, it's going to follow that path again, because it's like, okay, here's a conversation, and the last time the user asked me to do a thing, I made the change, I ran the test, test broken, fixed the test, and then I told the user. But if I say make a change and it makes a change, it doesn't run the tests, then I'm on a different trajectory. And if I say, okay, make another change, basically they're autoregressive. So they're predicting what's the next message in this conversation.
And so the example we talked about in No Vibes Allowed was of course, hey, the model makes a mistake and then you yelled at it, and then it made another mistake and then you yelled at it, and then it's like, cool, what's the next message in this conversation? Well, look, if I read the history, I should probably make another mistake so the human can yell at me. So I was like, okay, that's a great example of time to start over.
Let's talk about some observations on how software engineering is changing. One thing you talked about recently on the evolution of the coding meta is going from token harder to token smarter. Can we talk about what you mean by token harder and token smarter?
Yeah. So token harder is, I mean, I'm in a group chat called hyperengineering and it's all people trying to max out their Claude subs.
Oh wow. Okay.
It's just like
That sounds like a fun, is it a fun place?
It's a fun place, but it's all token harder. It's like, look at all the side projects I built. Look at everything that I've gotten my Claude token. I've got six Claude Code accounts. I've gotten all of them maxed out every 5-hour period. I've timed it out so I always use all the tokens and it starts up immediately when the limit resets.
And so it's like, I mean, getting into Eli Goldratt and The Goal, it's optimizing for utilization and efficiency of one node in your factory rather than the end-to-end goal of how do we ship value and things that people like, that are stable and will last a long time. But that's my idea of token harder, and it's the same thing with the dark factory thing, it's like, hey, if you remove humans from code review, you can push more tokens through the system.
So we talk about software factories, but what is the dark factory?
Ah, so the dark factory, this comes from this idea of, there are factories where everything is automated by robotics. So you can imagine a car factory where it's all robots building the cars, and they don't have lights because there's no humans.
Oh, so that's where it comes from.
The dark factory. Yeah. You walk in, there's no lights. There's not even light switches.
So, it will be the fully automated software factory where it will be no human input, basically.
No human input. Raw materials go in, cars come out.
Yep.
And I think in a micro, you can have many loops that are dark in your thing of, hey, if the code review agent comes back with a problem, you loop that back to the builder agent, it fixes it and comes back, and that's dark. You don't need a human loop for that. But the full dark factory where you don't read any code, yeah, it's a good way to maximize your token utilization. And it's like, if your belief is, my job is to extract as much intelligence out of the machine god as I can because that's how I get the most value and the most leverage on my time, then token harder.
And my take is basically what we talked about before. Token smarter is like, okay, how do I move faster? How do I get as much value out of AI as I can without having to turn the lights off, while still maintaining control and taste and judgment and understanding the system architecture, and applying my hard-won opinions through 10 years of software engineering to the design of the program, so that I can feel confident that the code's going to get better and more maintainable over time.
It's the same thing of, you look at the SRE team inside Google. They brought out this book, Site Reliability Engineering, and the whole take was like, hey, we're going to go from one data center to five data centers and we need the same six-person team to be able to manage five data centers, and we need the same six-person team to be able to manage 50 data centers next year. And it's basically, how do we apply software to this problem so that instead of scaling linearly of, okay, every data center needs five DevOps people so we need to scale the people with the things, how do we continually automate the parts that we don't need? So a little bit orthogonal and maybe even contradictory to what I just said, but this idea of how do you find leverage, and the way, the way
Well, I think what you were saying there is, when Google did that, they never sought to remove those SREs from the process at all. They just said, look, can we think ahead and scale yourselves, and they actually grew the team. It wasn't actually six people. It was more like, I think Google specifically said, "Okay, we have five data centers. Next year we'll have 50. There's six of you. We do not want to have 60 people," and then management layer and all that. It's like, how can we do it with like 12 or like 10, and then when we'll have 500. And now actually their SRE has grown, but
Of course, yeah.
But they never, you know, I think as engineers we feel pretty threatened when someone says, all right, we just want to have zero engineers. I mean, that's not a fun place to work at, but what it sounds like
It's not a possible place to work at. If they have zero engineers, neither of us can work there, right?
But do I understand the token smarter is like, let's keep humans in the loop, let's keep adding value and figure out what are the parts which are not as relevant, boring, where we don't need it. And so one developer can probably do more than before, but you are built to be part of this whole thing, and the lights are on in a factory.
Yeah. And basically, I think what I'm trying to get to is the connection here is SRE built a thing where headcount scales at a square root function or a logarithmic function, whereas their output scales linearly, and the way you do that is with good architecture and good program design. And so in order to avoid this problem where you have to throw more people or more tokens at the problem, if you design good software in such a way that it gets more maintainable and more scalable over time, and just today it doesn't feel like, basically you need humans in the loop to be able to do that.
Let's talk about AI slop. At one point you wrote, "Yeah, AI can write your code, but it can also write your specs and PRDs. But the same rule is always slop in, slop out. If you outsource your thinking, you're gonna get garbage."
Yep. So yeah, that's basically the idea. The way we think about getting high-quality outputs is, yeah, you could write the code by hand, or you could sit with a model and work back and forth and go maybe a little bit faster, and you have control, and every time it makes a change, you go read the change, and if it's bad, you tell it, nope, we want it like this, and you kind of incrementally, slowly.
This is kind of the stage two or stage three version of working with agents, where the agent's writing all your code, but you're very much in the loop. And this will make you go faster, but it won't make you go that much faster. It won't make you go anywhere near, there's that level, and then there's the maximum speed you can go while still caring about the code, and then there's the maximum speed you can go if you turn the lights off.
And so we always think about it in terms of leverage. It's like, okay, everything starts with a sentence or a voice note ramble, like, I want to build this thing, it's going to work like this, or whatever it is. Let's say on average two sentences: I got to fix this thing, or there's a support ticket, I got to fix this thing. If you can turn that with AI into a one-pager, and make sure that's correct, and then turn that one-pager into a three-pager and make sure that's correct, and then turn that three-pager into a 10-page detailed outline, then you can write 100 pages' worth of code. And it's maybe not perfect, you shouldn't sweat over these documents and make sure they're perfect, but you're increasing the chance that, you're decreasing the uncertainty of the outputs. You can think of it like you have a line of where it's going, and then you have the probabilities of where it might go in that range. If you are reviewing along the way as you get more and more detailed into what you're building and how you want it to be built, you kind of collapse the uncertainty and the set of end states that you could land in.
That's me doing the physics thing of, you got to superimpose all these probabilities. And I don't know, I have this thing that I think people who really like playing real-time strategy games are probably going to be really good with AI, because you kind of have to, I don't know. Matt Pocock was just talking about fog of war and things that are at the frontier of, there's stuff we don't know about this problem
yet. How can we find that out and how can I make the best decision now knowing what I have seen? I've seen a couple pieces of information and so there's a 30% chance it's this and there's a 40% chance it's this. How could I get more information? So in my head I can recalculate those probabilities and decide what's the most likely path that's going to lead us to success.
Speaking of the most likely path that leads you to success, let's talk about your company that you've just come out of stealth with, HumanLayer. What is HumanLayer and what is the probability that you're setting up for success?
That's a good question. 100%. 100% probability, maybe 110. But no, so HumanLayer is an AI IDE, it's a collaboration platform, and it is building blocks for your software factory. And the basic pitch is engineers solving hard problems in complex codebases. Basically there's two categories of builders: there's vibe coders building side projects, and then there's people building production software where the stakes are high and if something breaks we're going to get fined millions of dollars or, you know, we're going to lose millions of dollars of money for the company. And there's a whole spectrum in between there. But if you're kind of in the left half of that spectrum, you're building software that matters and it has to last and be around for a while, then helping people like that solve problems two to three times faster without descending into slop is like, how do you maintain that near human level of quality and move two to three times faster?
And what were the ideas that you built and that you came with?
One idea that we're really excited about right now, I mean, it all comes from this RPI and this using specs to... I mean, I've kind of been hinting at it this whole time, right, of okay cool, start really high level and zoom in layer by layer and resteer and find that leverage that helps you move faster and increase the chance that your agent's going to build exactly what you want or something that's really high quality.
The other thing I think that's really interesting, where I just posted yesterday, I said, "Hey chat, should we kill the pull request?" And that's something I can't talk too much about, but basically the idea is the IDE of the future needs to be rethought from the ground up for agents. And it might not even be a... I don't know, a lot of editors kind of started with the text field and bolted on an agents tab. And then eventually you've seen Cursor 3. I can't even find the text editor. I know it exists. People have told me you can get to a text view of files, but it's also very agent first.
And so we started from the ground up of what is an IDE designed for helping a developer interact with and manage the work of agents. And then we zoomed out and said how do we make this collaborative and build in a sync engine and durable streams and all of these pieces of tech that enable me to get human input and feedback on what I'm doing with agents in real time rather than waiting for the pull request time.
And great engineering teams have been doing this for decades of, hey, we're gonna have a design review where we're going to talk about how we're going to build the thing as a two-page Google doc or whatever, 10-page, however,
BRD,
architecture requirements document, and then you go to sprint planning and you break it down into little tickets and you decide who's going to do what. It's like AI can help with all of this. If you're just using AI to write the code, you're missing out on a lot of the benefits that AI can bring to your SDLC. And a lot of people say, "Well, we don't need any of those meetings anymore because we have the loop. We have the dark factory. Things just fly around the loop." But it's like, "Okay, but if you want to actually move faster and maintain quality, then you should have these checkpoints before you go to actually write the code and you should use AI to help with that."
So, we built this cloud platform that kind of has a Google Doc style component where you can comment and the agent can surface mockups and Mermaid diagrams and HTML and all these things. So, basically, how do we make agents Figma style? Everything's in the cloud. Everything's collaborative. I see all my coworkers' sessions. They see all of mine. It's almost like the benefit that Slack had over email was that you didn't have to be in every conversation to know what was happening. You could see all these channels light up. You could check on them. Okay, I don't care about any of that. But if you saw a conversation that you cared about, you could jump in on that.
And it's like, how do we do that for engineering work, versus we really had these very strict, even when we called it agile, it's very waterfall: PRD, ARD, tickets, everyone goes and builds for a day, and then you get the PR back and then one person reviews it. How do you create this more just soup, and what is the data model for that world where you have agentic traces, you have documents, you have tasks and projects that group these things, you have actual git diffs being streamed everywhere, where it's like, why would I review all the code at once when everybody's work lives in a shared environment that anyone can go interact with?
I mean, what it reminds me of is what GitHub did. Software teams before GitHub and its competitors, you might have a tracker somewhere, but most teams were just kind of, inside the company, you didn't know what one team was... I remember pre-GitHub, you had individual teams, some of them had a board with stickers, but no one else in the company knew what they were doing. They were all working in isolation. And now when you have GitHub, or even the internal version of GitHub inside a company, you can always see, when you go to a team, you see the pull requests flying, you can join in, you have history, it is all kind of connected, and it came together. And now, for a very long time, I was like, duh, you're going to use GitHub, or people will copy it. So do I sense that you're trying to build something like this workflow for when you have the software factories, which are like dark factories and loops at a bunch of places? How can we have this new way of working, which will feel natural, but coming up with it is hard work and it's counterintuitive.
How can we do something that accomplishes what GitHub did but 10x better, more specifically more continuous and more real time and more collaborative than these discrete units of work that is the pull request.
Well, now I'm starting to understand why you're saying maybe we should kill the pull request, because the pull request was invented by GitHub, right? It is not part of Git, but they did it as a way for you to do a code review merge before it goes in and be able to modify it or just reject it, etc.
And it's probably a lot better than whatever we had before, which I guess was emailing your git patch to Linus and asking him to merge it into the kernel or whatever.
They still do that. It works for them. That's the point. But it only works for them.
Yeah. I don't know anybody else who does that. I mean, I'm sure even before GitHub for you, you guys had what, CVS or
CVS? So, if you had a lot of money, for Microsoft,
They made us use Subversion in undergrad because the guy who invented Subversion was a UChicago guy. The year after I graduated, they switched everybody to Git. And I was like, damn, I learned a useless thing just for somebody's ego.
Specifically for AI startups, or startups building on top of AI or building AI products, how important do you think location and network is, especially as you are based in the Valley? We see research that AI startups are more frequently funded from here than normal startups as well. Do you see this advantage, and also do you see some disadvantages of being in a specific place, may that be Silicon Valley or elsewhere?
I don't have really strong opinions on this. Actually, Paul Graham gave a talk in Sweden about why SF is cool. Rather than just regurgitate that, I will forward people onto that one. We can put it in the show notes or whatever, but he talks about all of the dynamics of Silicon Valley and the pay-it-forward culture and how people take you way more seriously just because you're based here.
I lived in Chicago for a long time. I have a lot of really good friends from high school, from college, from growing up in LA. And never before have I felt so locked in with my people more. Never have I felt more seen, more connected. There's just so many people here. Again, talking about the founder thing, people who care deeply, who are incredibly competent, who, we have all the same types of problems. We love all the same types of things. I don't do LAN parties where we play video games, but all my buddies will come over and we'll sit in the office till 11. We'll just do co-working and hack on cool fun projects and stuff. And you can't do that anywhere else. There's not enough critical mass for that to just happen organically everywhere you go. And I absolutely love it. I wouldn't trade it for anything.
Yeah, I think critical mass nails it on the head. When it comes to hiring, what types of folks are you hiring for specifically? Because I'm interested in how hiring changes and what a standout engineer means for you and how you are trying to, you know, confirm that those traits exist.
In general, we are looking for people who have really strong software fundamentals. So, understand distributed systems, understand the core fundamentals of CS and operating systems and these kind of things. I mean, you don't have to be a PhD in freaking kernel design or whatever, but it's a lot easier. We can teach somebody, I think, to be a really good AI developer in a few months. You can build enough intuition where you are, you know, accelerated off the ground and you can keep growing there. It's really hard to teach someone a CS undergrad program in 3 months.
And what's a problem space that you're excited about in software engineering, or even product engineering or building products, that you think in the next few years is going to be one of the interesting things that you're going to be attacking?
My co-founder could talk more about this, but there's a lot of interesting things happening in real time, in cloud and sandboxes, in sync, and kind of using these new building blocks that have gotten really solid in the last couple years. We're big fans of the ElectricSQL team, where users have durable streams. It's like, how can you build systems that kind of are a lot more spread out and distributed and almost decentralized? This is really interesting for coding because you want to be able to run coding agents anywhere. You want to be able to run them for a short time, for a long time, on demand, on a schedule, all these things, and have them all be part of this kind of brain. So I don't know, parts of what we're doing are really boring, like all our data is in Postgres, and then parts of what we're doing are really interesting. But there's a lot of distributed systems problems. There's a lot of infrastructure problems. We are building tools for AI, but there's a lot of problems in building collaboration platforms that are really, really hard, and there's a lot of new tech that makes it easier and more interesting, but it's still by far from an easy problem.
It sounds like what you're saying is the infra layer is, to some extent, a new infra being built, and it'll take some time, but it'll be just new blocks and it will eventually become the primitives. Like for cloud, we have primitives already, but it took a freaking decade to get those together, or more.
Yeah. You had AWS in what, 2008? 2006. Yeah. And then you got Kubernetes a decade later.
Yep. And as closing, what's a book or reading that you would recommend? Something that you personally enjoyed.
Nowadays, we talk a lot about Refactoring by Martin Fowler, classic. I think it's because we spend a lot of time improving the design of existing code and trying to figure out how to get models to build code that is easy to maintain and easy to read and easy to understand and easy to build on. I feel like I probably have a better answer than that, but that's what's top of mind these days. We're reading a lot of classics of software engineering. Refactoring, Clean Code, The Pragmatic Programmer, all that stuff I think is more relevant now than it has ever been.
Love it. Well, Dex, thanks so much. This was fun.
This was a blast, dude. Thanks for having me on. This was great. I had a lot of fun.
I don't know about you, but I really enjoyed this conversation. Dex is such a big believer in agentic coding. Yet he's the one warning us that if you stop reading the code, you have about 3 to 6 months before your codebase becomes easier to rewrite than to fix. And this comes from firsthand experience. His team built a light software factory, ran it, and then had to shut it down.
I also like the idea of the slow loop. Loop engineering feels like a somewhat meaningless term to me. What Dex's team does is actually pretty boring. A cron job runs every night, fixes one issue or one anti-pattern, and opens one small pull request. The team wakes up to a codebase that's a little bit better every morning, and a dev still needs to review and approve it. This is a practice that honestly any engineering team could just adopt today.
Finally, I really enjoyed the history lesson. The term software factory comes from a NATO conference in 1968. The idea of software used to build software, with analogies to a factory, is more than 60 years old, and every generation of our industry has tried to automate more of the loop of building software. AI agents are just yet one more attempt, although probably the most successful one.
Do check out the show notes below for the related The Pragmatic Engineer deep dives that go even deeper into AI engineering and other related topics. If you enjoy this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks and see you in the next
Article published
