Dex Horthy on Context Engineering, and Why HumanLayer Shut Down Its Own Lights-Off Software Factory

Open on YouTube ↗
Overview

Dex Horthy, CEO and cofounder of HumanLayer, is credited with coining the term "context engineering." His team also tried running a software factory in which agents shipped code that no human read. They started it in July 2025 and shut it down by November. In this conversation on The Pragmatic Engineer podcast, Horthy explains what context engineering means and why context windows degrade as they fill. He describes how his team uses automated loops, why they believe the "dark factory" fails, and the workflow HumanLayer is building instead. His position is that he is heavily invested in agentic coding and still keeps humans reading the code. He argues that current models are not trained to write code that stays maintainable over months.

35 min read

From moon-crater pathfinding to "the thing that builds the thing"

Horthy studied physics as an undergraduate and decided against academia. He saw three exits from physics: a PhD, finance, or programming. His first programming experience came at 17, during a high-school internship with NASA researchers at the Jet Propulsion Laboratory. The team had a new high-resolution topographic map of the Moon's south pole. Craters there are deep enough that some have never seen sunlight and may hold frozen water. His task was to find a route between two points that stayed within a rover's maximum incline and decline. He had never opened a CS textbook, so he wrote what he calls a "really naive bad version" of Dijkstra's algorithm.

After half a CS minor, he joined an API platform team at Sprout Social in Chicago. Within a few months he concluded that the most valuable people in the company were the ones building the developer platform: CI/CD, sandbox environments, preview infrastructure. He says he has been "obsessed with software factories" ever since. He is surprised when developers say they hate CI/CD. In his view, lazy engineers should want to build "the thing that builds the thing," because that is the highest-leverage work.

A stint at the consumer fintech Aspiration followed. The VP of engineering who hired him left within three months, and Horthy served as acting CTO for a while. He came away convinced he is "a B2B guy." He then spent about four years at Replicated. The company built its own container orchestrator before Kubernetes and Docker Swarm matured. The goal was to let SaaS vendors ship installable versions of their products that customers could run in their own data centers or VPCs, much like GitHub Enterprise. After two years in core engineering, he says he kept arguing with the CTO about the build pipeline. His manager told him to stop fixing CI and do his tickets.

He moved into a customer-facing engineering role. Replicated's customers included HashiCorp, DataStax, Puppet, TravisCI and CircleCI. In his first three months he met every stalled prospect in the pipeline, and he recounts closing about 12 deals. He grew that organization to around 25 people. The team later shrank after zero interest rates ended, and the company shifted toward a more product-led, self-service model. Horthy then became a product manager, carrying a "laundry list" of usability fixes from his customer work.

He also gives a personal reason for the customer-facing years. He felt introverted and socially awkward through his twenties. His uncle, a music producer who had worked with Randy Newman, once told him that greatness comes only from making something the only thing you do. Horthy decided that instead of reading self-help books, he would "make it my freaking job to just talk to people." He says the move also forced him to become organized, which did not come naturally. He recommends that every engineer spend a year or two in a customer-facing role.

Metalytics, the pivot, and 12-factor agents

Around 2020, Horthy and a friend in Chicago started Metalytics, a data engineering company. HumanLayer is technically the same company after a pivot. He describes the data-tooling market of that era as a funding boom. By 2021–2022, people realized the total addressable market was smaller than assumed, and both fundraising and customer acquisition became hard.

When his cofounder burned out and left on good terms, Horthy started experimenting with AI agents. Frameworks such as LangChain and CrewAI were popular then, with large Discord communities and plugin ecosystems. He saw a missing piece: controlling which tools an agent calls. In a chatbot, you can show approve and deny buttons. Horthy cared more about "outer loop" or proactive agents that run in the background and are triggered by events. He cites OpenClaw's heartbeat model as the biggest current example. He would not trust such an agent unless it could reach him through Slack or iMessage, and unless he could deterministically approve, deny, or deny with feedback. HumanLayer entered Y Combinator in fall 2024 with an API platform for this. He compares it to PagerDuty, but routing approvals rather than on-call incidents.

He then talked to about a hundred engineers who were shipping AI into enterprises on six-figure contracts. Most had tried the popular frameworks for a month or two and dropped them. They wrote LLM calls by hand and built systems that looked more like pipelines and workflows than open-ended "tools in a loop." A major influence was his friend Vaibhav at Boundary, which built what Horthy calls "protobufs for AI." Vaibhav framed every step in an AI workflow as tokens in and tokens out. The engineer's job is to choose the input tokens that maximize the chance of good output.

Horthy distilled these conversations into twelve principles and published them as a GitHub repository. It stayed on the Hacker News front page for about two days. The factors include: natural language to tool calls, own your prompts, own your context window, tools are just structured outputs, unify execution and business state, launch/pause/resume with simple APIs, contact humans with tool calls, own your control flow, compact errors into the context window, small focused agents, trigger from anywhere, and make your agent a stateless reducer. He notes someone corrected him that the last one is technically a "transducer." Of all of them, factor three, "own your context window," is the one he considers most durable. Whether the system is agentic or a single pipeline step, the only lever on output quality is careful crafting of the input.

Naming context engineering

Horthy wrote the 12-factor document in March, published it in April, and gave a talk on it at AI Engineer in early June. About one to two weeks later, Shopify's Tobi Lütke posted that he liked the idea of "context engineering." A week after that, Andrej Karpathy argued for thinking in terms of context engineering rather than prompt engineering. Horthy jokes that depending on the day, Gemini will credit him, Lütke, or Karpathy. He adds a caveat: he did not invent the practice. He saw what the hundred builders he interviewed had in common and put a name on it. He considers clean vocabulary important in a field full of hype and jargon.

His definition is that context engineering is "de-abstracting" the layers built on top of LLMs. RAG, memory, agentic history and structured output are all, in the end, different ways of passing tokens into a model and usually getting structured output back. Off-the-shelf agent and memory frameworks can get you to about 80% and a good demo. To reach 95% or 99%, you have to go down a level. You decide what goes into the context window and in what order, which can vary by model.

Asked why the idea caught on when it did, Horthy says context always mattered. It took time for enough builders trying to sell reliable products to discover that token-level thinking breaks through the quality ceiling. That means chains of LLM calls and workflows rather than a single model with tools and a system prompt. The approach needs more work and deeper intuition about LLMs, but it offers many more levers.

On cost, he advises "make it run, make it right, make it fast." First check whether the best available model solves the problem, and whether people want the result. Engineering time is the real bottleneck. Building evals and tuning prompts costs more than simply paying for a smarter model, until you reach something like millions of requests. At that point, context engineering lets you split a task into several calls. Some calls can run on far cheaper models; he mentions an open-weights model at roughly a thousandth of the cost of Opus. The expensive frontier model is then reserved for the steps that need it. He connects this to Eli Goldratt's The Goal: find the actual bottleneck. Latency and cost may become the bottleneck someday, but usually not at the start.

Harness engineering: inner and outer harness

Around October or November 2025, Horthy posted about what he called harness engineering. He later learned that Viv, now at LangChain, had written about it a few weeks earlier. His framing was this: you use context engineering when you build an agent, and harness engineering when you use one. The question is how to engineer against the integration points of a harness like Claude Code or Codex — commands, MCPs, skills, codebase organization — so every turn produces the best possible result. It "raises the floor."

The term quickly blurred. Some people took it to mean building a harness, and others took it to mean building around one. Horthy prefers Martin Fowler's framing. The inner harness is what a product like Claude Code, Codex or Amp exposes: tool definitions and integration points. The outer harness is what the human adds to adapt it to their codebase and languages.

He is surprised that "context engineering" still means roughly the same thing a year later, given how little AI writing from 15 months ago still holds up. His explanation is that it rests on how transformer attention works. It will stay relevant until post-transformer or linear-attention architectures arrive, and he says no one knows when that will be.

The physics of context: instruction budgets and the "dumb zone"

Horthy acknowledges that longer context windows let you talk to a model longer. But a million-token version of a model, such as the long-context variants of Opus, is not a smarter model. The model's intelligence determines how well it can attend to everything in context and pick out what matters for the next tool call. He cites a 2025 study finding that frontier models could follow about 150 to 250 instructions before adherence dropped off quickly. He adds that the models were older and a later study reportedly showed improvement, though he has not looked at that data himself.

He divides context engineering into two budgets. The information budget is the familiar one: use RAG to fetch the relevant pages instead of loading the whole book. The instruction budget is less discussed. Too many instructions, especially conflicting ones, cost the model computation. That includes changes of direction mid-conversation ("actually, I don't want any of that"), since the model has to work out what to ignore. When those instructions sit far back in the window and get only partial attention, the chance of the model honoring exactly what you said 100,000 tokens ago drops significantly. He does not claim a mathematical proof. He says he is "not a PhD in machine learning," but attention is quadratic, and more content spreads it thinner.

This leads to his "smart zone" and "dumb zone." In November his team spoke of the first 40% of the window. Once million-token windows arrived, they switched to absolute numbers. As training wheels, he suggests about 100K tokens for smaller models and about 200K for heavier models like Codex and Opus 4.8. Past that, results may degrade. He personally sometimes works up to 250K–300K tokens, occasionally 400K, when his intuition says the task can tolerate it. He attributes the core insight to Geoffrey Huntley's Ralph Wiggum work: the less context you use, the better the outcome.

His clearest warning sign is a model deep into a session trying to make a test pass. It tries increasingly extreme things, up to something like deleting your env file. At that point he writes everything done so far to a file, or uses built-in compaction. He then starts a fresh session at 30K–50K tokens with an explicit instruction to fix the test without being "stupid about it."

He treats this as a skill learned through suffering, not from textbooks. He quotes his friend Jake from Netflix, who said at AI Engineer that you learn bad patterns by debugging them at three in the morning.

Trajectory, and why "You're absolutely right" means start over

The host raises Horthy's advice that when a model says "You're absolutely right," you should start over. Horthy says models produce this phrase when a user angrily points out a mistake, and then they often keep making the mistake. (He jokes that Opus's newer version is "You're right to push back on that.")

He lists four properties of a context window: its size; whether it contains incorrect information, such as a reasoning trace that concluded something false; whether information is missing; and its trajectory. Trajectory is the history of what the agent has actually done. Models are autoregressive and predict the next message in the conversation. If the agent previously made a change, ran the tests, fixed them and reported back, it will likely follow that path again. If it skipped the tests, it is on a different trajectory. In his "No Vibes Allowed" talk he gave this example: the model makes a mistake, the user yells, it makes another mistake, the user yells again. The most plausible next message is another mistake. That is when to reset.

Loop engineering: back pressure and slow loops

Horthy traces current interest in loops to Ralph Wiggum. About a year before the recording, Geoffrey Huntley visited San Francisco and showed that he had run Sonnet around the clock. He spent about $6,000 over six weeks and built an entire Gen Z–themed programming language, including a stage-two compiler written in the language itself. Horthy's main lesson from it is back pressure: letting the model check its own work by automating feedback through linters, unit tests and similar tools. A programming language worked well for this because it is endlessly verifiable. If you can make a problem verifiable, you can treat it largely as a black box and let it loop.

His simple example is speeding up slow CI during a release. He gives the model a goal and a fixed set of steps: change code, commit, push, launch a subagent to watch the job, then decide what to do next. He compares this to the goal commands in Claude Code and Codex and to auto-research, where a prompt keeps the model trying until results improve.

He warns against getting carried away and rebuilding everything as a grand agent-first "dark factory" meant to be your infrastructure for five years. Few teams can accept a 60,000-line PR from a three-day Ralph loop that fixed every lint error, because no one will review or sign off on it. What excites him are "iterated" or "slow" loops. A GitHub Actions cron job runs every night with a simple instruction: run this linter, fix one thing, commit and push. The team wakes up each morning to one PR that makes the codebase a little better.

This can grow in two ways. You can add checks: React Doctor for the frontend, or a pattern with no deterministic tool at all. For one such pattern, his colleague Kyle wrote a description of good and bad examples. It is "prop narrowing," which makes unnecessarily optional props required so code is easier to reason about. The team now wakes up to about four PRs, one per check. As confidence grows, you can also widen the scope from one fix to four. Kyle has shipped a skill so others can build these loops. More broadly, Horthy says a good loop is one that no human has to start. It is triggered by a Sentry alert, a support ticket, a PM's ticket, a failing test or a cron schedule, and it follows a defined workflow.

The lights-off factory that HumanLayer shut down

The host quotes a Horthy tweet that did "a lot of numbers." It predicts a one-to-three-year period in which things break at 3 a.m., loops are expected to fix them, no one understands what's under the hood, and companies face existential risk. Horthy grants that with today's tools you might get away with not reading the code. The problem is that loops eventually produce more code than anyone can read. He places the strong "dark factory" position, and what he describes as Ryan Lopopolo's version of harness engineering — spend as many tokens as possible — in this category.

HumanLayer tried it. The team built a lights-off factory in July 2025 and shut it down by November. Horthy estimates it takes three to six months of shipping unread code before you realize things are getting much worse and starting over is easier than fixing. At one point they hit a bug that Opus 4.1 could not root-cause, no matter how carefully they prompted. The team spent about three weeks re-onboarding into a codebase they had stopped reading three months earlier. After several days of digging, they found that a primary key was being routed through the whole system and needed to become a different type of object. At the time, he concluded it was still worth it: skip reading code most of the time and occasionally pay two weeks of manual repair. He no longer believes that. He thinks the volume of code teams can now generate has grown something like 10x or 100x, so the problem keeps getting worse.

That is why HumanLayer's loops improve code quality rather than ship user-facing features, and why the team reads all the code. Horthy distinguishes system architecture from what he calls program design: where the interfaces and seams are, how dependency injection works, and whatever else keeps a change in one place from breaking another. Software engineering as a discipline, he says, came about in the 1970s to avoid the "giant ball of spaghetti."

The host compares this to how engineers become senior: over years, you learn how small mistakes snowball into disasters, and that tech debt can let competitors overtake you. Horthy allows that "GPT-7" might solve it. But he describes a company that sends every user complaint, crash, PM ticket, and "obnoxious essay" from the CEO in Slack to agents. It replaces human review with agentic testing and agentic review because PR review became the bottleneck. In that setup, none of the remaining components have any feel for architecture, because that intuition has not been trained into them.

Why models don't learn maintainability

Horthy says he always tries to understand one layer below where he works. For software factories, that meant spending recent weeks studying reinforcement learning with verifiable rewards (RLVR) and coding benchmarks. In the typical setup, a model gets a small problem. Its test changes are reverted, a test patch is applied, and the run is scored on whether the tests pass. SWE-bench-style benchmarks take a real commit and issue from Django or another repository and check whether the model reproduces the fix.

In his view, reinforcement learning is what made Claude Code good. The model and harness were trained together, so the model became very good at that harness's tools for reading, writing and searching files. But that is one dimension. The cost of bad architecture cannot be measured by running unit tests, because it shows up three to six months later when the software has become hard to change. No lab has shown a benchmark that tells him whether a model makes a codebase better or worse over time, and benchmarks tend to reflect lab priorities.

The best he has seen is Cognition's Frontier Code. It checks whether tests pass, then uses one judge model to assess whether the patch is functionally equivalent to the reference answer and another to review code quality. He also mentions a "marathon"-style benchmark. He calls these better but insufficient. He is also skeptical of agentic code review. It raises the floor, but models are sycophantic. Ask a model whether code is good and it will praise it. Tell it a coworker wrote the code and it will find problems. So he finds it hard to trust models to judge code quality. His own benchmark idea is to have a model build 20 features in sequence without knowing what comes next, as a real product team would. The benchmark should be hard enough that most frontier models fail by feature six or seven.

Software factories, from 1968 to agents

Horthy says "software factory" first appeared at a NATO conference in 1968. The idea was a system of steps like a factory floor — coding, testing, validation, integration — at a time with no CI/CD and barely any version control. Companies including Toshiba adopted it. DevOps was the next wave: Chef, Ansible and Puppet replaced people resizing disks by hand or clicking through the AWS console. A disk hits 90%, Nagios fires an alert, Chef enlarges the disk. Around 2018, Nicolas Chaillan, then chief software officer of the U.S. Air Force, wrote a roughly 100-page document calling for a "DevSecOps factory" at the Department of Defense. It covered Jenkins, code-quality and security scanning, and CI/CD. The goal was to ship daily instead of every three months or once a year, with automation catching about 90% of issues.

He then describes the pre-AI factory. A work tracker such as Linear or Jira holds items in stages. People plan, pick up tickets, open PRs, get reviews, pass CI and ship. Users complain to support, crashes reach monitoring, and both feed back into the tracker for prioritization. The host notes the long latency in each stage: a bug report can take months to reach a developer and much longer to be fixed. Even at a company like Netflix or Meta, which can deploy many times a day, the work itself still takes hours or days.

In the agentic version, an agent replaces the human builder. It needs orchestration, a sandbox, an LLM, an inner harness and an outer harness. He cites Cursor's background agents as an example, since they add browsers and video recording. Building then takes about ten minutes, so code review becomes the bottleneck, and teams add AI review and agentic testing. Next, support tickets and Sentry or Datadog crashes flow straight to agents that open PRs. He calls this "the Ramp Inspect thing." Finally, faced with too much code to review, some teams turn the lights off. If users don't complain, the system is treated as working.

Horthy agrees with the host that every team is changing its factory at a different pace. His advice is to build one loop at a time and keep each small and contained. "Everything except stop reading the code is really good advice." Turning support tickets into tracker tickets, and possibly into PRs, is fine.

Three options, and seeking leverage

Horthy states his three commitments: cut through hype by trying things and talking to people who use them; protect useful words from "semantic diffusion," another Martin Fowler term, as has happened to "agents"; and understand one layer down. From there he lays out three options for teams going all in on agentic factories:

  1. Turn the lights off. Let everything flow, and hope you don't create too much slop before better models arrive and you're left with "a giant pile of ash."
  2. Slow down and read every PR and every line. He says this yields modest gains, and from his work with teams he would expect roughly a 30–50% productivity lift.
  3. Find leverage points. Identify where an hour of human planning saves four hours of rework in implementation. With the right leverage points, he says, teams can move two to three times faster while staying close to 99% faithful to what careful humans would have written.

He explains that reviewing PRs of hundreds or thousands of lines is costly. It is especially costly when the direction is wrong, because once an agent has committed to a direction, restarting beats steering. His goal is to spend an hour of human-agent discussion up front so the resulting PR takes 20 minutes to read, rather than six hours of back-and-forth.

Research, Plan, Implement — and what it got wrong

Horthy first presented RPI (Research, Plan, Implement) in August 2025. In the research step, agents read large amounts of code with parallel subagents, without being told what task was coming. They produced a markdown document explaining how the systems work and connect. Perhaps 100,000 tokens of reading became a 10K-token document. A fresh context window then produced a plan, and another implemented it.

Looking back, he thinks planning became popular in mid-2025 because it made agents work longer. Ask for "a B2B SaaS for burrito delivery" and you got a homepage. Ask for a plan first, then hand it over, and the agent kept going until the plan was done. But he now calls those plans "actually terrible." They contained every line of change as diff blocks. His team told people to read them, and read their own. Eventually he was only skimming them, so they no longer served to steer. People who reviewed both the plan and the code spent 20 minutes on each, and the two differed. That doubled the reading instead of reducing it — "anti-leverage."

The host compares RPI with spec-driven tools such as Amazon Kiro and GitHub's workflow, which he says looked good but were mostly abandoned. Horthy notes that some people call RPI spec-driven development, since for some the term just means markdown files used while coding. One OpenAI researcher proposed never reading code and treating specs as source that "compiles" into code. Horthy says that never materialized, "maybe with GPT-7." He is on a GitHub issue in Spec Kit that has been open for a year, where people keep complaining that specs and code drift apart. Two sources of truth stop being useful. His team now treats RPI documents as tactical and disposable. They are created for a task, then thrown away, and research is regenerated from scratch next time. Tokens are cheap, his time is expensive, and stale research is risky. He has seen people keep docs and code in sync but has never met anyone who found it worth the effort. He treats the code as the source of truth.

He describes the current workflow as "frequent intentional compaction," a direct application of context engineering meant to keep work in the smart zone. Research is compacted into a document; he says he doesn't read research documents, since models are good at describing how code fits together, though not at opinions or bug-hunting. The ticket and intent become a design document: current state, desired end state, and the model's design questions, which he calls a thorough, perhaps over-engineered plan mode. Planning then happens in a fresh session with both documents. Each step exists because models have a weakness at that stage. Models are weak at designing end-state architecture and program design, so a human reviews the design.

Models also favor what he calls "horizontal" plans: database, then services, then API, then frontend. In an existing codebase, that means 2,000 lines of changes across the system with nothing testable until the end. He would build vertically instead. He starts with a mock API endpoint with fake data, shapes the frontend, wires a services layer, adds the migration, then business logic, then error handling. He would rather read five small, manually verifiable diffs than 2,000 lines that don't work for reasons no one can locate.

Token harder vs. token smarter

Horthy belongs to a group chat called "hyperengineering," where members try to max out their Claude subscriptions. Some run six Claude Code accounts timed so every five-hour window is fully used. That is "token harder." In Goldratt's terms, it optimizes the utilization of one node in the factory instead of the end-to-end goal of shipping stable, lasting value. The dark factory is the same idea. It is named after fully robotic factories that need no lights because no people work there: raw materials in, cars out. Removing humans from review pushes more tokens through the system. He grants that small internal loops can safely be "dark." A review agent can send problems back to a builder agent with no human involved. The full dark factory, where no one reads code, maximizes token use. It suits people who see their job as extracting as much intelligence from the "machine god" as possible.

"Token smarter" means getting as much from AI as possible while keeping control, taste, judgment, understanding of the architecture, and years of hard-won opinions about program design. He compares it to Google's SRE book, where a small team had to manage growth from a handful of data centers to many more without growing headcount linearly. The host adds that Google never removed SREs and the team did grow. Horthy agrees and restates the point: headcount should scale like a square root or logarithm while output scales linearly. Good architecture and program design make that possible, and today that requires humans in the loop. Whatever the aim, he adds, a company with zero engineers is not a place either of them could work.

Slop in, slop out

On AI slop, Horthy repeats his line that AI can write your specs and PRDs as well as your code, but slop in means slop out if you outsource your thinking. He describes three speeds. Working closely with an agent and reading every change is faster than hand-coding, but not by much. Above that is the maximum speed at which you can still care about the code. The fastest is lights-off. HumanLayer targets the middle through a staged expansion. A two-sentence request or a voice note becomes a one-pager, which is checked. That becomes a three-pager, which is checked, then a ten-page outline, and finally around 100 pages of code. The documents need not be perfect. Each check narrows the range of possible outcomes; he calls this his physics habit of "collapsing" probabilities. He also suspects that people who like real-time strategy games will do well with AI. He cites Matt Pocock's discussion of "fog of war": estimating probabilities from partial information and seeking the information that best reveals the likely path to success.

HumanLayer and the idea of killing the pull request

Horthy describes HumanLayer, recently out of stealth, as an AI IDE, a collaboration platform, and building blocks for a software factory. It is aimed at the production end of the spectrum, where failures can cost millions, not at vibe coders building side projects. The goal is to help those engineers solve hard problems in complex codebases two to three times faster while staying near human-level quality.

Its core ideas extend RPI: start at a high level, zoom in layer by layer, and re-steer where it gives leverage. The day before the recording he had posted, "should we kill the pull request?", and he says he can't say much about it yet. His broader argument is that the IDE must be rethought for agents. Many editors started as text fields with an agent tab added on. He says that in Cursor 3 he can't even find the text editor, though he's told it exists. HumanLayer started with an IDE for directing and managing agent work, then made it collaborative using a sync engine and durable streams. The aim is for humans to give feedback on agent work in real time rather than at PR time.

He points out that strong teams have long held design reviews and sprint planning. AI can help with all of that, and using it only to write code misses much of the benefit. People who say those meetings are unnecessary because the dark factory runs everything are, in his view, giving up quality. The platform includes a Google Docs–style commenting layer where agents can show mockups, Mermaid diagrams and HTML. Everything is in the cloud and shared, in the style of Figma, so everyone sees everyone else's sessions. He compares this to Slack's advantage over email: you can see channels light up and join only the conversations you care about. He wants to move beyond discrete units like the PR and even "agile" processes that are still waterfall in practice. The open question he names is the data model for a world of agent traces, documents, tasks, projects and continuously streamed git diffs.

The host compares this to how GitHub made teams' work visible across a company. Horthy agrees that the goal is something like what GitHub did, but more continuous, real-time and collaborative than the PR. The PR, he notes, was a GitHub invention, and probably better than emailing patches, which the Linux kernel still does and which the host says works only for them. Horthy recalls learning Subversion as an undergraduate at the University of Chicago, reportedly because Subversion's creator came from there. The school switched to Git the year after he graduated.

Silicon Valley, hiring, and reading the classics

On location, Horthy points listeners to Paul Graham's talk on why San Francisco matters rather than repeating it. He says he has never felt more "seen" or connected than there. People are competent, care deeply, and share the same problems. Friends co-work in the office until 11 p.m. on side projects, which he believes needs a critical mass not found elsewhere.

HumanLayer hires for strong fundamentals: distributed systems, operating systems, core CS. His reasoning is that he can teach someone to be a good AI developer in a few months but cannot teach a CS degree in three. The problems he finds most interesting are real-time systems, cloud sandboxes and sync. He mentions admiring the ElectricSQL team's durable streams. Coding agents need to run anywhere, briefly or for a long time, on demand or on schedule, while staying part of one shared "brain." Some of the stack is boring, such as data in Postgres. He calls collaboration platforms hard distributed-systems problems, even with newer building blocks. The host compares it to cloud primitives taking about a decade to mature, from AWS to Kubernetes.

For a book, Horthy recommends Martin Fowler's Refactoring. His team spends much of its time improving the design of existing code and trying to get models to write maintainable, readable code. He adds that classics such as Clean Code and The Pragmatic Programmer are "more relevant now than [they have] ever been." That fits the unresolved problem running through the conversation: models are trained to make tests pass, and the question is how to get them to write code that is still easy to change three months later.