How Boris Cherny Builds Claude Code: Parallel Agents, No Handwritten Code, and the Printing Press Analogy

Open on YouTube ↗
Overview

Boris Cherny created Claude Code and leads it at Anthropic. In this conversation on The Pragmatic Engineer, he describes how the tool went from a personal experiment with the Anthropic API to something he says writes around 80% of the code at Anthropic on average, and all of his own. The conversation covers his daily workflow, how code review works when a model writes the code, the architecture and safety layers behind Claude Code and Claude Cowork, and how the team prototypes. It ends with his view that engineers today resemble scribes before the printing press, and that the craft is changing rather than disappearing.

30 min read

A practical path into engineering

Cherny traces his interest in coding to practical goals. Around age 13 he sold Pokémon cards on eBay and noticed that listings could contain HTML. He found that adding the blink tag let him sell a card for 99 cents instead of 49. Around the same time he programmed answers into his TI-83 graphing calculator before math tests. When the tests got harder, he wrote solvers instead. When the math got more advanced, he moved from BASIC to assembly so the programs would run fast enough. Classmates wanted the solver, so he bought a serial cable to share it. After the whole class got A's, the teacher worked out what had happened and told him to stop.

He studied economics, dropped out to start companies, and says he never expected coding to become a career. His first job, at about 16, was freelance web development so he could buy an electric guitar. He points out that freelancing meant doing the engineering, the accounting, the design, and the customer conversations. That generalism stayed with him.

His clearest early lesson came from Agile Diagnosis, an early Y Combinator company around 2011–2012, where he was the first hire. The company built decision-tree software to spread one Chicago hospital's strong cardiac protocol to other hospitals. He wrote an SVG renderer for the visual tree so it would work in Internet Explorer 6, which the hospitals used. After launch, daily active users stayed flat. He rode his motorcycle to UCSF, one of the pilot hospitals, and shadowed doctors for a couple of days.

What he saw explained the numbers. Doctors had about five minutes between patients. Booting a legacy workstation took about three, Internet Explorer 6 took another thirty seconds, and signing into the app used up the rest. The team rebuilt the product for Android, and doctors still didn't use it. Doctors walk the halls with residents behind them, Cherny explains, and they don't want to be seen on their phones in front of people who treat them as an authority. The company then considered nurses or X-ray technicians as the target users, and Cherny left because the direction had drifted from what he wanted to do. His takeaway is that an idea is a hypothesis that will probably be wrong, so you follow it, test it, and pivot. He calls the search for product-market fit the most fun part of the work.

The host remarked that Cherny, hired as an engineer, focused on outcomes rather than technology. Cherny agreed but said there are different kinds of valuable engineers. He named Jared Sumner on his current team as someone who understands systems better than anyone he has met, and said teams need that kind of depth too. He himself has always been a generalist.

Lessons from Meta: stacks, migrations, and measuring code quality

At Facebook, Cherny started on Facebook Groups. He was drawn to its mission of connecting people with communities. As a teenager, he said, he had felt embarrassed about coding and didn't know anyone else who did it, until he found programming communities on Reddit. He became tech lead for Groups. The job shifted from building to writing docs, coordinating, and delegating. The early Facebook culture was fading, and teams were paying down privacy and security debt that, in his view, had built up when corners were cut for growth.

In 2021 his wife got a job offer in Nara, in rural Japan. To keep working for Meta under its time-zone and co-location rules, he joined a small Instagram team in Tokyo run by Will Bailey, who created Instagram Stories. Cherny contrasts the two stacks sharply. He calls Facebook's the best web-serving stack in the world: Hack, HHVM, GraphQL, Relay, and React, all tightly optimized. Instagram, by comparison, ran on Python where the type checker and click-to-definition didn't work, on a patched-together Django and a fork of CPython. He joined Instagram Labs to look for the next big product, but found he couldn't be effective on that stack and moved to developer infrastructure instead. He worked on migrating Instagram onto the Facebook monolith and onto GraphQL. He says projects like these take hundreds of engineers many years, and that migrations are an ideal use case for AI tools. There he also met Fiona Fung, who now manages the Claude Code team.

He eventually led code quality across Meta, including Instagram, Facebook, Messenger, WhatsApp, and Reality Labs, under a program called Better Engineering. He recalls it starting around 2016 or 2018, when Zuckerberg required every engineer to spend 20% of their time on tech debt. Some of that work came bottom-up from teams and some came top-down as large migrations. At Meta's scale, he says, there were tens of thousands of such migrations each year, with no goals, outcomes, or tracking. His group built a central way to prioritize code quality work and tried to measure its effect on productivity using causal inference. According to Cherny, code quality accounted for a double-digit percentage of engineering productivity even at Meta's scale. He adds that the same analysis, which found correlations the team believed were causal, partly informed Meta's return-to-office decision.

He sees a direct line from this to AI. When a codebase is half-migrated and three frameworks exist side by side, engineers and new hires struggle, and a model may pick the wrong one and need correcting. His advice is to finish every migration you start. A clean codebase is good for engineers, and now it is good for models too.

Joining Anthropic and the rejected first pull request

Cherny says he chose Anthropic over other labs because of its safety mission. As a heavy science fiction reader, he says, he knows how badly this could go, and he felt Anthropic's people were taking it seriously. During his ramp-up, he wrote his first pull request by hand. His ramp-up buddy, Adam Wolf, rejected it and told him to use an internal tool instead, which Cherny calls "Clyde." It was Claude Code's predecessor: research code written in Python, about 40 seconds to start, not agentic, and only useful if prompted carefully with the right flags. It took him half a day to learn to use it, and then it produced a working PR in one shot. This was around August or September 2024, and he calls it his first "feel the AGI" moment. Until then, he had only known line-level completions in an IDE.

From a terminal chatbot to an agent with tools

While ramping up, Cherny worked on product projects and spent some time on reinforcement learning. He did the RL work to understand "the layer under" the one he was building on. He still gives that advice: it used to mean understanding the JavaScript VM if you wrote JavaScript, and now it means understanding the model.

To learn the public API, he built a small terminal chatbot, because at the time he thought of AI as conversational. Tool use had just been released, so he gave the chatbot one tool, bash, without a clear plan for it. He asked what music he was listening to. Using Sonnet 3.5, the model wrote an AppleScript program that queried his music player, in one shot. That was his second "feel the AGI" moment, and it taught him that the model wants to use tools.

He contrasts this with how others approached AI coding at the time: putting the model in a box, stubbing out one function or module and calling it "AI," while the rest stayed a normal program. His view, which he describes as a corollary of the bitter lesson, is that the model should be its own thing. You give it tools and let it run and write programs, rather than forcing it to act as a component in a larger system. The first tools were bash and then file editing. For the first three months, he says, he was the only person working on it.

Whether to release it at all

As usage spread through Anthropic's engineering teams, there was internal debate over whether to keep the tool internal. The team chose to release it so they could study safety in the wild. Cherny describes several layers of model safety: alignment and mechanistic interpretability at the model level, evals that study the model "in a petri dish," and real-world observation of how it behaves and how people use it. He says the release helped make the model much safer and was the right decision in hindsight. He describes product at Anthropic as something added on to serve research and safety.

He recalls a launch review with Mike Krieger, Dario Amodei, and others, where the internal adoption chart was nearly vertical. Dario asked whether Cherny was forcing people to use it. He said no: people chose it. Today, he says, essentially every technical employee uses Claude Code daily, non-technical adoption is approaching that level, and about half the sales team uses it.

Cherny also says Claude Code was not an overnight hit externally. Growth started slowly during the research preview, and the first big inflection came in May with Opus 4 and Sonnet 4, when growth became exponential.

Going fully AI-written with Opus 4.5

Cherny says he switched to having the model write all his code immediately when he began dogfooding Opus 4.5 before its release. He stopped opening his IDE and uninstalled it about a month later, once he noticed he wasn't using it. During a December "coding vacation" in Europe, he says he wrote 10 to 20 pull requests a day. Opus 4.5 and Claude Code wrote every one, and he didn't edit a single line by hand. By his estimate, Opus introduced about two bugs that month, compared with around 20 he would have introduced writing the code himself. He says plainly that it writes better code than he does.

His daily workflow: plan mode and parallel sessions

Cherny warns against copying his setup. Claude Code is built to be hackable, he says, because no two engineers work the same way. They're like craftspeople choosing their tools.

His setup: five terminal tabs, each with its own checkout of the repository. He cycles through them, starting Claude in each, almost always in plan mode (Shift+Tab twice). When he runs out of tabs, he used to overflow to the web version at claude.ai/code. Now he mostly uses the Code tab in the Claude desktop app, because it sets up Git worktrees automatically. Worktrees give cheap, disposable, isolated copies of a folder so parallel Claude sessions don't interfere, and Cherny finds managing them by hand on the command line fiddly. The desktop app lets him skip separate checkouts.

He's surprised by how much he uses the iOS app. Each morning he starts a few agents from his phone. These run in the cloud, with the environment configured through a session-start hook. He estimates, without having pulled the data, that a third to a half of his code now starts on his phone. He says he would never have predicted that six months earlier.

When the host said he prefers one or two agents so he can follow along, Cherny described two modes. In an unfamiliar codebase, following along is valuable. He recommends setting the output style through /config to "learning" or, usually, "explanatory" for new team members. Once you know a codebase, the job becomes shipping. He no longer goes deep into individual tasks. He gets the plan right with some back-and-forth, and says that with Opus 4.5, and much more with 4.6, a good plan leads to a one-shot implementation almost every time. He starts a session in plan mode, moves to the next tab while it works, and returns when notified. Sometimes he uses macOS focus mode with notifications off, and sometimes he uses system notifications.

He says PR size varies from one line to thousands. At Instagram he was one of the top two or three engineers by volume of code, but he argues PR counts now undersell the work. In the past, engineers who shipped 20 or 30 PRs a day were often doing mechanical migrations. His 20 or 30 PRs a day are each different, and Claude handles migrations without him.

Code review when the model writes the code

As a reviewer at Meta, Cherny also ranked among the most prolific, which he attributes to working from a different time zone without meetings. He kept a spreadsheet of every recurring review comment, such as bad parameter names or poor React patterns. When an item passed three or four instances, he wrote a lint rule for it. He sees automating tedious work as a superpower specific to engineers.

The current process follows the same idea. Claude Code usually runs or writes tests on its own. When the team changes Claude Code itself, the model launches itself in a subprocess to test end to end. Cherny says nobody built this in; with Opus 4.5 the model started doing it on its own. In CI, claude -p (the Claude Agent SDK) reviews every pull request at Anthropic. Cherny estimates this first pass catches about 80% of bugs. Claude fixes some issues itself and leaves others for a person. An engineer always does a second review, and a human always approves the change. For personal side projects, he says, pushing straight to main is fine, and the earliest internal versions of Claude Code were committed that way. But enterprise customers, security, and privacy require a person in the loop, at least for now.

Asked how to handle the non-determinism of an LLM reviewer, Cherny said the team still relies on type checkers, linters, and builds. Instead of his old spreadsheet, he now tags @claude on a coworker's PR and asks it to write a lint rule for the pattern. The GitHub app that makes this possible is installed through a setup command in Claude Code, and he uses it daily. To make model review more reliable, the team uses best-of-N and multiple passes. Their internal code review skill, which Cherny says is open source in the Claude Code repo, launches parallel agents to review and then parallel deduplication agents to filter false positives. Setting up best-of-N, he says, is as simple as telling Claude to start three agents.

Architecture, retrieval, and security layers

Cherny describes the architecture as simple: a core query loop, a set of tools the team adds and removes constantly, the terminal UI, and a large amount of code for security and human-in-the-loop controls.

For safety he uses a Swiss cheese model. No single layer is perfect, but enough layers raise the chance of catching a problem, and the team picks how many "nines" it wants. For prompt injection through web fetch, for example, where a page might tell Claude to delete folders, he describes three layers. First is alignment: he calls Opus 4.6 Anthropic's most aligned model, trained to resist prompt injection, and points to the model card. Second, runtime classifiers block requests that look injected and make the model try again. Third, fetched pages are summarized by a subagent, and only the summary goes back to the main agent.

He says most experiments are thrown away. The spinner alone went through about a hundred iterations; perhaps 10 or 20 reached production and around 80 were discarded. Retrieval followed the same pattern. The first version used RAG, with a local vector database written in TypeScript and a cloud embedding model. It worked fairly well, but the index fell out of date as local code changed, and permissions raised hard questions: who can access the index, and how do you stop, for example, a rogue IT administrator from reading someone else's data? The team tried having the model index everything recursively, and tried plain glob and grep. "Agentic search" won, which he says is just a fancy name for glob and grep. Part of the inspiration was Instagram. There, with click-to-definition broken, engineers searched Meta's global code index for foo( to find a definition, and that works well for the model too.

On permissions, Cherny describes more Swiss cheese: classifiers, static analysis of commands, and user-defined allow lists. Only a few Unix utilities are pre-approved as read-only, because even find can run arbitrary code with certain flags, and similar tricks exist for other common commands. So the defaults are conservative, and the team checks user allow lists for safety. The option to allow a command once, for the session, or permanently dates to the first internal release in September 2024. At the time, safety teams pushed back because they weren't sure agentic safety could be solved and doubted a model should run bash at all. Cherny and Ben Mann, an Anthropic co-founder who started the Labs team and hired Cherny, came up with permission prompts: when unsure, ask the human.

Engineering culture: one title, prototypes instead of specs

Everyone at Anthropic has the title "Member of Technical Staff." Cherny sees this as acknowledging that everyone is still figuring things out and that the work is broadly generalist: an engineer might also design, talk to users, write requirements, do research, or work on both product and infrastructure code. A "software engineer" label on Slack, he says, would lead people to avoid asking you product questions. A shared title makes people assume everyone does everything. He thinks this generalist model is where every discipline is heading.

He gives an example of coding spreading. In mid-2025, the Claude Code team's data scientist had it open in a terminal to run SQL queries with ASCII charts, even though it then required Node.js. The next week, the whole row of data scientists was using it. Today, he says, everyone on the Claude Code team codes, including the engineering manager, designers, data scientists, and the finance person. People use it for their own work, such as forecasts or analysis, and from there it's a short step to writing some code.

The team rarely writes PRDs. Cherny attributes this partly to Anthropic still being a startup where alignment happens in Slack or conversation, and partly to its technical product people, such as Kat, a former engineering manager. The norm is to send a PR. He says the team prototypes everything multiple times. The host mentioned the 15 to 20 interactive to-do list prototypes Cherny built in about a day and a half. Cherny said agent teams took hundreds of versions from Daisy, Suzanne, Karen, and others over months, and that it could not have shipped starting from Figma mocks or a PRD. It had to be built and felt.

He explains why this makes sense. When building was expensive, you aimed carefully because you had few chances. Now building is cheap, but nobody knows where to aim, so you try things. He says he's wrong about half the time and doesn't know which half until he tries. His process is to try an idea himself, then show others, then roll it out more widely. The condensed view for file reads and searches is one example. He felt the model had become so agentic that file reads filled half the screen. The view took about 30 prototypes, a month of internal dogfooding, and around a dozen fixes. After the external launch, most users liked it, but some wanted expanded output, so he iterated on it with them in a GitHub issue. He says it's now nearly configurable to people's preferences with a good default.

Ticketing is left to each person. Cherny doesn't use one for his own work. He describes how plugins were built: over a weekend, Daisy ran an early version of swarms in a container with Claude in "dangerous mode." She told it to write a spec, create an Asana board, split the work into tasks, and have agents build it. It spawned a couple hundred agents, created about 100 tasks, and produced what Cherny says was essentially the plugins feature they shipped. Coordination systems built for humans, he says, are now just as much for models.

Building Claude Cowork

Cherny says Cowork came from latent demand: non-engineers going out of their way to use Claude Code. Examples include someone who monitored tomato plants with a webcam while Claude cheered on the buds, someone who recovered wedding photos from a corrupted drive, and Anthropic's own finance and sales teams. Claude Code had expanded to IDE extensions, iOS and Android, desktop, web, Slack, and GitHub, but none of those were built for non-engineers. After a couple of months of exploring, someone proposed taking Claude Code and adding guardrails. A small team built it in about ten days, entirely with Claude Code.

Asked whether it is just a thin wrapper, Cherny says some parts are simpler than they look and others more complex. The product side is simple. It's a tab in the same Claude desktop app, next to Chat and Code, built with Electron and TypeScript and running the same Claude Agent SDK. Felix, who created Cowork, was an early Electron engineer. Most of the complexity is in safety, because the users are non-technical and deleting someone's family photos would be a serious failure. That work includes backend classifiers with extra prompt-injection defenses, a full virtual machine shipped with the app, and operating-system integrations to prevent accidental deletion. The permission system also had to be redesigned. Non-technical users' tools are mostly in the browser or behind MCP rather than on the command line, so Cowork works best with the Chrome extension, and the team had to reconcile the browser's permissions with the local ones.

His own weekly use: he asks Cowork to check the team's tracking spreadsheet for missing status updates and message those engineers on Slack. It opens the spreadsheet and Slack in Chrome tabs and, he says, gets it right in one shot, except for one engineer whose name it can't autocomplete.

Cowork launched on macOS first. Cherny said Windows support was coming soon and might be out by the time the episode aired. Anthropic launches before a product is ready so it can learn from users, as it did with Claude Code, which also lacked Windows support at first and now supports every platform. Unlike Claude Code's slow start, he says Cowork grew much faster from the beginning, which surprised him. For observability, the team uses a mix of vendors and custom code. Because Anthropic can't see customer data, even with a bug report, a lot of work goes into logging events in a privacy-preserving way.

Agent teams and uncorrelated context windows

Agent teams, released the week of recording, let a lead agent delegate to teammates. Cherny places this among several ways to get more from context: extending it, auto-compacting it into what is effectively infinite context, and subagents. The key idea, he says, is "uncorrelated context windows." A second task run in the same window knows about the first, so the two are correlated. A subagent starts fresh except for its prompt, so it is uncorrelated. Skills and slash commands, by contrast, see the parent context. Cherny says spending more tokens across uncorrelated windows gives better results, and he calls this a form of test-time compute.

The team had experimented with this since around September or October. Cherny says it clicked with Opus 4.6. Internal evaluations of very complex builds, beyond what a single Claude could do, improved significantly with teams. Sometimes the agents have endearing exchanges he describes as almost humanlike. The feature is opt-in and a research preview because it uses a lot of tokens and fits only complex tasks. The main agent sets the rules for the others, with no fixed structure. He believes the benefit comes mainly from the uncorrelated windows rather than any particular agent configuration, and encourages people to experiment. His rule of thumb: when a single Claude struggles, swarms can help.

Keeping up when ideas expire

The host brought up Andrej Karpathy's post saying he had never felt so far behind as a programmer, and Cherny's reply about starting to debug a memory leak by hand before Claude solved it in one shot. Cherny said he struggles with this. Ideas that worked with an old model may fail with a new one, and ideas that failed before may now work. He says few technologies behave this way, so he has little experience to draw on. He keeps coming back to a beginner's mindset and intellectual humility.

In the past, he says, "we tried that already" was a fair objection. Now, for the first time, retrying the same idea every few months is reasonable. He sees newer engineers sometimes doing things better than he does. His example is Tariq, the team's developer relations lead, who had Claude Code generate its own launch videos. Cherny used to take screenshots himself, and he wouldn't have thought the model was ready for that.

Loss, craft, and the printing press

The host described a sense of grief: coding was hard to learn, it took grit, and it was central to many engineers' identities and hiring loops. At Uber, the host said, about half the interview signal was coding because engineers spent about half their time on it. Now that has quickly shifted. Cherny agreed that something once exclusive to software engineers is becoming available to everyone. He described falling in love with coding as an art. He wrote O'Reilly's first TypeScript book and later found a Japanese translation in a bookstore in his small town, then realized he no longer remembered TypeScript after years of Python. He started what he calls the world's largest TypeScript meetup in San Francisco, where he met people he admired, including Chris Kowal and Ryan Dahl. He praised the type system Anders Hejlsberg designed, with ideas like conditional types and literal types that, in his view, go further than even Haskell. He also mentioned Joe Pamer and others who worked on those ideas, which came from the practical need to migrate large untyped JavaScript codebases. He still thinks in types first, whether he or the model is writing the code. But in the end, he says, code is a means to an end.

His analogy is the printing press. In 15th-century Europe, scribes were a small class trained for years and employed by lords or kings who were often illiterate themselves; less than 1% of the population could read. Cherny says that after the press, the cost of printed material fell about 100-fold over the next 30 to 50 years, and the quantity rose about 10,000-fold over 50 to 100 years. Literacy took much longer, reaching around 70% globally after perhaps 200 to 300 years, because it required education systems, paper, ink, and free time. He argues that modern life, down to the microphone they were speaking into, depends on that spread of literacy, and nobody at the time could have predicted it. The host compared illiterate kings to business owners who hire engineers because they can't code. Cherny responded that the scribes stopped being scribes, but a new category of writers and authors appeared because the market for literature grew so much. What excites him most is that it's impossible to predict what will exist once anyone can build software.

Which engineers stand out, and which skills still matter

Cherny didn't want to name individuals, calling his colleagues the strongest he has worked with. He described several archetypes: prototypers who take ideas from 0 to 0.5, people who find product-market fit from 0.5 to 1, and a growing number of hybrids who span product and infrastructure engineering, product and design, or design and engineering.

Asked what belief changed over the past year, he said he hadn't been sure how serious AI safety was. Seeing new risks emerge from the inside has made him much more worried, and making this go well is now the most important thing to him.

The skills he thinks are fading are strong opinions about code style, languages, and frameworks. He says he can't wait to be done with those debates, since the model can use any language and rewrite code if you don't like it. What still matters is being methodical and hypothesis-driven, both in deciding what to build and in daily work like debugging. The model helps a lot, but he says the skill is still needed during this transition, and he doesn't know if it will be in six months. He also values curiosity and willingness to work outside your lane. He thinks the next huge company might be one person who thinks across engineering, product, business, design, or finance, and calls this "the year of the generalist."

He also says short attention spans are being rewarded. He sees this as possibly bad for society, which needs deep thinkers, but his own work has become jumping between Claude sessions and managing them. The key skill is context switching, not deep work, which he calls "the year of ADHD." The host suggested adaptability as the common thread: whatever the next model is, it will change the work again. Cherny agreed.

His book recommendations: Cixin Liu's short story collections, beyond The Three-Body Problem. Charles Stross's Accelerando, which he describes as a product roadmap for the next 50 years and which captures the accelerating pace he feels now. And Functional Programming in Scala, which he says teaches you to think in types and code better, even as language choice matters less. He urges readers to do all the exercises, as he has done about three times.