How Boris Cherny Builds Claude Code: Parallel Agents, No Handwritten Code, and the Printing Press Analogy
The Pragmatic EngineerBoris Cherny created Claude Code and leads it at Anthropic. In this conversation on The Pragmatic Engineer, he describes how the tool went from a personal experiment with the Anthropic API to something he says writes around 80% of the code at Anthropic on average, and all of his own. The conversation covers his daily workflow, how code review works when a model writes the code, the architecture and safety layers behind Claude Code and Claude Cowork, and how the team prototypes. It ends with his view that engineers today resemble scribes before the printing press, and that the craft is changing rather than disappearing.
A practical path into engineering
Cherny traces his interest in coding to practical goals. Around age 13 he sold Pokémon cards on eBay and noticed that listings could contain HTML. He found that adding the blink tag let him sell a card for 99 cents instead of 49. Around the same time he programmed answers into his TI-83 graphing calculator before math tests. When the tests got harder, he wrote solvers instead. When the math got more advanced, he moved from BASIC to assembly so the programs would run fast enough. Classmates wanted the solver, so he bought a serial cable to share it. After the whole class got A's, the teacher worked out what had happened and told him to stop.
He studied economics, dropped out to start companies, and says he never expected coding to become a career. His first job, at about 16, was freelance web development so he could buy an electric guitar. He points out that freelancing meant doing the engineering, the accounting, the design, and the customer conversations. That generalism stayed with him.
His clearest early lesson came from Agile Diagnosis, an early Y Combinator company around 2011–2012, where he was the first hire. The company built decision-tree software to spread one Chicago hospital's strong cardiac protocol to other hospitals. He wrote an SVG renderer for the visual tree so it would work in Internet Explorer 6, which the hospitals used. After launch, daily active users stayed flat. He rode his motorcycle to UCSF, one of the pilot hospitals, and shadowed doctors for a couple of days.
What he saw explained the numbers. Doctors had about five minutes between patients. Booting a legacy workstation took about three, Internet Explorer 6 took another thirty seconds, and signing into the app used up the rest. The team rebuilt the product for Android, and doctors still didn't use it. Doctors walk the halls with residents behind them, Cherny explains, and they don't want to be seen on their phones in front of people who treat them as an authority. The company then considered nurses or X-ray technicians as the target users, and Cherny left because the direction had drifted from what he wanted to do. His takeaway is that an idea is a hypothesis that will probably be wrong, so you follow it, test it, and pivot. He calls the search for product-market fit the most fun part of the work.
The host remarked that Cherny, hired as an engineer, focused on outcomes rather than technology. Cherny agreed but said there are different kinds of valuable engineers. He named Jared Sumner on his current team as someone who understands systems better than anyone he has met, and said teams need that kind of depth too. He himself has always been a generalist.
Lessons from Meta: stacks, migrations, and measuring code quality
At Facebook, Cherny started on Facebook Groups. He was drawn to its mission of connecting people with communities. As a teenager, he said, he had felt embarrassed about coding and didn't know anyone else who did it, until he found programming communities on Reddit. He became tech lead for Groups. The job shifted from building to writing docs, coordinating, and delegating. The early Facebook culture was fading, and teams were paying down privacy and security debt that, in his view, had built up when corners were cut for growth.
In 2021 his wife got a job offer in Nara, in rural Japan. To keep working for Meta under its time-zone and co-location rules, he joined a small Instagram team in Tokyo run by Will Bailey, who created Instagram Stories. Cherny contrasts the two stacks sharply. He calls Facebook's the best web-serving stack in the world: Hack, HHVM, GraphQL, Relay, and React, all tightly optimized. Instagram, by comparison, ran on Python where the type checker and click-to-definition didn't work, on a patched-together Django and a fork of CPython. He joined Instagram Labs to look for the next big product, but found he couldn't be effective on that stack and moved to developer infrastructure instead. He worked on migrating Instagram onto the Facebook monolith and onto GraphQL. He says projects like these take hundreds of engineers many years, and that migrations are an ideal use case for AI tools. There he also met Fiona Fung, who now manages the Claude Code team.
He eventually led code quality across Meta, including Instagram, Facebook, Messenger, WhatsApp, and Reality Labs, under a program called Better Engineering. He recalls it starting around 2016 or 2018, when Zuckerberg required every engineer to spend 20% of their time on tech debt. Some of that work came bottom-up from teams and some came top-down as large migrations. At Meta's scale, he says, there were tens of thousands of such migrations each year, with no goals, outcomes, or tracking. His group built a central way to prioritize code quality work and tried to measure its effect on productivity using causal inference. According to Cherny, code quality accounted for a double-digit percentage of engineering productivity even at Meta's scale. He adds that the same analysis, which found correlations the team believed were causal, partly informed Meta's return-to-office decision.
He sees a direct line from this to AI. When a codebase is half-migrated and three frameworks exist side by side, engineers and new hires struggle, and a model may pick the wrong one and need correcting. His advice is to finish every migration you start. A clean codebase is good for engineers, and now it is good for models too.
Joining Anthropic and the rejected first pull request
Cherny says he chose Anthropic over other labs because of its safety mission. As a heavy science fiction reader, he says, he knows how badly this could go, and he felt Anthropic's people were taking it seriously. During his ramp-up, he wrote his first pull request by hand. His ramp-up buddy, Adam Wolf, rejected it and told him to use an internal tool instead, which Cherny calls "Clyde." It was Claude Code's predecessor: research code written in Python, about 40 seconds to start, not agentic, and only useful if prompted carefully with the right flags. It took him half a day to learn to use it, and then it produced a working PR in one shot. This was around August or September 2024, and he calls it his first "feel the AGI" moment. Until then, he had only known line-level completions in an IDE.
From a terminal chatbot to an agent with tools
While ramping up, Cherny worked on product projects and spent some time on reinforcement learning. He did the RL work to understand "the layer under" the one he was building on. He still gives that advice: it used to mean understanding the JavaScript VM if you wrote JavaScript, and now it means understanding the model.
To learn the public API, he built a small terminal chatbot, because at the time he thought of AI as conversational. Tool use had just been released, so he gave the chatbot one tool, bash, without a clear plan for it. He asked what music he was listening to. Using Sonnet 3.5, the model wrote an AppleScript program that queried his music player, in one shot. That was his second "feel the AGI" moment, and it taught him that the model wants to use tools.
He contrasts this with how others approached AI coding at the time: putting the model in a box, stubbing out one function or module and calling it "AI," while the rest stayed a normal program. His view, which he describes as a corollary of the bitter lesson, is that the model should be its own thing. You give it tools and let it run and write programs, rather than forcing it to act as a component in a larger system. The first tools were bash and then file editing. For the first three months, he says, he was the only person working on it.
Whether to release it at all
As usage spread through Anthropic's engineering teams, there was internal debate over whether to keep the tool internal. The team chose to release it so they could study safety in the wild. Cherny describes several layers of model safety: alignment and mechanistic interpretability at the model level, evals that study the model "in a petri dish," and real-world observation of how it behaves and how people use it. He says the release helped make the model much safer and was the right decision in hindsight. He describes product at Anthropic as something added on to serve research and safety.
He recalls a launch review with Mike Krieger, Dario Amodei, and others, where the internal adoption chart was nearly vertical. Dario asked whether Cherny was forcing people to use it. He said no: people chose it. Today, he says, essentially every technical employee uses Claude Code daily, non-technical adoption is approaching that level, and about half the sales team uses it.
Cherny also says Claude Code was not an overnight hit externally. Growth started slowly during the research preview, and the first big inflection came in May with Opus 4 and Sonnet 4, when growth became exponential.
Going fully AI-written with Opus 4.5
Cherny says he switched to having the model write all his code immediately when he began dogfooding Opus 4.5 before its release. He stopped opening his IDE and uninstalled it about a month later, once he noticed he wasn't using it. During a December "coding vacation" in Europe, he says he wrote 10 to 20 pull requests a day. Opus 4.5 and Claude Code wrote every one, and he didn't edit a single line by hand. By his estimate, Opus introduced about two bugs that month, compared with around 20 he would have introduced writing the code himself. He says plainly that it writes better code than he does.
His daily workflow: plan mode and parallel sessions
Cherny warns against copying his setup. Claude Code is built to be hackable, he says, because no two engineers work the same way. They're like craftspeople choosing their tools.
His setup: five terminal tabs, each with its own checkout of the repository. He cycles through them, starting Claude in each, almost always in plan mode (Shift+Tab twice). When he runs out of tabs, he used to overflow to the web version at claude.ai/code. Now he mostly uses the Code tab in the Claude desktop app, because it sets up Git worktrees automatically. Worktrees give cheap, disposable, isolated copies of a folder so parallel Claude sessions don't interfere, and Cherny finds managing them by hand on the command line fiddly. The desktop app lets him skip separate checkouts.
He's surprised by how much he uses the iOS app. Each morning he starts a few agents from his phone. These run in the cloud, with the environment configured through a session-start hook. He estimates, without having pulled the data, that a third to a half of his code now starts on his phone. He says he would never have predicted that six months earlier.
When the host said he prefers one or two agents so he can follow along, Cherny described two modes. In an unfamiliar codebase, following along is valuable. He recommends setting the output style through /config to "learning" or, usually, "explanatory" for new team members. Once you know a codebase, the job becomes shipping. He no longer goes deep into individual tasks. He gets the plan right with some back-and-forth, and says that with Opus 4.5, and much more with 4.6, a good plan leads to a one-shot implementation almost every time. He starts a session in plan mode, moves to the next tab while it works, and returns when notified. Sometimes he uses macOS focus mode with notifications off, and sometimes he uses system notifications.
He says PR size varies from one line to thousands. At Instagram he was one of the top two or three engineers by volume of code, but he argues PR counts now undersell the work. In the past, engineers who shipped 20 or 30 PRs a day were often doing mechanical migrations. His 20 or 30 PRs a day are each different, and Claude handles migrations without him.
Code review when the model writes the code
As a reviewer at Meta, Cherny also ranked among the most prolific, which he attributes to working from a different time zone without meetings. He kept a spreadsheet of every recurring review comment, such as bad parameter names or poor React patterns. When an item passed three or four instances, he wrote a lint rule for it. He sees automating tedious work as a superpower specific to engineers.
The current process follows the same idea. Claude Code usually runs or writes tests on its own. When the team changes Claude Code itself, the model launches itself in a subprocess to test end to end. Cherny says nobody built this in; with Opus 4.5 the model started doing it on its own. In CI, claude -p (the Claude Agent SDK) reviews every pull request at Anthropic. Cherny estimates this first pass catches about 80% of bugs. Claude fixes some issues itself and leaves others for a person. An engineer always does a second review, and a human always approves the change. For personal side projects, he says, pushing straight to main is fine, and the earliest internal versions of Claude Code were committed that way. But enterprise customers, security, and privacy require a person in the loop, at least for now.
Asked how to handle the non-determinism of an LLM reviewer, Cherny said the team still relies on type checkers, linters, and builds. Instead of his old spreadsheet, he now tags @claude on a coworker's PR and asks it to write a lint rule for the pattern. The GitHub app that makes this possible is installed through a setup command in Claude Code, and he uses it daily. To make model review more reliable, the team uses best-of-N and multiple passes. Their internal code review skill, which Cherny says is open source in the Claude Code repo, launches parallel agents to review and then parallel deduplication agents to filter false positives. Setting up best-of-N, he says, is as simple as telling Claude to start three agents.
Architecture, retrieval, and security layers
Cherny describes the architecture as simple: a core query loop, a set of tools the team adds and removes constantly, the terminal UI, and a large amount of code for security and human-in-the-loop controls.
For safety he uses a Swiss cheese model. No single layer is perfect, but enough layers raise the chance of catching a problem, and the team picks how many "nines" it wants. For prompt injection through web fetch, for example, where a page might tell Claude to delete folders, he describes three layers. First is alignment: he calls Opus 4.6 Anthropic's most aligned model, trained to resist prompt injection, and points to the model card. Second, runtime classifiers block requests that look injected and make the model try again. Third, fetched pages are summarized by a subagent, and only the summary goes back to the main agent.
He says most experiments are thrown away. The spinner alone went through about a hundred iterations; perhaps 10 or 20 reached production and around 80 were discarded. Retrieval followed the same pattern. The first version used RAG, with a local vector database written in TypeScript and a cloud embedding model. It worked fairly well, but the index fell out of date as local code changed, and permissions raised hard questions: who can access the index, and how do you stop, for example, a rogue IT administrator from reading someone else's data? The team tried having the model index everything recursively, and tried plain glob and grep. "Agentic search" won, which he says is just a fancy name for glob and grep. Part of the inspiration was Instagram. There, with click-to-definition broken, engineers searched Meta's global code index for foo( to find a definition, and that works well for the model too.
On permissions, Cherny describes more Swiss cheese: classifiers, static analysis of commands, and user-defined allow lists. Only a few Unix utilities are pre-approved as read-only, because even find can run arbitrary code with certain flags, and similar tricks exist for other common commands. So the defaults are conservative, and the team checks user allow lists for safety. The option to allow a command once, for the session, or permanently dates to the first internal release in September 2024. At the time, safety teams pushed back because they weren't sure agentic safety could be solved and doubted a model should run bash at all. Cherny and Ben Mann, an Anthropic co-founder who started the Labs team and hired Cherny, came up with permission prompts: when unsure, ask the human.
Engineering culture: one title, prototypes instead of specs
Everyone at Anthropic has the title "Member of Technical Staff." Cherny sees this as acknowledging that everyone is still figuring things out and that the work is broadly generalist: an engineer might also design, talk to users, write requirements, do research, or work on both product and infrastructure code. A "software engineer" label on Slack, he says, would lead people to avoid asking you product questions. A shared title makes people assume everyone does everything. He thinks this generalist model is where every discipline is heading.
He gives an example of coding spreading. In mid-2025, the Claude Code team's data scientist had it open in a terminal to run SQL queries with ASCII charts, even though it then required Node.js. The next week, the whole row of data scientists was using it. Today, he says, everyone on the Claude Code team codes, including the engineering manager, designers, data scientists, and the finance person. People use it for their own work, such as forecasts or analysis, and from there it's a short step to writing some code.
The team rarely writes PRDs. Cherny attributes this partly to Anthropic still being a startup where alignment happens in Slack or conversation, and partly to its technical product people, such as Kat, a former engineering manager. The norm is to send a PR. He says the team prototypes everything multiple times. The host mentioned the 15 to 20 interactive to-do list prototypes Cherny built in about a day and a half. Cherny said agent teams took hundreds of versions from Daisy, Suzanne, Karen, and others over months, and that it could not have shipped starting from Figma mocks or a PRD. It had to be built and felt.
He explains why this makes sense. When building was expensive, you aimed carefully because you had few chances. Now building is cheap, but nobody knows where to aim, so you try things. He says he's wrong about half the time and doesn't know which half until he tries. His process is to try an idea himself, then show others, then roll it out more widely. The condensed view for file reads and searches is one example. He felt the model had become so agentic that file reads filled half the screen. The view took about 30 prototypes, a month of internal dogfooding, and around a dozen fixes. After the external launch, most users liked it, but some wanted expanded output, so he iterated on it with them in a GitHub issue. He says it's now nearly configurable to people's preferences with a good default.
Ticketing is left to each person. Cherny doesn't use one for his own work. He describes how plugins were built: over a weekend, Daisy ran an early version of swarms in a container with Claude in "dangerous mode." She told it to write a spec, create an Asana board, split the work into tasks, and have agents build it. It spawned a couple hundred agents, created about 100 tasks, and produced what Cherny says was essentially the plugins feature they shipped. Coordination systems built for humans, he says, are now just as much for models.
Building Claude Cowork
Cherny says Cowork came from latent demand: non-engineers going out of their way to use Claude Code. Examples include someone who monitored tomato plants with a webcam while Claude cheered on the buds, someone who recovered wedding photos from a corrupted drive, and Anthropic's own finance and sales teams. Claude Code had expanded to IDE extensions, iOS and Android, desktop, web, Slack, and GitHub, but none of those were built for non-engineers. After a couple of months of exploring, someone proposed taking Claude Code and adding guardrails. A small team built it in about ten days, entirely with Claude Code.
Asked whether it is just a thin wrapper, Cherny says some parts are simpler than they look and others more complex. The product side is simple. It's a tab in the same Claude desktop app, next to Chat and Code, built with Electron and TypeScript and running the same Claude Agent SDK. Felix, who created Cowork, was an early Electron engineer. Most of the complexity is in safety, because the users are non-technical and deleting someone's family photos would be a serious failure. That work includes backend classifiers with extra prompt-injection defenses, a full virtual machine shipped with the app, and operating-system integrations to prevent accidental deletion. The permission system also had to be redesigned. Non-technical users' tools are mostly in the browser or behind MCP rather than on the command line, so Cowork works best with the Chrome extension, and the team had to reconcile the browser's permissions with the local ones.
His own weekly use: he asks Cowork to check the team's tracking spreadsheet for missing status updates and message those engineers on Slack. It opens the spreadsheet and Slack in Chrome tabs and, he says, gets it right in one shot, except for one engineer whose name it can't autocomplete.
Cowork launched on macOS first. Cherny said Windows support was coming soon and might be out by the time the episode aired. Anthropic launches before a product is ready so it can learn from users, as it did with Claude Code, which also lacked Windows support at first and now supports every platform. Unlike Claude Code's slow start, he says Cowork grew much faster from the beginning, which surprised him. For observability, the team uses a mix of vendors and custom code. Because Anthropic can't see customer data, even with a bug report, a lot of work goes into logging events in a privacy-preserving way.
Agent teams and uncorrelated context windows
Agent teams, released the week of recording, let a lead agent delegate to teammates. Cherny places this among several ways to get more from context: extending it, auto-compacting it into what is effectively infinite context, and subagents. The key idea, he says, is "uncorrelated context windows." A second task run in the same window knows about the first, so the two are correlated. A subagent starts fresh except for its prompt, so it is uncorrelated. Skills and slash commands, by contrast, see the parent context. Cherny says spending more tokens across uncorrelated windows gives better results, and he calls this a form of test-time compute.
The team had experimented with this since around September or October. Cherny says it clicked with Opus 4.6. Internal evaluations of very complex builds, beyond what a single Claude could do, improved significantly with teams. Sometimes the agents have endearing exchanges he describes as almost humanlike. The feature is opt-in and a research preview because it uses a lot of tokens and fits only complex tasks. The main agent sets the rules for the others, with no fixed structure. He believes the benefit comes mainly from the uncorrelated windows rather than any particular agent configuration, and encourages people to experiment. His rule of thumb: when a single Claude struggles, swarms can help.
Keeping up when ideas expire
The host brought up Andrej Karpathy's post saying he had never felt so far behind as a programmer, and Cherny's reply about starting to debug a memory leak by hand before Claude solved it in one shot. Cherny said he struggles with this. Ideas that worked with an old model may fail with a new one, and ideas that failed before may now work. He says few technologies behave this way, so he has little experience to draw on. He keeps coming back to a beginner's mindset and intellectual humility.
In the past, he says, "we tried that already" was a fair objection. Now, for the first time, retrying the same idea every few months is reasonable. He sees newer engineers sometimes doing things better than he does. His example is Tariq, the team's developer relations lead, who had Claude Code generate its own launch videos. Cherny used to take screenshots himself, and he wouldn't have thought the model was ready for that.
Loss, craft, and the printing press
The host described a sense of grief: coding was hard to learn, it took grit, and it was central to many engineers' identities and hiring loops. At Uber, the host said, about half the interview signal was coding because engineers spent about half their time on it. Now that has quickly shifted. Cherny agreed that something once exclusive to software engineers is becoming available to everyone. He described falling in love with coding as an art. He wrote O'Reilly's first TypeScript book and later found a Japanese translation in a bookstore in his small town, then realized he no longer remembered TypeScript after years of Python. He started what he calls the world's largest TypeScript meetup in San Francisco, where he met people he admired, including Chris Kowal and Ryan Dahl. He praised the type system Anders Hejlsberg designed, with ideas like conditional types and literal types that, in his view, go further than even Haskell. He also mentioned Joe Pamer and others who worked on those ideas, which came from the practical need to migrate large untyped JavaScript codebases. He still thinks in types first, whether he or the model is writing the code. But in the end, he says, code is a means to an end.
His analogy is the printing press. In 15th-century Europe, scribes were a small class trained for years and employed by lords or kings who were often illiterate themselves; less than 1% of the population could read. Cherny says that after the press, the cost of printed material fell about 100-fold over the next 30 to 50 years, and the quantity rose about 10,000-fold over 50 to 100 years. Literacy took much longer, reaching around 70% globally after perhaps 200 to 300 years, because it required education systems, paper, ink, and free time. He argues that modern life, down to the microphone they were speaking into, depends on that spread of literacy, and nobody at the time could have predicted it. The host compared illiterate kings to business owners who hire engineers because they can't code. Cherny responded that the scribes stopped being scribes, but a new category of writers and authors appeared because the market for literature grew so much. What excites him most is that it's impossible to predict what will exist once anyone can build software.
Which engineers stand out, and which skills still matter
Cherny didn't want to name individuals, calling his colleagues the strongest he has worked with. He described several archetypes: prototypers who take ideas from 0 to 0.5, people who find product-market fit from 0.5 to 1, and a growing number of hybrids who span product and infrastructure engineering, product and design, or design and engineering.
Asked what belief changed over the past year, he said he hadn't been sure how serious AI safety was. Seeing new risks emerge from the inside has made him much more worried, and making this go well is now the most important thing to him.
The skills he thinks are fading are strong opinions about code style, languages, and frameworks. He says he can't wait to be done with those debates, since the model can use any language and rewrite code if you don't like it. What still matters is being methodical and hypothesis-driven, both in deciding what to build and in daily work like debugging. The model helps a lot, but he says the skill is still needed during this transition, and he doesn't know if it will be in six months. He also values curiosity and willingness to work outside your lane. He thinks the next huge company might be one person who thinks across engineering, product, business, design, or finance, and calls this "the year of the generalist."
He also says short attention spans are being rewarded. He sees this as possibly bad for society, which needs deep thinkers, but his own work has become jumping between Claude sessions and managing them. The key skill is context switching, not deep work, which he calls "the year of ADHD." The host suggested adaptability as the common thread: whatever the next model is, it will change the work again. Cherny agreed.
His book recommendations: Cixin Liu's short story collections, beyond The Three-Body Problem. Charles Stross's Accelerando, which he describes as a product roadmap for the next 50 years and which captures the accelerating pace he feels now. And Functional Programming in Scala, which he says teaches you to think in types and code better, even as language choice matters less. He urges readers to do all the exercises, as he has done about three times.
You wrote the first ever TypeScript book with O'Reilly. Yeah. I found that book translated in Japanese in this little town in Japan. That was just the coolest moment. And then I realized I don't remember TypeScript at all. Now we're at the point where Claude Code writes, I think, something like 80% of the code at Anthropic on average. I wrote maybe 10, 20 pull requests every day. Opus 4.5 and Claude Code wrote 100% of every single one. I didn't edit a single line manually.
Andrej Karpathy posted that he's never felt as much behind as a programmer as he is now. This is something I really struggle with. The model is improving so quickly that the ideas that worked with the old model might not work with the new model. One metaphor I have for this moment in time is the printing press in the 1400s because there was a group of scribes that knew how to write. Some of the kings were illiterate who were employing the scribes. And if you think about what happened to the scribes, they ceased to become scribes, but now there's a category of writers and authors. These people now exist. And the reason they exist is because the market for literature just expanded a ton.
What happens when you join one of the top AI labs in the world and your first pull request gets rejected? Not because the code was bad, but because you wrote it by hand. This is exactly what happened to Boris Cherny when he joined Anthropic. Boris is the creator and engineering lead behind Claude Code. Before joining Anthropic, he spent seven years at Meta where he led code quality across Instagram, Facebook, WhatsApp, and Messenger and was one of the most prolific code authors and code reviewers at the company.
In today's episode, we cover how Claude Code went from a side project to one of the fastest growing developer tools and the internal debate at Anthropic whether to release it at all. Boris's daily workflow of shipping 20, 30 pull requests a day with zero handwritten code and how code review works when AI writes everything. Why Boris believes we're living through a time as transformative as a printing press and which engineering skills matter more now and which ones do not. If you want to understand how one of the people closest to AI coding agents actually build software today and what that means for the rest of us engineers, this episode is for you.
This episode is presented by Statsig, the unified platform for flags, analytics, experiments, and more. Check out the show notes to learn more about them and our other season sponsors, Sonar and WorkOS.
How did you get into tech, software engineering, and coding in general?
It starts a while back. I think there was kind of like two parallel paths that crossed. So, when I was maybe 13 or something like this, I started selling my old Pokémon cards on eBay. And I realized that on eBay you can actually like write HTML. And I was looking at other people's Pokémon card listings, and I realized like some of them have like big colors and fonts and stuff like this. And then I discovered the blink tag.
And I... What was the name? Was blink tag?
I put the blink tag on it, I could sell my card, you know, for like 99 cents instead of 49 cents or whatever. So, I kind of learned about HTML this way, then I got an HTML book and kind of learned about HTML.
And then the second thing was this was also, I think, sometime in middle school. We had these old TI-83 graphing calculators. And we used them for math. And what I realized is I can get a better answer on the math test if I just program the answers to the math test into my calculator. And so, I wrote these little programs. You just program the answers, and then the test got harder, so then I had to program solvers instead of the actual questions cuz I didn't know what, you know, the coefficients and stuff would be ahead of time.
And then the math got more advanced like the next year. And so, I had to drop down from BASIC to assembly to just make the program run a little bit faster. Oh, so you got in the high school you dropped down to assembly? I think this is like middle school or high school. It may be like eighth or ninth grade or something like this.
Then the thing I realized is everyone in my class was starting to realize that I had the solver, and they got kind of jealous. And so, I bought this little serial cable so I can give it to them, too. And then the next math test, everyone in the class just got A's. And the teacher was like, "What's going on?" And then eventually she realized it. It was just like, "Okay, you get away with it once, and knock it off."
But for me it was very practical. So, you know, in school I studied economics. I actually dropped out to start startups. And I never thought that coding would be a career at all. It was always very practical to me. Coding is a means to build things and to make useful things.
This startup, the first one was I think it's like my friends and I were trying to get weed. And so we started this like weed review startup. We made like a website. We called kind of different dispensaries, I think. And then we just tried to get kind of like weed samples so we could like review it for them. And it actually kind of blew up. And then I actually got more interested in... at the time no one was like testing this stuff. And so I got into kind of the like chemical testing, kind of chemical analysis. And then after this I kind of did a bunch of other startups. And then I joined YC actually pretty early. And I was the first hire of this YC startup up in Palo Alto after.
How did you decide to go to one startup after the other? Kind of vibes. Vibes, I'd say. Cuz you know, like, you know, startups, it's never a linear path. You always kind of pivot, pivot, pivot. You have to figure out what the market wants and what users want. And it's never the thing that you think. You always try things, but the idea is always a hypothesis and then almost always you have to pivot once, twice, three times.
You know, at this medical software company, this is called Agile Diagnosis. This was kind of an early YC company. This was back in maybe 2011, 2012, something like that. It was medical software for doctors. And the idea was there's these like clinical decision protocols. They vary a lot hospital to hospital. And our idea was there's one hospital in Chicago that had a really great protocol specifically for cardiac symptoms. And so we're like, wouldn't outcomes be great if every hospital in the US would use the same protocol?
And so we tried to standardize it. And we made this like decision tree software for doctors to use. And I wrote, you know, some of the software. The team was like, it was just a few of us. It was a pretty small team. And I wrote the software. It was in a web browser. And I remember this was back in the like the Internet Explorer 6 days. That's what hospitals were using.
And I wrote this like SVG renderer because it was this visual decision tree. And we launched it and then we had a DAU chart and the DAUs were flat and couldn't figure it out. And we were piloting it with a few hospitals at the time. And at the time we were based in Palo Alto, we were piloting it with, you know, a few hospitals including UCSF. And I rode a motorcycle at the time. So I rode my motorcycle up to, you know, UCSF and I shadowed doctors for a couple days just to see how do they actually use this.
And I realized that actually doctors don't have time to sit down and use a computer because you're seeing a patient, then you have maybe 5 minutes until the next patient. And in those 5 minutes you have to walk down the hall, you have to go to the computer station, you have to open up this totally legacy computer. By the time it boots up that's like 3 minutes. Then you open up Internet Explorer 6. That takes like 30 seconds. Then you have to open up this like app that we built. You have to sign in. And your 5 minutes are up. You don't even have time to use it. And so we rewrote everything to run on Android and they still weren't using it.
And the thing we realized is doctors are walking around with a bunch of residents behind them. In this kind of situation it's like a social situation, right? Like the thing that matters is they're seen as an authority. They don't want to be seen on their phones.
And then we pivoted again. So at that point we were like, okay, so maybe the doctor isn't the target user. Actually we want it to be used by maybe nurses or x-ray technicians or something like this. At that point I left because I was like, this is actually pretty far off from kind of what I wanted to do.
This is like the most fun thing for me, is finding this product market fit cuz it's always surprising. You can't have one big idea because the idea is probably going to be wrong. So, you kind of form hypotheses, you follow down, and you see what's right.
Also, I find it's so interesting how you're telling us this story, cuz I feel behind a lot of sort of success stories, we hear the success story, we hear the path of how it went. But, first of all, a lot of startups are like this, and second of all, what struck me is you were hired as a software engineer, right? And this was back before product engineers or anything was a thing, which we're now talking about. But, you just like rode your motorbike, and you went there, and you shadowed the people, and you understood how they're using it, why they're not using it, getting ideas. I feel, you know, this is what makes a great software engineer back then and even today, right? You weren't... Doesn't seem to me that you were focused on the technology, you were focused on the outcome, though.
Yeah, I mean, look, there's different kinds of engineers, and there's different ways to do it. And, you know, even on our team right now, I look at an engineer like Jarred Sumner, and he's just an incredible technical mind. He understands systems better than anyone I've met. And, you know, you need people like this. You need people with this kind of depth. For me, engineering has always been a practical thing. And, you know, for me, I've always been a generalist. And like, it doesn't matter if I'm doing, you know, like design, or, you know, if I'm doing engineering, or user research, or whatever.
The investment thesis for AI and software engineering is straightforward. As AI writes more code, more code needs to be verified. But, there's a catch. AI-generated code is, on average, harder to verify than human-written code. This is why there's Sonar, the makers of SonarQube. As a critical verification layer for the AI-enabled world, Sonar ensures that speed and volume with AI does not compromise your codebase.
Sonar's competitive position is built on 17 years of specialized expertise that no foundational model can replicate. We're talking about deep analysis engines, like symbolic execution and cross-repository data flow tracking that simulate how code actually behaves, not just what it says. To bridge the divide between AI productivity and code quality, Sonar has released the SonarQube MCP server. This tool acts as a universal translator between AI applications and the SonarQube platform. By using the model context protocol, it gives AI tools like Claude Code, GitHub Copilot, and Cursor direct access to SonarQube's analysis capabilities.
Instead of context switching, your AI agent becomes a full-fledged code review and quality assurance copilot capable of analyzing code syntax for issues, filtering bugs by severity, and even checking your project's quality gate status before you ever commit code. Whether you're working with coding assistance or scaling up with full agentic workflows, Sonar provides the automated verification that 75% of the Fortune 100 rely on. It's about giving your developers the freedom to innovate without the fear of breaking the codebase. Head to sonarsource.com/pragmatic to learn more about how Sonar enables the confidence to develop at the speed of AI. With this, let's get back to Boris's career and what he learned working at startups.
My first job I ever had, I was like, I think I was 16. And I just wanted to buy an electric guitar. And so what I did was I just started freelancing. And so I was like, okay, I guess I'll make websites. And I think Fiverr was not a thing back then, so there were some other freelancing websites. So I just started, like I put up a website, I started bidding on stuff. And my first paycheck, I just spent the entire thing on an electric guitar. But it was very practical. Right? Cuz it's like when you're in this kind of setup, you have to do the engineering, you have to do kind of the accounting, you have to do the design, you have to talk to customers. So it's just always been like that for me.
After a couple of these startups, you ended up at Facebook, now called Meta. And there you spent 7 years. Can you just talk us through what you worked on there, what you've learned there? You've also had a very remarkable career growth in terms of four promotions over 7 years. And what do you take away from that experience?
Yeah, so I started on Facebook groups. That was the first time I worked on... Vlad Kolesnikov hired me. I think he's actually still at Facebook. I think he's on some other team now. And it was cool actually. There's a big group of people that I worked with that were these kind of early JavaScript people too. And, you know, like I did a bunch of JavaScript stuff and it's funny, like I kept crossing paths with these people. And so Vlad, he worked on Bolt JS, which was the framework that powered Ads Manager, which later became React JS. I kept crossing paths with these people. And later on, yeah, later on there was a bunch more people like this.
But anyway, so I was working on Facebook groups. I was really excited about it because of this mission of connecting people to their community. This is the thing that drew me in. And at the time I was a big Reddit user. I became a Reddit user back when I was a teenager because I didn't know anyone else that coded. Even in college I didn't really know anyone that coded. And honestly, I was always kind of embarrassed about it cuz I thought it was this nerdy thing. And I thought it was kind of this thing that I knew how to do, but I wanted, you know, I wanted to be like a cool kid and, you know, like I couldn't like tell people that I coded. It was like, it was very nerdy. And at some point I discovered it was some like programming community on Reddit. And I was just shocked. Like there's other people that are into this thing. It's like such a weird hobby. It's so niche. And it was just so exciting to find like-minded people like this and get this connection. And so I just wanted to work on this. I wanted to kind of contribute to this in some way.
So I worked on Facebook groups for a while. And then, you know, there's a bunch of different projects, happy to kind of get into details for any of these. Eventually I became the tech lead for Facebook groups. And kind of grew into this and the work grew. It changed from kind of building to a lot of like doc writing and coordination and kind of delegating to others. The culture was changing at the time. So, you know, this early Facebook culture was disappearing. The docs were coming in, the, you know, alignment meetings were coming in. There was a lot more work around this kind of foundational stuff like privacy, security, things like this that I think honestly early on a lot of corners were cut in order to grow, but at some point you just have to pay that debt. And that was the time when that happened.
Then I spent a few years at Instagram after. And that was also a funny story. My wife got a job offer and she was just really excited about it and she came to me and was like, "Hey, like I got this offer, but we're going to move. Is that okay?" And I was like, "Yeah, that's fine, you know, like I work in tech. We can work remotely anywhere. Where's the job?" And she was like, "It's in Nara." And I was like, "Where's that?"
And Nara is like rural Japan. And this was a
Different time zone, yeah.
This was 12 hours or something difference or something like that. Something like that, yeah. It was like 2021.
Wow.
And then I tried to kind of find a team that would sponsor me, because there were these kind of arcane HR rules about the time zone you have to be in and the team you have to be co-located with and so on. And so there was a little kind of nascent team for Instagram in Tokyo. And Will Bailey was running the team. He was also the guy that made Instagram Stories. And so he was my manager for a while. And so we decided to grow that team together, and I worked remotely from Nara and then most of the team was in Tokyo.
And during this time I was sort of hacking on Instagram, and the stack was just insane. Like, Facebook was the single best web serving stack in the world. The way that everything is optimized, from the Hack language to the HHVM runtime to GraphQL as the transport layer to the client libraries like Relay and all the stuff, and React, it was just amazing. There's no other dev stack in the world that was this good. And it is just fully optimized. And then I went to Instagram and it's like, you know, Python where the type checker didn't work. Click to definition didn't work. And it was this kind of hacked together Django, and then like a fork of, you know, the CPython runtime. And just nothing really worked.
And so, I came to Instagram, I joined the Labs team, you know, in Japan. And the idea was to find the next big thing for Instagram. We tried some stuff, but what I very quickly realized is that I was just not effective at working on the stack because it was such a terrible stack. And so, I just went and started working on Dev Infra because we needed to fix it.
And there's a few projects that we worked on. So, one was migrating from Python to the big Facebook monolith. Another one was migrating from Instagram GraphQL. And these projects are actually in progress, you know, like these are things that involve... It takes hundreds of engineers many years to do this. It's a big code base. It's a big migration.
Now it's faster. Yeah, with these tools that we have, the AI tools, and migrations are a pretty good use case for them, though.
Yeah, it's the perfect use case for it. And then I just started getting kind of deeper into this. And by the time I left Instagram, I was working on Dev Infra and kind of leading a bunch of these migrations. That's also where I intersected with Fiona Fung, who is now the manager for the Claude Code team. I just worked with her, and she was just such an amazing leader. This incredible depth and kind of history in tech. And I just thought there's no better manager for this team.
And then I also started working on code quality. And so, the work on Instagram kind of expanded a bit. And by the time I left, I was leading code quality for all of Meta. And so, I was responsible for the quality of the code bases across Instagram, Facebook, Messenger, WhatsApp, Reality Labs, kind of all these code bases.
At Meta, it was this program called Better Engineering. And the idea was, I think it started like 2016 or 2018 or something, but Zuck mandated that every engineer at the company, 20% of their time has to be spent fixing tech debt.
Oh, interesting.
And we call this Better Engineering.
Mhm.
And some of this is kind of bottom-up where, you know, a team knows best the tech debt that they have to fix, and then some of it is top-down where you need to do, you know, very big migrations, you need to migrate to new language features, new frameworks, things like this. And at Facebook scale, you know, there are tens of thousands of these migrations every year.
And so I was sort of leading all this, and I realized very quickly that it just needed a little bit more order to it. There were no goals, no one knew kind of what the outcomes were, there wasn't any tracking. And so we developed a bunch of stuff. One of the ideas was a centralized way to prioritize the different kind of code quality efforts. The second thing was figuring out the impact of code quality on engineering productivity, which turned out to be significant.
How did you measure? What did you find there?
There was a bunch of stuff. I think some of this has been published. I don't know if all of it has, but essentially you try to do causal analysis and causal inference. This is the methodology. You try to figure out what are the factors that make it so engineers are more productive. Some of it is code quality, some of it is outside of code quality. So for example, Meta went back to, you know, return to office instead of work from home. That was partially driven by this. Because we just found some, you know, fairly strong correlations that we thought were causal.
Yep.
About this. The code quality actually contributes, you know, double-digit percent to productivity, it turns out, even at the biggest scale.
This is kind of comforting to hear, because I think it's rare to have a place where you actually measure this, but I think we feel it. Like when you have a clean code base, modular, it can get easier to work with, and I think, you know, reasoning, could it also be easier for LLMs to work with it, and my hint would be yes, it should be, right? But I think there's just very little data, but that's the feeling that I would have.
Yeah, I think a lot of the big companies have published about this. Like I think Facebook published something, Microsoft publishes a bunch about this, Google does. But yeah, totally. If every time that you build a feature, you have to think about do I use framework X or Y or Z? These are all options that you can consider because the code base is in a partially migrated state where all of these are around the code somewhere. As an engineer, you're going to have a bad time. As a new hire, you're going to have a bad time. As a model, you might just pick the wrong thing. And then, you know, the user has to course correct you. So, actually, you know, the better thing to do is just always have, you know, a clean code base. Always make sure that when you start a migration, you finish the migration. And this is great for engineers, and nowadays, it's great for models, too.
And then you joined Anthropic, and I've heard the story, which you can confirm or give more color to, that your first pull request was rejected by Adam Wolf.
He was my ramp up buddy. So, I joined Anthropic. I was trying to figure out kind of what to do next, and you know, I met a bunch of people at all the different labs, and Anthropic was just the obvious choice for me because of the mission. This is the thing that personally I know that I need the most. And also just kind of seeing all this change that's happening, it's important to have some sort of framework to think about this and to think about our role in it. I'm also a really big sci-fi reader. Like, that's definitely my genre. I'm a big reader. I have like, you know, a giant bookshelf at home and stuff. And I just know how bad this thing can go. And I just felt like this is a place that has serious thinkers. People are taking this very seriously and thinking about what can we do to make this thing go better.
So, when I joined Anthropic, I did a bunch of ramp up projects, just, you know, various stuff that I was hacking on. And I wrote my first pull request by hand because I thought that's how you write code.
That used to be how you write code.
That used to be how you write code. But even at the time at Anthropic, there was this thing called Claude, and it was the predecessor to Claude Code. It was super janky. It was Python, you know, it took like 40 seconds to start up. It was research code. It was not agentic. But if you prompt it very carefully and hold the tool just right, it can write code for you.
And so, Adam rejected my PR, and he was like, "Actually, you should use this Clyde thing for it instead." And I was like, "Okay, cool." It took me like half a day to figure out how to use this tool, because you have to pass in a bunch of flags and use it correctly. But then it spit out a working PR. It just one-shotted it.
Oh.
And this was like 2024. It's like September 2024 or August. Something like that. And I think for me this was my first feel-the-AGI moment at Anthropic, because I was just, oh my god. Like I didn't know the model could do this. I was used to these kind of tab completions, line level completions in an IDE. I had no idea that it could just make a working pull request for me.
Boris just talked about how he had a true wow moment at work using their AI model. A very different wow moment is when you use a tool at work that makes things so much easier than before. And this leads us nicely to our presenting sponsor, Statsig. Statsig offers engineering teams tooling for experimentation and feature flagging that used to require years of internal work to build. It's the kind of tool that was so complex to build that only large companies like Meta or Uber had their own custom advanced tooling for it.
Here's what Statsig looks like in practice. You ship a change behind a feature gate and roll it out gradually, say to 1% or 10% of users at first. You watch what happens. Not just did it crash, but what did it do to the metrics you care about? Conversion, retention, error rates, latency. If something looks off, you turn it off quickly. If it's trending the right way, you keep it rolling forward. And the key is that measurement is part of the workflow. You're not switching between three tools and trying to match up segments and dashboards after the fact. Feature flags, experiments, and analytics are all in one place using the same underlying user assignments and data.
This is why teams at companies like Notion, Brex, and Atlassian use Statsig. Statsig has a generous free tier to get started, and pro pricing for teams starts at $150 per month. To learn more and get a 30-day enterprise trial, go to statsig.com/pragmatic. And with this, let's get back to Boris and the origin story of Claude Code.
Yeah, and then when you joined Anthropic, we've covered this in a deep dive, but we could recap briefly on how Claude Code came to be out of what seemed like a side project or just a cool hack.
So, yeah, I started hacking on a bunch of different stuff. I was working on some things in product. I worked on reinforcement learning for a little bit just to kind of understand the layer under the layer at which I was building. This is still advice that I give to a lot of engineers: always understand the layer under. It's really important because that just gives you the depth, and you kind of have a little bit more levers to work at the layer that you actually work at. This was the advice 10 years ago. It's still the advice today. But the layer under is a little bit different now. You know, before it was like, if you're writing JavaScript, understand the JavaScript VM and frameworks and stuff.
Yeah.
Now it's like understand the model. So, I was hacking on a bunch of different stuff. Some things shipped, some things didn't ship. And at some point I just wanted to understand the public Anthropic API, because I'd never used it before. And I didn't want to build a UI. I just wanted to, you know, hack something up quite quickly, because we didn't have Claude Code back then. We were still writing code by hand. And I wrote this little bash tool that, all it did was it hit the Anthropic API, and it was essentially a chat-based application, but just in the terminal, because that's what AI used to be.
And you know, I still think about it like engineers are the first adopters. And so, when we started to move out of conversational AI to agentic AI, it took a little bit, but engineers understood it pretty quick. And I think now when you ask non-engineers about what is AI, they would say it's this conversational AI. It's a chatbot or something. And that's why I'm actually very excited for, you know, Cowork, this new product that we launched, because it's going to bring the same thing that engineers saw very early to everyone else. But when I think about, you know, Cowork, I think back to this moment that we're talking about, very early on.
Claude Code originally wasn't Claude Code. It was a chatbot. Because that's what I thought AI was. But we had to kind of figure out what is the next thing. And so at the time I built this chatbot. It was somewhat useful, but it was just a chatbot. And the next thing that I tried was I wanted it to use tools. Because tool use just came out and I didn't know what it was. And I was like, what's the experiment?
And I gave it a single tool, which was the bash tool, and I didn't know what to do with the bash tool. And so I asked it, you know, I actually didn't know if it could even do this, but I asked it, what music am I listening to? And it just wrote a little AppleScript program using like said or whatever to open up my music player and then query it to see what music it's listening to. And it just one-shotted this with Sonnet 3.5. This was actually my second feel-the-AGI moment, very quickly after the first one.
And the model just wants to use tools. That's just what I realized. Like, if you give it a tool, it will figure out how to use it to get the thing done. And I think at the time, when I think about the way that people were approaching AI and coding, everyone essentially had this mental model of you take the model and you put it in a box. And you figure out, what is the interface? How do you want to interact with this model? What do you need it to do? Essentially, it's like if you have a program, you stub out some module, stub out some function, and you say, "Okay, this is now AI." But otherwise, the rest of the program is just a program.
And so this is just not the way to think about the model. The way to think about it is the model is its own thing. You give it tools. You give it programs that it can run. You let it run programs. You let it write programs. But you don't make it a component of this larger system in this way. And I think this is a version of the bitter lesson. The bitter lesson is a very specific framing, but there's many corollaries to it. One of the corollaries is just let the model do its thing. Don't try to put it in a box. Don't try to force it to behave a particular way.
What are the first ways you saw? It was giving it tools, giving access to bash, and then later to the file system, and then to more tools, right?
That's right. Yeah, we give it bash, then... I say we, it was just me the first 3 months, but then the team grew. So, it was bash, and file edit was the second one.
And one of the interesting things we talked about last time for the deep dive is when you built it and it started to actually write code with all the tools that you had, you had an internal debate inside Anthropic: should we just keep it to ourselves? Because suddenly it spread across engineering and it was making all of you a lot more productive, right?
Yeah, that's right. In the end, the decision was to release so that we can study safety in the wild. Because when you think about safety, and you know, I keep talking about the word safety, the reason Anthropic exists as a lab is safety. This is the reason it was founded. This is the reason it exists. If you ask anyone at Anthropic why they chose it, it's because of safety. And so, if you think about model safety, you know, there's different layers at which to think about it. There's kind of alignment and mechanistic interpretability. This is at the model layer. Then there's evals, and this is kind of like it's kind of putting the
model in a petri dish and synthetically studying it in this way. And then you can study it in the wild and you can see how it actually behaves. You can see how users talk about it. You can see what are the risks in the wild, and you actually learn a lot this way. And by doing this, we've been able to make the model much safer. So, in hindsight, it was totally the right decision.
It's amusing to hear about it from your perspective because from the outside what I saw and what a lot of engineers saw was like, oh, Anthropic released Claude Code. Oh, wow. This, for the first release, it was with the Sonnet 4 release. Did it come out with Sonnet 4 originally or Sonnet 4.5? I think it was 4. It was 4. That was the general availability in February, but I think it was research preview before that.
Yeah, but when it came out, my interpretation was like, "Oh, this thing can write code pretty well." And over time it became a lot more capable. So, from our perspective, it was like this really capable coding tool that we just started to adopt and use for all sorts of increasingly productive parts. And it has become, I believe, one of the fastest-growing developer tools. And I'm always surprised to hear the story that it actually comes from research and the goal to understand how people use the model. Because on the other hand, some startups have been trying to build developer tools deliberately to get adoption. And yet this research tool is getting a lot more adoption.
I mean, this is a, you know, Anthropic, we're a research lab. We're a safety lab. And, you know, product is this kind of thing tacked onto the side. Product exists so that we can serve research better and so we can make the model safer. And this is kind of how we think about everything.
There's also this funny moment early on when we had this launch review. And we were deciding whether to launch it. I remember this moment cuz we were in the room. There was Mike Krieger, there was Dario, there were some other folks in the room, and we were deciding what should we do. We were looking at the internal adoption chart, which was just vertical. This was just insane. It was, you know, like nowadays it's 100%, right? Just 100%. Like nowadays every technical employee at Anthropic uses Claude Code every day. It's pretty much 100%. For non-technical employees, it's actually getting quite close to 100%. It's increasing very quickly. Like, you know, half of the sales team uses Claude Code. And I think that's increasing. It's just crazy. Dario had this question about how did it grow this fast? Are you forcing people to use it? And I was like, "No." We offer this tool. People vote with their feet. And, you know, we just let people use the tool that they prefer.
Yeah, you don't seem like the person who's exactly forcing people to use your tool.
Yeah. I mean, the way we did it, we just launched the thing, and then we just listened to the users, and we talked to people, we saw how they use it, we followed up, we made it better. And yeah, I mean, now we're at the point where Claude Code writes, I think, something like 80% of the code in Anthropic on average. And, you know, it writes all my code, for sure.
Yeah, and this started for you, the first time you mentioned, I think it was in November when it started to write all of your code. When did that switch come? And what happened to make you trust it to write your code, or how much do you trust it? How much do you review that code, for example?
So, the switch was instant when we started using Opus 4.5. This was before it came out, you know, we were dogfooding it for a little bit. And it was just right away. It's such a more capable model, I just found that I didn't have to open my IDE anymore. I just uninstalled my IDE, cuz I just didn't need it at that point. I actually did that like a month later, cuz I just didn't even realize that I wasn't using it anymore.
Yeah, a lot of us had similar experiences once Opus 4.5 was out in the public, and especially over the winter break. I had a similar experience. I just realized that this thing actually writes, if I'm being honest with myself, as good code as I would have written in the stack that I'm very familiar with, in my code base, my side projects where I know it, and just a lot better than what I could for a code base that I'm not as familiar with, or technologies that I'm not as familiar with.
Yeah, I'll be honest, it writes better code than I do. I don't want to go there. I still like to keep my pride, but probably true.
Yeah. I realized this cuz also in December I was traveling a little bit. I was on a coding vacation. We were talking about this before, but I went to Europe. We were just in a different time zone, kind of nomading around. And it was so fun, cuz I was just coding all day, every day, which is my favorite thing to do. And I wrote maybe, you know, like 10, 20 pull requests every day, something like that. Opus 4.5 and Claude Code wrote 100% of every single one. I didn't edit a single line manually. And I realized at the end of that month, Opus introduced maybe two bugs. Whereas if I'd written that by hand, that would have been, you know, like 20 bugs or something like that.
Can we talk about your development workflow? You have written some threads about this, which is awesome. It's on social media, on Threads and on X. But can you tell us how you use Claude Code today in terms of, you know, parallelism and tips and tricks that you and the team have kind of learned and share across the team?
Yeah, I mean, look, there's no one right way to use Claude Code. So I can share some tips and things, but I think the wrong conclusion to draw would be to just copy these and use it. The way we built Claude Code is we built it to be hackable. Because we know every engineer's workflow is different. There's no one way to do things. There's no two engineers that have the same workflow. Every engineer is different.
Workstation setup, right? Like keyboards, monitor placement, all that, everyone has it differently.
Yeah, it's like we're craftspeople, right? You choose your tools. We care deeply about it. So there's no one right way to do it.
So for me, the way that I do it generally is I have five terminal tabs. Each one of them has a checkout of the repository. So it's five parallel checkouts. And usually I'll kind of round robin and start Claude Code in each one. Almost every time I start in plan mode, so that's like shift-tab twice in the terminal.
And I also overflow as I run out of tabs. There's only so many terminal tabs. I used to use web a lot for this, so like claude.ai/code. That's the place that I overflow to. Nowadays I actually use the desktop app. It's more convenient. So Claude Code, you know, it's been in our desktop app for many months. It's just a code tab in the Claude app. And I actually really like it cuz it has built-in worktree support. So that's existed for a while. And that's quite nice for parallelism. So, you don't need multiple checkouts. You just have one, and then we automatically set up Git worktrees for you. So, you get this kind of environment isolation. The reason I do that is I actually just really hate fiddling with Git worktrees on the command line cuz it's kind of fiddly. Like, you need to know to CD and
worktree for those who are not as familiar with it? It's when you can check out and, instead of having a separate local folder, it's almost like it checks out a separate branch, right? And then you can work on it separately, but only have the conflicts at merge time.
That's right. Imagine that you have a folder, but Git makes maybe five copies of that folder in a way that's very cheap and kind of easy to throw away. So, you get this kind of isolation. You can work in parallel and the Claudes don't interfere.
Yeah, so you now have support for this, which I think you recently added, like native support, but for your workflow, you just stuck with the old one of checking out on separate folders, right?
Yeah, exactly. Over time, I'm using the desktop app more and more for this, just cuz I don't need these separate checkouts and, you know, I just have a bunch of Claudes running in parallel and I don't have to think about it. The other surprise hit is the iOS app for me. Every day, I wake up and I just start a few agents on my phone. Oh, the native one, yeah? The native one, yeah. It's the Claude app. It's the code tab in the Claude app, and it's the same exact Claude Code. Yeah, except it runs in the cloud, right? It runs in the cloud. Yeah, so you have to kind of configure the environment. Like, your environment's pretty simple, so, you know, we just use hooks for it. So, you just use the session start hook and configure it. This is kind of one of the benefits of making Claude Code really hackable, is it's very easy to do this kind of configuration.
And this is something, honestly, I would never have predicted because, you know, I code on a computer. If you told me 6 months ago I'd be writing, I don't know, I haven't pulled the data, maybe like a third, half, something like this of my code on a phone, that's crazy. But that's what I'm doing today.
And you're using parallel agents. At what point did you start using them and how has it changed your work? Cuz one thing that I noticed on myself, I don't really use that many parallel agents. Maybe like two at a time, but I'm someone who, well, I like to be in charge, and especially with Claude. Claude is a tool that you can follow along. It tells you what it's doing. You can also have, for example, learning mode, which was shipped a lot earlier, where you can actually follow along. It gives you tasks. I feel that staying in one tab and following along, the model's pretty fast as well, I can kind of keep in touch. I'm assuming at some point you must have done this, but then what happened when you changed to parallel? Do you feel you're losing any control or it doesn't really matter that much?
Yeah, I think there's kind of like two modes to think about, or two kinds of workflows to think about. So, when you're new to a code base, learning mode is awesome. Highly recommend it. For people that are onboarding to the Claude Code team, people that onboard to Anthropic, the thing that we recommend is, for people that haven't tried it, you do /config in Claude Code, you pick the output style, and you can do learning or explanatory. We usually recommend explanatory cuz that tends to be better for new code bases that you kind of haven't been in before.
For me, once you're familiar with a code base, you just want to be productive, right? You just want to ship as much as you can and you want to kind of be effective doing that. So, the role really switches. I don't really go deep into tasks anymore. I start a Claude in plan mode. I'll have it kick something off. With Opus 4.5, I think it got there. With 4.6, it just really does it. Once there's a good plan, it will one-shot the implementation almost every time. So, the most important thing is to go back and forth a little bit to get the plan right.
So, what I do is I start one. I enter plan mode. I give it a prompt. As it's chugging along, I'll go to my second tab and I'll start the second Claude, also in plan mode. Get it chugging along, then go to the third tab, go to the fourth one. Then maybe I'll go back to the first one when I get notified that it's done. And then I'll kind of keep track of that.
On, or do you turn them off?
I actually operate in both modes. Sometimes I do, you know, focus mode on the Mac, so I just have it off, but also sometimes I use the system notifications.
And you're very productive with PRs. I mean, I think it was very visible even around the holiday breaks on social media. You actually were responding to, I think someone reported a bug or a feature request, I'm not sure which one it was. And then an hour or two later, it was done cuz you did it. You've also talked about the number of pull requests you've done in a day, not to show off, but just as context. What does a pull request typically involve in terms of complexity? Are some super trivial, or are some actually larger pieces of work as well?
Yeah, pull requests, each one varies a lot. Sometimes it's a few lines, sometimes it's a few hundred or a few thousand lines. They're all just very, very different. It's changed so much. Back when I was at Instagram, I think I was one of the top two, maybe top three most productive engineers at Instagram just by volume of code written. Oh, wow. So I've always, you know, for me, I've always just coded a lot. Coding is a way that I can express myself and it's a way that my brain thinks also. And so now I just get to do it, but I think with Claude Code, the kind of code that you write, if you are very productive, the number of PRs sort of undersells what's happening.
Because I think people that used to be very productive in the old days before AI assistants, a lot of the code maybe was like code migrations or something like this. So people that shipped, you know, 20, 30 PRs every day, a lot of it was pretty, you know, like a one-liner or kind of migrating A to B or whatever. Nowadays, I ship, you know, 20, 30 PRs every day, but every PR is just completely different. Some of them are thousands of lines, some of them are hundreds, some of them are dozens, some of them are one-liners. None of these are kind of code migrations, cuz actually Claude just does those and I don't need to be part of that.
Shipping this much code, or this much to production, the obvious question that comes up for any software professional is, well, the review. The way teams used to work, and I'm not sure if Instagram did this, but a lot of other companies did this, is you make a pull request, you put it up there, there's a mandatory human reviewer. At Google, there's actually two, cuz there's one on code quality as well. How has this workflow changed? How does the Claude Code team think about code review and how has it changed over time?
Yeah, I'll start by talking about how code review used to work for me. So, the way that I used to do it is every time I, I also used to be one of the most prolific code reviewers. Oh, okay. So, both. Yeah, writers and code reviewers. That's actually one of the benefits of being in a different time zone. I'm not superhuman, I just didn't have any meetings.
And the way that I approach code review is every time that I would have to comment about something, I would drop it in a spreadsheet. And I would describe the issue. So, let's say, you know, someone named a parameter in a function badly, I would put that in a spreadsheet. If someone did some bad React pattern or something, I would put that in a spreadsheet. And then over time, I would just kind of tally up the spreadsheet, and anytime that a particular row had more than three or four instances, I would write a lint rule for it. So, just automate it with kind of static analysis. And so, that's what it used to look like.
For me, I've always tried to automate myself away because there's just so many things to do. And this is one of our superpowers as engineers, is we are able to automate all of the tedious work. There's very few other fields where you're able to do this thing. This is a thing uniquely that we're able to do. And this is a thing that I've just
always enjoyed because it gives me more free time. And I get to do the work I actually enjoy. And so, today the way this looks is a little different, but it mirrors this a little bit. So, when Claude Code writes code, it generally will run tests locally, and this is something Claude just often decides to do when it's relevant, or it'll write new tests. So you kind of do this kind of verification.
When we make changes to Claude Code, Claude will also test itself. So it'll launch itself kind of in a subprocess, it'll verify itself and it'll test itself end-to-end.
This is for your internal Claude Code implementation. So you have like this test suite so it can test itself.
Yeah, that's right. That's right. But it'll literally launch itself just in a bash process and kind of just see like, "Hey, do I still work?" So it'll do this. And this is something that we just didn't code in. Like, with Opus 4.5 especially, it just started spontaneously doing this. It just wants to kind of check. So we do this, and then we also run claude -p. So this is the Claude Agent SDK in CI. So every pull request at Anthropic is code reviewed by Claude Code. And that actually catches maybe like 80% of bugs, something like this.
And it's the first round of code review. Claude will automatically address some of these. Some of them it'll leave to a human because it's not sure what to do. There's always an engineer that does the second pass of code review. And you know, there always has to be a person in the loop approving the change.
So on the team, before anything goes into production, if you will, an engineer does look at it?
Yes.
As you think of code review, would you do this for every type of project, or is this specifically because you now know that this actually has real-world impact, people depend on it, you know, there's a lot of users? Let me put it the other way around. Can you see places where you would just not have an engineer review code? What situations would that be in? I think it depends how it's used.
Yeah, I'd agree with that. Like, you know, if you're building some personal side project, you can just YOLO straight to main, you know.
Even before AI, you would have not reviewed. You just trust yourself, or you know, you just shipped to production or SSH into production and do some changes. You get that kind of stuff, right?
Exactly. Exactly. The very first versions of Claude Code that were internal, you know, I committed straight to main. But then, you know, as soon as you have users, and you know, for Anthropic, our main customer base is enterprises, this is what we care about the most. For us, for safety reasons, security is really important, privacy is important. These are all related. It's also very important for our customers. And so, because this is an enterprise product, it has to be secure. We have to make sure that it meets a certain bar. So, we definitely use a lot of automation, but at least for now, there has to be a human in the loop just to make sure.
One thing that is just known about LLMs is they're non-deterministic. And by putting an LLM as a reviewer, Claude, doing a review, it will give good feedback, but how would you deal with the fact that you can't be sure it's always giving the feedback? You cannot be sure that even if it's capable of catching an issue, it will necessarily catch that. Are you doing anything in this loop to do deterministic things? For example, linting is very deterministic, as you will very well know. Have you thought of marrying some of these ideas, or are you using, for example, linters on the codebase, or have you found no need for it?
Yeah, absolutely, absolutely. Yeah, we have type checkers, we have linters, we run the build. Claude is actually so good at writing lint rules. So, actually, what I do now, I used to tally stuff up in a spreadsheet. Now, what I do is when a coworker puts up a pull request, and I'm like, this is lintable, I'll just be like, @Claude, please write a lint rule for this, in that PR, on their PR. And you just run, I think it's like /install-github-app or something like this. You can do this in Claude Code, and it'll install the GitHub app, which then makes it so you can tag @Claude on any pull request, any issue. I use this every single day. So, very, very useful.
So, you want these deterministic steps. Also, though, there are ways to get Claude to be a little bit more deterministic. So, for example, you can do best of N, you can have it do multiple passes. And this is actually quite easy to do. So, you know, for example, the code review skill that we use internally, it's open source, and it's available in the Claude Code repo. And so, all we do is, you know, we launch parallel agents to do stuff, and then we launch parallel deduping agents to check for false positives. But essentially, best of N, the way you implement it is all you say is, "Claude, start three agents to do this." And that's it.
Boris just talked about building that enterprise infrastructure layer. The auth, the permissions, the security, that has to all work before you can ship to real customers. This makes it a great time to speak about our season sponsor, WorkOS. If you're building any SaaS, especially an AI product, then authentication, permissions, security, and enterprise identity can quietly turn into a long-term investment. SAML edge cases, directory sync, audit logs, and all the things enterprise customers expect. It's a lot of work to build these mission-critical parts, and then some more to maintain them. But you don't have to. WorkOS provides these building blocks as infrastructure, so your team can stay focused on what actually makes your product unique. That's why companies like Anthropic, OpenAI, and Cursor already run on WorkOS. Great engineers know what not to build. If identity is one of those things for you, visit workos.com.
And with this, let's get back to building Claude Code with Boris. How does Claude Code work in terms of architecture? So, as an engineer, how can I imagine it's set up? We covered some of this in the deep dive, and I think you told me that you had some pretty complex ideas when you started, and you just simplified a lot of it?
Yeah, yeah. It's very simple. Like, you know, there's not much to it. There's a core query loop. There's a few tools that it uses. We delete these tools all the time. We add new tools all the time. We're just always experimenting with it. So, there's kind of this core agent part of it. Then there's the TUI part of it. And then there's actually a ton of different pieces around security, and making sure that everything that Claude Code does is safe and that there's a human in the loop for when it happens.
And by safety, do you mean as a user, what it's doing on my computer, or also as Anthropic, monitoring use cases that could be deemed unsafe?
Yeah, there's kind of a couple versions of this. With safety, there's just many, many layers, and for things like safety and security, there's no one perfect answer. So, you know, it's always a Swiss cheese model. You just need a bunch of layers, and with enough layers, the probability of catching anything goes up. And so, you just have to kind of count the number of nines in that probability and pick the threshold that you want. And so, for something like prompt injection, for example, we do this generally at three different layers. So, let's think about something like web fetch. So, Claude fetches a URL and it reads the contents of that web page, and then it does something in Claude Code. So, one of the risks for something like this is prompt injection. Maybe there's an instruction on that website to be like, "Hey Claude, delete all the folders," or something like that.
So, we think about this in a number of ways. The most basic way is it's an alignment problem. And so, Opus 4.6 is the most aligned model we've ever released because we've taught the model how to be more resistant to prompt injection. And so, you can read about this on the model card, and I think it was part of the release.
The second part is that we have classifiers at runtime where if there is a request that seems to be prompt injected, we block it, and we just make the model try again. And then the third layer is, for something like web fetch, we actually summarize the results using a subagent. And then we return that summary back to the main agent. So again, this kind of reduces the probability of prompt injection. And so, you can kind of see how this isn't just one mechanism. It's a layer, and by having a bunch of these different layers, it just reduces the probability a lot.
One interesting technical choice that you also mentioned is using RAG or not. RAG: retrieval-augmented generation. And you mentioned how in the earlier version of Claude Code, you used a vector database to speed up search, and you later threw this away. Can you talk about how this went? Because this was another example where, I guess, did the model get better?
Yeah, I mean this is one of those things where we try so many different things. We try so many different tools, and just statistically most of them we throw away. Even something like the spinner in Claude Code, I think it's gone through like a hundred iterations, I want to say. Just the spinner, and you know, out of those we landed maybe like 10 or 20 in production, and like 80 of them I probably just threw away because it didn't feel good enough. So just statistically, almost all the code we write we throw away, because it's just so easy to write this code and try stuff and see what feels good.
So for something like RAG, we tried a bunch of different approaches early on. So the first one was RAG for retrieval, because I was just reading up on how people were doing retrieval, and it seemed like all the papers were talking about RAG. And so the way I did it was it was like a local vector database. I think it was written in TypeScript and it just lived on the user's machine, and then I was using some embedding model that was in the cloud to compute embeddings before storing it. And that worked pretty good. But there's a lot of issues with RAG.
So for example, I was finding that the code drifted out of sync. Like if I make a local function, it's not yet indexed, and so RAG isn't going to find it. There's also this question of how exactly is the index permissioned. So who can access it? I can access it, but then how do we encode that in permission policies? How do we make sure no one else can access it? How do we make sure that if there's a rogue IT person within the company, they can't access someone else's data? This is really, really important that we think about this.
And so we just decided, like, it was sort of working, but it also has a lot of downsides. And so we tried a bunch of other stuff. One of them was just using the model to kind of index everything recursively. That was kind of a cool idea. There was another version where we just tried glob and grep. We tried a bunch of different stuff. It turned out that agentic search just outperformed everything.
And what is agentic search?
It's just a fancy word for glob and grep. That's all it is.
Nice. So the model both got good enough and you realized that it can use these tools pretty efficiently.
Yeah. And this was partially inspired, honestly, by my experience at Instagram. Because at Instagram, click-to-definition didn't work because the dev stack was just broken like half the time. And I think now it's better. And so, what engineers used to do instead is, let's say you're looking for the definition of the function foo. Instead of click-to-definition, what you would do is you would use the global index, which is quite good at Meta, and then you would search for foo opening parenthesis. And this worked pretty well. And it's funny because this works for the model pretty well, too.
Interesting how one idea from one area can come to the other. One of the more advanced parts of Claude Code that we also previously talked about is the permission system. Can you talk about what was complex about it? And also you recently open-sourced sandboxing, right?
Permissioning is really complex. Like everything else that has to do with security, it's a Swiss cheese model. There are a number of classifiers that run to make sure the command is safe, and there's also static analysis that we do to make sure the command is safe. As a user, you can also allowlist particular patterns that you know to be safe. So, for example, some standard Unix utilities we pre-allow because we know they're read-only, because we know they can't exfiltrate data or anything like this. So, we just won't prompt you for permission. But actually quite few tools fall into this category, because even something like the find command, there's actually a way to execute arbitrary code as part of that command, because there's system flags you use for this. Or even something like the sed command, there's ways to use this. So, there's just all this arcana about these various Unix utilities where it's actually not as safe as you think. And so, we want to be by default fairly conservative about what we allow by default. As a user, though, you can configure an allowlist. So, you can say, for example, these patterns are allowed, these patterns are not allowed. And so, we let you define that, and we also check this allowlist to make sure that it's safe.
Yeah, and then you have this neat permission system where every time you run a command that needs permission, you can decide to run it once, to run it for either the session or whatever makes sense, or just globally allow it going forward, right?
That's right. This is a funny artifact. This was actually in the very, very first version of Claude Code. This is the way permissions worked. This is the very first release. This was like September 2024, the first internal release. I remember at the time we weren't sure whether agentic safety could even be solved. And so, there was actually a lot of pushback internally from safety teams, because they were like, "Okay, you can't just let the model run bash commands. That's unsafe. So, what do you do? This is not a solvable problem. So, we can't launch this." I brainstormed with Ben Mann, and Ben started the Labs team. He's one of the founders at Anthropic. He's actually the person that hired me to Anthropic. We just came up with permission prompts as the way to do this. If you're not sure, just ask the human, and then they can decide.
Yeah. I want to ask you about how software engineering is done in general at Anthropic. And one of the first questions, which is, I guess, a more formal one, or from the outside, is titles, or lack of them. Everyone at Anthropic has the same title, member of technical staff. Why did this happen, and what does this result in? This kind of like everyone basically has no titles, right? Except for one.
I think it's kind of an acknowledgement that everyone just is figuring stuff out. And if you kind of squint and look at the work people are doing, it's all quite similar, and it's kind of quite generalist. And if you talk to the average software engineer, they might not just be doing coding, they might also be doing a little design. They might also be talking to users. They might be writing their own product requirements. They might be writing software and also, you know, doing research. They might be writing product code and also infrastructure code. At Anthropic there's a lot of generalists. This is also, you know, from my background, one of the reasons that I gravitated towards it. And I think member of technical staff just kind of encodes this in the way that people talk to each other, even if they don't know each other. Without this title, the
default would have been I see your name on Slack and under your name it says software engineer. And then I'm like well okay I guess you're like you're the coding person and so I'm not going to ask you like product questions. But when everyone's title is member of technical staff by default you assume everyone does everything.
And so it kind of inverts this relationship between people even if you don't know each other well. In a way it's kind of this like optimism built into the structure.
I think it's also a glimpse of the future because I think this is where software engineering is going. I think this is where every discipline is going is more of this generalist model.
Definitely it feels like it in software engineering and I heard this a funny comment by Marc Andreessen how he said that there's this Mexican standoff happening in the tech world where the designers are saying that they're actually now doing like PM and engineering work. The engineers are saying that we're doing design and like everyone thinks they're doing the work of the others and they're kind of standing there like I'm doing your work as well. But in the reality is everyone's role is expanding most of it thanks to AI because it makes easier for an engineer to do product work or for product person to engineer work and so on. So just what you've said.
I remember back in June or July of last year I walked into the office and the data sci- There's a row of data scientists that sit right next to the Claude Code team, at least at the time. And I walked in and our data scientist for the Claude Code team had Claude Code up on his monitor.
And he was using it and I was like, "This is interesting cuz you're a data scientist. Did you have like why are you using a terminal? Like you didn't have Node.js installed cuz we depended on Node.js back then. I was like, "Are you dogfooding it? Like are you just like trying to like figure out how this thing works or something?" He's like, "No, no, like I'm using it to run queries." He was just like using it to run SQL and it had like little like ASCII visualizations in the terminal. And then the next week the entire row of data scientists had Claude Code running on their computers.
And this expanded. And so if you look at the team today, on the Claude Code team, everyone codes. The engineers code, our engineering manager codes, designers code, data scientists code, our finance guy codes. Everyone on the team codes.
And I think part of it is Claude Code just make it so easy. So you don't really have to understand the code base, you can just like dive in and kind of make small changes quite easily. But I think another thing is people are able to use Claude Code to do their jobs more, whether it's, you know, financial forecasts or, you know, data science or whatever. And by doing this it's actually quite an easy crossover to just use it to write a little bit of code also. So it's just a way to dip your toe in the water.
One other interesting thing about how you work is Kat was talking about she is I guess your title is the same but people might gravitate for role a little bit more and I understand she's a little bit more on a product role. But you said that PRDs are just not really written inside Anthropic. And PRDs, product requirement document, it's a well-known artifact across big tech and increasingly over larger startups where you write a spec and the idea is that you write down your thoughts, people align, you send it over, and now you know what to build. But apparently you're not doing much of this or at all.
Some of this I think is because Anthropic is still, you know, it's still a startup. So you don't actually have to align with that many people. Usually you can just kind of talk about it or do it in Slack or whatever. But yeah, also part of it is, you know, like Kat used to be an engineering manager. She's extremely technical. And I think this is the way that, you know, our product team thinks about it, too, is you know, better just send a PR.
You're doing a lot of prototyping instead. So like that's also something where when we talked about how you were building Claude Code early on, you were showing actually you had a whole thread about the number I think you did like 15 or 20 prototypes for the to-do list and all of them interactive working. And what surprised me compared to my past tech experience, and you said that well, you did this in like a day and a half. All 20, tried it out, got a feeling for it, which incomprehensible for me. It would have taken a week or 2 weeks and people would have not done 20. They would have done three. Yeah. So like are you seeing this is there an increase in prototyping and in building and showing instead of, you know, writing things?
Yeah, absolutely. I mean on our team the culture is we don't really write stuff. We just show. It's a little hard to reflect back on the time before cuz I think now just prototyping everything is so baked into the way that we build.
It just everything is prototyped multiple times. Like you know, we launched agent teams over this week. This is our implementation of swarms. It's very exciting cuz it just lets Claude do more work for longer more autonomously. You have a bunch of different uncorrelated context windows and you have this kind of communication between agents. They can just do more. This is something that Daisy and Suzanne and other folks on the team and Karen they prototyped this for months. And they tried I all in all probably hundreds of versions of this before they got a user experience that felt really good.
It was just really really hard to get right. There's just no way we could have shipped this if we started with, you know, like static mocks in Figma or if we started with a PRD or something like this. It's a thing that you have to build and you have to feel and you have to see how it feels.
And to me one of the big takeaways even from there was like we probably should prototype more and just be more daring or just release your priors of how long it took to build a prototype or who needed to build. Back then it was always an engineer that needed to build, but it's probably not true anymore.
Yeah, that's right. I mean, we're in this world right now also where we just we don't know what the right answer is. You know, like I think back in the old way of building you the cost of building was high and so you had to actually spend a lot of effort to aim very carefully before you take your shot. Because after you take your shot, it's very hard to course correct. You can only take so few shots. But now it has changed. The cost of building is very low, but also we don't know where we're aiming. So we just have to like we have to try and we have to see what feels good.
And it's just very very exploratory. And I think also a big part of it is humility where, you know, personally I'm wrong like half the time. I'd say like most of my ideas are bad. At least half of them are bad. And I don't know which half until I try it.
Mhm. And then get feedback from others as well sometimes.
That's right. It's like I have to try it myself and then I have to see what others think cuz, you know, my intuition does not always match others.
When you were showing these prototypes of just how the tasks were built, you were telling me that you built the prototypes and then your process was always you first like looked at it, you tried it out, you got a feel for it, and then for the ones that you felt were good, you showed it to others and sometimes they give you feedback like nah, this doesn't work. And then sometimes when it felt good, then you shared it even broader. So I feel like, you know, like it's a mix, right? Where like sometimes you can decide already and then sometimes you get feedback and then eventually some good ideas come out of it.
Yeah, and there's a lot of examples of this. Like we launched this kind of condensed view for file reads and file search just cuz the model is just so agentic now, like I felt like half the screen is these like file reads and I actually don't care. Like I, you know, I read a thing. I don't really care what it is. And so, we condensed this down to make the output a little bit more readable. I really liked it after probably 30 prototypes or something like this. It took so much effort to make that feel really good and clean.
We rolled it out to employees at Anthropic for about a month, and we had everyone dogfood it, and I fixed another probably dozen bugs, dozen tweaks based on all this feedback. We launched it externally, and you know, almost all users liked it, but there were a few users that didn't because they want more expanded output. And so, on the GitHub issue, I was just going back and forth with people to be like, you know, what like what don't you like? And people give a lot of feedback. I shipped another version. Then, some people liked it, some people didn't. And so, I iterated it again, and kind of made it good. And it's actually I think almost there where people can configure it the way that they want, but still the default is really good.
But, this is just the process. You know, we get it right some of the time. We have to learn from our users. We want to hear from people so we can get it right.
Do you use ticketing systems for your work? Where you know, where you capture like, all right, here's the work. Okay, I want to or do you just pretty much do the work as it comes in?
So, at Anthropic, we leave it up to teams. On the Claude Code team, we leave it up to every person. Different people use this differently. For example, I don't use a ticketing system. Some people like to use Asana or notes or something like this.
One of the coolest things that I saw, this was maybe like 3 months ago or something, we launched plugins. And the way we launched that is Daisy for a weekend. She had a very early version of Swarms. And she let the swarm run, and she told it, "Your job is to build plugins. You have to come up with a spec, then you have to make a Asana board and split up into tasks. And then, all the different agents have to build it." And she set up a container, and she set up a Claude in dangerous mode. And she let it run for the entire weekend. It spawned a couple hundred agents. They made 100 tasks on the Asana board. And then, they implemented it. And that's pretty much the version of plugins that we shipped. These kind of coordination systems used to be for humans, but I think nowadays it's just as much for models.
Let's talk about Claude Cowork. It's one of the very important things about this it looks great. So I tried it out. It's inside Claude you have the Cowork tab there and then you can I feel it's a lot more visual way of running agents and interacting with them. One of the surprising things I heard that it was built in 10 days. Can you take us through like what it took to build it and what does actually mean? Was it from the idea or like from the decision of building it and how big was a team building it?
The team was really small. It was just a few people. For a long time we felt that there is some product to be built for non-engineers. The reason we felt this is for a long time people that were using Claude Code are non-engineers. And so you know in the product world when you see latent demand, you see people jumping through hoops to use a product that was not designed for them. That's a really good sign it's time to build another product that is built just for them.
There's all these people on Twitter that there's this one guy that was using Claude Code to like monitor his tomato plants. I just I loved this. It was like he had like a webcam set up and the Claude was like, "Oh my god, I'm so happy that our plant is budding." And because it had like a webcam and just like everyday was like monitoring it and it was so happy that the tomatoes were growing. There was someone that was using Claude Code to you know recover photos off of a corrupted hard drive and it was like his wedding photos. Wow. You know like I said our entire finance team at Anthropic uses Claude Code, our sales team uses Claude Code.
So there's all these people that are non-engineers that were using it. And at that point Claude Code is available in a lot of form factors, right? Like we started in a terminal. Then we expanded and we added support for IDEs. So we have extensions for you know every VS Code based IDE, every JetBrains based IDE. There is also iOS and Android apps. There's the desktop app. There's web. So then like Slack and GitHub apps.
So, we can expand it to all these places to make Claude Code easier for engineers. But, ultimately none of these are built still for non-engineers. And so, Claude Code evolved a lot, but it still felt like there's kind of a gap and there's a product that could make this even easier for people.
And so, for the last couple months the team was kind of hacking around and just seeing like what is the right product. And at some point someone came up with this idea of like what if we just take Claude Code, add some guardrails. So, for example, Cowork works with a virtual machine. This is one of the many ways that we make sure it's really safe. Especially for non-technical users that don't want to read like bash commands to figure out what it's doing.
And they were hacking on this, I think it was something like 10 days and 10 or something. It was just fully built with Claude Code. And then we shipped it.
And can you give us a sense of like the complexity behind an app like this and if we can walk through like what parts needed to be built because from the outside it's a little bit hard to tell like is this just a nice UI wrapper or that's you know like I don't know like a few hundred lines of code. I'm just being obviously I'm provocative here or behind the scenes it's actually really complex piece of software. And the reason I ask is like Uber is a great example where people look at the app it looks really simple. I worked there and I know it's really really complex because you don't see a lot of the complexity. There's a lot of regional things. There's a lot of back-end things that are all hidden. So, from just from looking at it Claude Cowork it's hard to tell how much of this is additional business logic that needed to be carefully thought out versus it's actually just a nice little thin wrapper on top of the model.
In some places I think there's less complexity than you would think in some places there's more complexity. So, on the product side it's quite simple cuz it's just the Claude desktop app. So, you know, you download the Claude app. It's a single desktop app. It has a tab for Cowork. It has a tab for code. It has a tab for chat. So, it is just one app and we're able to inherit a lot of that product logic. There's some UI rendering code. Under the hood, you know, it's just the same Claude Code running. It's the same Claude Agent SDK that powers Claude Code.
A lot of the complexity actually is about safety. Because we know, like I said, we know the user is non-technical, and so we just want to make sure they have a good experience.
And so, for example, if someone launches the app and then, you know, like they delete a bunch of family photos, that's really not good. And so, we wanted to make sure that we protect against this, so you can't accidentally do that. And so, that's where a lot of the guardrails came from. So, there's a bunch of classifiers running on the back end. This is for safety and again, extra mitigations for things like prompt injection and, you know, risks like this around security. On the front end, there's an entire virtual machine that we ship. There's a bunch of operating system system-level integrations to make sure people don't accidentally delete things. So, just around safety, there's a lot there.
And then, we also have to rethink the permission system because we inherit the permission system from Claude Code. But also, for Cowork, actually a big
Part of the value is not just running locally, but it's using all of your tools the way that Claude Code uses it. But the thing is, for non-technical users, your tools aren't really available as CLIs. Some of them are available over MCP. Many of them are available in a browser. And so, Cowork is really, really good when you pair it with a Chrome extension. And this is the way that I usually use it.
So, you know, for example, I use it every week to do project management for the team. We have like... We have a spreadsheet that tracks kind of at a really high level what everyone's working on. And this is kind of my personal way of project managing. You know, other people, like I said, use Asana or other people use notes or whatever. For my own tasks, I don't use anything, but kind of for the team overall, I have the spreadsheet.
And I have Cowork kind of check in. And I just ask Cowork every week, "Hey, can you look at the rows for any status that has not been filled out? Can you just ping the engineer on Slack?" And so, it'll open one tab in Chrome for the spreadsheet. It'll open another tab with Slack. And then, it'll just start messaging engineers in Slack. And it just one-shots it. There's like one engineer's name, for some reason, it can't auto-complete. But everything else it just gets.
And so this is actually like, from a safety point of view, we also thought pretty deeply about this Chrome extension and how this works and how the permissioning model should interact with this local permissioning model. So there's also a bunch of code to kind of make sure that that feels smooth.
And what's the tech stack behind this? I assume a lot of it will be similar to the Claude app, but is it Electron, TypeScript, those kind of things, or something else?
Yeah, yeah. Just Electron and TypeScript. Actually, some of the people working on it are early Electron folks. So Felix, who's, you know, the creator of Cowork, he was a really early engineer on Electron and he helped build it.
Oh, amazing. And Cowork launched macOS only. What was the reason both for choosing this platform first and for now only choosing this platform?
Yeah, so Windows coming soon. I think probably by the time this podcast comes out, we will have Windows support. We just wanted to start early and start learning. You know, like everything we do at Anthropic, it's kind of like the way that I told my own story, one of the things I like about Anthropic is it just really, really matches the way that people here think about it. You know, back to this point where like we don't have high certainty about the things that we build. And our intuition is often wrong, and so we just have to like learn from users and figure out what people actually want, and you spend a lot of time listening to people and understanding the feedback deeply. This is the way that we build a product.
And so we always launch a little bit before it's ready. We did this for Claude Code. When we launched Claude Code initially, it didn't even support Windows. Also, it didn't support, you know, like a lot of different stacks, and then over the coming weeks we added support for every stack. Now Claude Code supports every single stack. You know, like Windows, whatever weird Linux distro you use, macOS, we support everything. And so for Cowork also, we just wanted to launch early. We wanted to start with Mac as that was just the starting point. But yeah, it's going to support everything.
One thing you mentioned is getting feedback. I'm curious, both for Claude Code and for Claude Cowork, how do you go about things like observability, monitoring? When you are rolling out, do you use any feature flags? And I'm more interested in, like, did you build custom tools for this or did you decide to use certain vendors? Because especially for observability, I'm sure that this is both important, but it also sounds like pretty high scale in terms of the number of users that we can derive, or is this will not be a small operation?
Yeah, there's some off-the-shelf vendors that we use, there's some custom code that we use. So, it's actually a mix of both. There's nothing too surprising about it. There's one thing about Anthropic that's kind of interesting: because we're an enterprise company and we care a lot about privacy and security, we can't see people's data. And so, you know, like if someone reports a bug, I actually can't pull up your logs to kind of see what's going on. A lot of work goes into kind of figuring out how to log events and things like this in a privacy-preserving way. This is just very important to the way that we operate.
For Cowork, what kind of learnings have you had so far? It's been out for I think a few weeks now. Did you see something unexpected? Are you shaping the product based on feedback that you're getting?
Yeah, every day the team is landing so many fixes. The most surprising thing is just how much people are loving it, to be honest. When Claude Code first came out, it actually wasn't an overnight hit. This is something people think it was, but it was sort of a slow take-off at the beginning, and I think the first big inflection was in May when we released Opus 4 and Sonnet 4. That's when it really clicked and that's when our growth became exponential. But at the beginning, it was sort of a research preview, people didn't really know how to use it. Some people got it immediately, but most people didn't. It took a little while. For Cowork, it's a much steeper growth trajectory than Claude Code was at the beginning. So, it's just been an instant hit, and that's actually been very surprising. I didn't really expect that.
One of your new releases, which came out just very recently, it was I think yesterday or the day before when we're recording this podcast, was agent teams. And as I understand the idea with agent teams, agent swarms: instead of a single agent, you can have a lead agent and it can delegate to its different teammates. How do you start experimenting with this and how did you decide to ship it now?
We're always doing experiments, right? There's all sorts of ways to get more mileage out of Claude Code. One way you can do it is by extending context. Another way is auto-compacting context, so it's essentially infinite context, and that's what we have right now. Another way is using subagents, so you have multiple agents kind of working together. There's just like a lot of different approaches to get a little bit more mileage out of the context window.
There's this one idea called uncorrelated context windows. That's what we call it, and the idea is you have multiple context windows, but they essentially start fresh. So, they don't know about each other. And so, an example of this is, like, a correlated context window is if you have the model and it does a task and then you have it just do a second task in that same context window. And in this case the second task knows about the first one cuz it's in the same window. But for something like a subagent, it's uncorrelated because the main agent prompts the subagent, but the subagent's context window is fresh. Besides that prompt, it doesn't know what's in the parent context window.
And you can see this actually a little bit in, for example, like subagents versus skills. Because when you run a skill, you know, or slash command, it sees the parent context window, versus for a subagent, it doesn't. So, it's uncorrelated. There's some cases where you want that context. There's some cases when you don't. And there's this kind of interesting thing where uncorrelated context windows, and just throwing more context at the problem and throwing more tokens at it, when the windows are uncorrelated, gives you better results. It's actually a form of test time compute to do this.
And for something like teams, we've been experimenting with this for a while, I think since maybe like October or September or something like this. And it really just felt like with Opus 4.6, it clicked, where the model figured out really how to use this. And sometimes you see these kind of cute exchanges where the agents are talking to each other and they're like discussing something, and it's just very cool to see. It's very like humanistic in a way. But there's other times where you just get very good results. And so we had a bunch of internal evaluations, for example, where we have Claude build something very, very complex. Something more complex than what a single Claude would build. And we saw the results just really, really improve with Opus 4.6 with teams. And that's why we felt it's the right time to release it.
We also wanted to be careful. And the reason you have to opt into it, the reason it's a research preview, is it uses a ton of tokens. Cuz it's just a bunch of Claudes that are running. Not everyone wants this all the time. So it's just exciting to see how people use it and, you know, to hear the feedback. It's something you want for fairly complex tasks. You probably don't want this for every task.
The main Claude decides the rules for the sub-Claudes. We don't have a kind of a regimented way to do this. It's context-specific. I wouldn't say there's one right way to do it. I think actually a lot of the magic of this comes out of this idea of uncorrelated context windows. It's less about the specific configuration of the agents. But, you know, it's something that people should experiment with. I don't think there's a one-size-fits-all.
Have you seen use cases, even... I know it's still research, but have you seen use cases where it looks promising, this approach, this swarm approach?
Well, you know, I guess I said before plugins were fully built with swarms. There's a bunch of other features since that are built in this way. So yeah, I think for anything where you see a single Claude struggling, the swarms can help. It's interesting to look at.
Talking about change in general. With Andrej Karpathy you had a really interesting exchange back in December, when he posted that he's never felt as much behind as a programmer as he is now because of the progress with AI. And then you shared the story about how you started to debug a memory leak the old-fashioned way and then Claude just one-shot it. I think it was a reflection of how everyone is feeling that things are changing so fast, and in the holiday break I started to feel that things have really shifted. How did you, I guess, come to terms with this or start to embrace this change?
This is something I really struggle with. The model is improving so quickly that the ideas that worked with the old model might not work with a new model. The things that didn't work with the old model might work with a new model. And it's weird because there's just not a lot of other technologies like this. So I just don't really have a lot of experience to draw on to figure out how I should approach this. And it's been this new skill that I've had to learn. In a way it's like you just always have to bring this beginner mindset. Honestly, like, I'm using the word humility a lot, but you always just have to bring this kind of intellectual humility. Because just all of these ideas that were bad before are now good, and the inverse. I think that's honestly it. It's something I constantly have to remind myself about.
And back in the... it's funny, back in the old world, when someone tries an idea again and we've tried it in the past and it didn't work, usually the feedback is like, why are you doing this again? Yeah, yeah, the you should run. I mean, we used to call it a bit of gatekeeping, but it was somewhat valid, where, I know, with architecture someone came and said, like, why don't we do microservices, and someone said, we tried it and it didn't work. And if you tried it a year or two or three years ago, it was kind of valid, right, cuz not much has changed.
Yeah, that's right, that's right. And something like microservices is funny cuz it's like every 10 years it goes in and out of style. But yeah, now I think it's the first time ever where it's actually not crazy to just try the same idea every few months, because the model improves and it just works. And I actually see this with engineers on the team, like new people that are newer to the team, people that are newer to engineering, sometimes do things in a better way than I do. And I just have to like look at them and I have to learn and I have to adjust my expectations.
You know, like an example of this is, you know, when we release features, sometimes I'll like screenshot myself using them on, you know, on X or on Threads or whatever, just to kind of talk about it. But recently Tariq, our, you know, our devrel guy, he actually coded a lot. He's amazing. And he just started automating this. So, he's having like Claude Code generate its own videos for its launches, and he just started doing this. And you know, this is something like I thought would be, you know, maybe it's possible. It's not something I would have tried cuz I wouldn't have thought the model was ready, but he just did it and it just kind of worked.
One thing that I felt like just a bit odd about, and I think a lot of developers can relate, is I've come to terms with this starting from Opus 4.5, and also similar models, like I think GPT 5.2 gave me a similar vibe as well. The models have been just really good at writing code, and I realize that I don't think I will handwrite the code when I want to get stuff done. If I actually want to, you know, get the pleasure of writing it, I can still do it.
But one thing I reflected on is it's just been so much effort to get good at coding. I remember when I was learning, when I started from like kind of hacking around, to going into university, to learning C and then C++, and it was just bloody hard. And actually, you know, going through my first few jobs where I started to become better at it, I became better at debugging. And there's a point where like a lot of my identity was tied to being good at coding. That's how we used to get jobs or higher paying jobs. When I was an engineering manager, when we designed the interview loop at Uber, we had talks with managers about what we need to screen for, and we talked like, well, what do developers do most of their time? About 50% of the time they code. Therefore, we placed about 50% of the signal all about coding. So, there was a lot of things tied into coding because it is just hard. I think we all know that it takes grit. It takes some level of intelligence to get good at it.
And there's a sense of loss of like, well, I think it's great on one end that the model can do it, but it feels that something really quickly got taken away that I don't think I personally thought would happen this quickly. And I think a lot of other people are feeling that. So, some people move on a bit easier, but there's definitely this sense of grief. How did you think about it? Because again, you're an example of... you wrote so much code at Facebook, also outside of it. I know it was just a tool of doing it, but not many people could do what you did. And now the models can also work as good as you have, or if not better. That's the challenge.
Yeah, I think it's something that used to be a thing that we do as software engineers. It's becoming a thing that everyone is able to do. There was a moment, you know, like when I started coding, it was a very practical thing and it was a way to get things done. And at some point I just fell in love with the art of coding and like languages and kind of the tools themselves. And at some point I kind of fell down this rabbit hole. I wrote a book about, you know, a programming language.
TypeScript. You were the first ever TypeScript book with O'Reilly.
Yeah, yeah, yeah. That's right. It was funny actually. There was this like... There's this amazing moment for me in my little town in
Japan. I went to the bookstore and I found that book translated in Japanese. No. In this tiny town, and that was just like the coolest moment. And then I actually realized I don't remember TypeScript at all, cuz I was only writing Python for a couple years at that point.
Yeah, and then at some point I started the biggest TypeScript meetup in the world. That was in SF, and I got to meet kind of a lot of my heroes. There was Kris Kowal, who wrote General Theory of Reactivity. There was Ryan Dahl, the guy that made Node. One of the first times that I went really deep into this community and just the language itself and the tools themselves.
And for something like TypeScript, there's this beauty in the types, in the type system, cuz Hejlsberg is just brilliant. Like the idea of conditional types, and just like anything can be a literal type. And there's these very deep ideas that even the most hardcore functional languages do not have. Like even in something like Haskell, it doesn't go this far, and Anders just took it and he pushed it much further than it had been pushed. And, you know, Joe Pamer and a bunch of other folks kind of explored a lot of these ideas and thought of this. And I think for them it was also very practical, right? Because they had these large untyped JavaScript code bases. How do you gradually migrate to something typed? And you have to come up with these very beautiful ideas to do this. For me, Scala was another kind of rabbit hole that I fell into, and kind of this functional programming world, and still when I write code and when the model writes code, I always think in the types first. That's what matters, is what is the type signature. That matters more than the code itself, and getting that right.
So, there is this beauty to it. There's an art to it, for sure. But in the end, it's a practical thing. And in the end, this is a thing that we use to build things and, you know, it's a means. It's a means to an end. It's not an end to itself. I think one metaphor I have for this moment in time that we're in is the printing press in, you know, the 1400s or whatever.
Mhm. Because at that moment it was actually quite similar, right? There was a group of scribes that, you know, knew how to write. And it was, as I understand, of course we never lived it, but as I imagine it was a hard process to learn. You need to get the equipment. You probably needed some sponsorship or being selected. Practicing, because you needed to produce the same thing over and over again, and few people could do that, and I assume it was either high prestige or highly paid or who knows. Let's assume it was. But then the printing press came along.
Yeah. Yeah, and at least in Europe, you had to, like, a lord or a king or something had to employ you, and then you had to go through, you know, years of training. And there was this class of scribes that knew how to write. They were employed by someone like this. Often the king themselves or, you know, the queen was not literate. So, it was this very, very niche skill, and less than 1% of the population was literate in Europe, you know, back then.
And then the printing press came out, and what happened? So, the cost of printed material went down something like 100x over the next, I think, 30 years, 50 years, or something. The quantity of printed materials went up like 10,000x in the next 50-100 years. This was the first effect. Literacy, it took a little while for it to catch up. So, I think global literacy went up to something like 70%, but that took another 200 years, 300 years, because learning to read is just very hard. Learning to write is hard. It takes a lot of effort. It takes an education system. It takes, you know, infrastructure to have paper and ink and the free time to do this instead of working on a farm. So, it took the early stage of industrialization to actually get there.
But I think this effect of making it so this thing that was locked away in an ivory tower, and now it's accessible to everyone, this is just, you know, none of the things around us would exist today without this. If we weren't literate, if the people that built, you know, this microphone weren't literate, it would have just been very hard to have a modern economy. None of these things would exist.
And I just kind of think about back then, if people had to predict what would happen when the printing press came out, no one would have predicted that the microphone would become a thing. So, I just feel like this is the best analog for the moment that we're in right now.
And it's interesting that you say that some of the kings were illiterate who were employing the scribes, because if we're being honest with ourselves, we have business owners who know what they want to build and they're employing software engineers because they themselves cannot write code. And I think we like to mock the CEOs who are coming to the scene. They might even have a drawn prototype or whiteboard and saying this should be easy, but of course they don't understand how difficult it is. But there seems to be a bit of an analogy where there's a person who wants what they want, but until now they needed to hire a specialist who can build that, and there's always that disconnect between the idea and the person. And just like with the printing press, what would happen if they could actually express them? Like the king could actually read or write their own letters. They wouldn't need that middle man, and things become more efficient. I mean, of course for the scribe it's not the best news necessarily, but I mean the smart scribes can also do... yes, so someone needs to write the books, run the press, etc.
Yeah, exactly. And if you think about what happened to the scribes, right? They ceased to be scribes, but now there's a category of writers and authors. These people now exist. And the reason they exist is because the market for literature just expanded a ton.
And I guess also if we think about back then, a scribe's work was read by a few people, and with the printing press, there's a lot more authors, and some of them are not really read, but some of them have wider reach than they could imagine. There's new careers that exist because of that.
Yeah, I love the analogy. And the most exciting thing for me is it's just so impossible to say today what will happen after this happens and after this transition happens. Just, you know, the economy as we know it would not have existed without it. So what's next? What's the thing that we can't even predict today that will exist because anyone can do this?
Well, we cannot predict, but I think we can look at what is working right now. If you look around in your environment, may that be the team across Anthropic, who are software engineers or builders or members of technical staff, however we call them, who to you are standout? What are they doing? What skills have they built up, and how have they changed the way they work?
It's hard to name individuals because honestly, these are the strongest people I've ever worked with in my career. There's all sorts of different archetypes. There are some people that are really amazing prototypers. So, take something from 0 to 0.5. Just, you know, figure out what are some cool ideas, what is the technology unlock. There's other people that are amazing at finding product market fit. So, kind of 0.5 to 1, or maybe 0 to 1. There's other people that span different disciplines, and I'm just seeing more and more of these people. Like I said, people that span product engineering and infrastructure engineering, or, you know, product and design, or design and engineering. I think I'm just seeing a lot more of these hybrids.
What's a belief that changed from last year to this year? Something that, you know, you either believed, or a conviction that you had, that you've either revised or completely threw away?
I think one thing I wasn't sure about is how big a problem is safety, to be totally honest. I joined Anthropic because, like I said, I read a lot of sci-fi and I know how bad this thing can go if it goes bad. It wasn't something I was sure about. But seeing it from the inside, and then seeing the new risks that have arisen in the last year, it just makes me much, much more worried about it. So, I think it was kind of an important thing for me. Now, it's just the most important thing for me: how do we make sure this thing goes well?
I think it's safe to say you were a really great software engineer even before all the AI things started, and you seem to be a very productive engineer, of course part of a team as well, but also individually. What are some skills of, you know, before, being a software engineer, that are still as valuable or maybe even more valuable than before, and what are ones that are maybe just not as much and they're best left behind?
Probably... Okay, so the stuff that's best left behind is maybe very strong opinions about code style and languages and things like this. I can't wait to get past these endless language debates and framework debates and all this stuff. Because the model can just, you know, use whatever language and framework, and if you don't like it, it can just rewrite it for you. So it just doesn't matter anymore.
I think something that still matters a lot today is being methodical and hypothesis-driven. This matters both in product design, in this world where everything is being disrupted and we need to figure out what to build next, and this is something everyone is thinking about. But it also matters for engineering day-to-day, you know, something like debugging. You just have to be very methodical about it. And the model can do this and it can help a lot. But I think still we're in this transition point where you still need to have the skill. I don't know if you're still going to need to have it in 6 months.
Other skills that I think are more valuable are being curious and being open to doing things beyond your swim lane. So, you know, if you're working on engineering, but you really understand the business side, you can just build really awesome products. And I think the next, you know, billion-dollar product, you know, like after Claude Code, whatever the next startup is that, you know, becomes the next trillion-dollar startup, it might just be one person that has some cool idea and their brain just is able to think across, you know, engineering and product and business, or, you know, design and finance and something else. People are going to become more and more multi-disciplined, and this will become more and more rewarded. So, in some ways I think this will be the year of the generalist.
I think the other skill that's actually been rewarded is having a short attention span. I've seen rewarded now. Oh, yeah. It's, you know, like teenagers are using, you know, TikTok and all this stuff, and I think in some ways it's kind of dangerous for society, because you want people that can think deeply and can contemplate ideas and aren't just moving on to the next idea very quick. But in some ways I think this year is kind of the year that is going to reward it. It's like the year of ADHD. Because the work for me has become jumping between Claudes. It's become managing Claudes. And so it's not so much about deep work. It's about how good am I at context switching and, you know, jumping across multiple different contexts very quickly.
Could I add that, from what all you said, maybe you could add one thing, which is adaptability, because you're saying of course that ADHD and you can jump across, but of course earlier you were very good at focusing deeply on one thing as well. And what strikes me about you, and maybe this is true for other people as well, you're just kind of very open to adapting your working style and seeing what works well for this stage, especially when things are changing. I think the one certain thing we can be sure of is whatever the next model that comes out, it will change again, and you need to be curious and open to adapting how you work, right?
Yeah. And as closing, what's a book or books that you would recommend?
I've gone down a Cixin Liu rabbit hole. So, he's the Three-Body Problem guy, but he actually has a lot of other really good books. I really love his short stories. He has a couple books of short stories. I'm a big fan. For people that are new to sci-fi and you want a little bit harder sci-fi, I really love Accelerando by Stross. This is a book I would totally recommend. It's essentially the product roadmap for the next 50 years, with takeoff kind of starting to happen and kind of AI singularity. And then it ends up with this kind of group lobster consciousnesses orbiting Jupiter. And it's just amazing, and the thing that I think it really captures is just the pace, this quickening, quickening, quickening pace of how this feels. It really matches the feeling right now.
And then on the technical side, I would strongly recommend Functional Programming in Scala. Even if language choice just doesn't matter as much anymore, I think there is this art to functional programming that just teaches you how to code better. And it'll just teach you how to think in types. If you read this book, I think what's really important is to do the exercises also, and I've gone through and I've done all of them probably like three times over, and it's just amazing. It really just knocks this idea of functional types into your head, and it's just a thing you can't stop thinking about.
Boris, thanks so much. This was awesome. Yeah, thanks, Greg.
This was a really interesting conversation, and the thing that I keep coming back to is Boris's printing press analogy. The idea that medieval scribes were this tiny elite who could write, employed by kings who themselves were often illiterate, and that we software engineers might be in a similar position today. We are the scribes. We spent years mastering this craft, and now the printing press is arriving. But what Boris told me is that the scribes did not disappear. They became writers and authors, and the entire market for written work expanded beyond anything anyone could have predicted. I do find this hopeful and also appreciate that Boris didn't sugarcoat it.
The other thing that stuck with me is just how differently the Claude Code team built software. No PRDs, no mandatory ticketing system, designers and data scientists and finance people all writing code and building dozens or hundreds of prototypes before shipping a feature. And Boris is shipping 20 to 30 pull requests a day without editing a single line by hand. And there are different verification systems in place: Claude Code reviewing its code, automated lint rules, best-of-N passes, and human code review.
If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks, and see you on the next one.
Article published
