Building Codex: Tibo Sottiaux on Rust, Open Source, Harnesses, and Cheap Code Changes
The Pragmatic EngineerTibo Sottiaux was on the team when Codex started and has led the broader Codex team since. On the Pragmatic Engineer Podcast, the host asked how OpenAI's coding agent was built and how it keeps changing. They covered why the CLI is written in Rust and released as open source, why it works with other companies' models, how the harness and the model share the work, and what happens to code review, maintenance, and architecture when changing code becomes cheap. Throughout, Sottiaux argues that much of the scaffolding engineers build around models is temporary, and that the durable skills are clear intent, good abstractions, and staying close to users.
From a Belgian village to applied mathematics
Sottiaux traces his interest in computers to his parents' decision to leave Brussels for a village of about 200 people. At around eight years old, with few people nearby he wanted to spend time with, he turned to computers and the early internet to learn about things. He credits the move with leaving him "no choice" but to get interested in computers.
He studied applied mathematics, finished early, and ran small consulting work for banks and on supply chain problems while still a student. He says the appeal of applied mathematics was taking sophisticated theory and using it to improve the world, rather than doing mathematics for its own sake. After university he co-founded a startup that optimized pharmaceutical supply chains for clinical trials. It decided how much medicine to produce and where to dispatch it to reduce waste. The company used non-ML techniques: Monte Carlo simulation and stochastic multi-stage optimization. It later applied the same methods to the steel industry and the European electrical grid. He says the company still exists and is now changing a lot because of modern AI.
Google: a canceled project, Maps, and DeepMind
Sottiaux joined Google in London in 2015. His first project was not Maps. It was an effort to make websites faster on mobile, run by a small group inside the ads organization. The goal was to offset ad revenue lost as traffic moved from desktop to mobile. He worked on it for about two years and calls it the most fun he has had solving hard technical problems. Then a VP flew in from California and canceled it: the product had only hundreds of users, which was "clearly not Google scale."
He says people on the team were surprised, and that the surprise was itself the lesson. The project lacked product-market fit, the right users, and the right feedback loop, and the team had trusted a product manager who said things were going well when they were not. The lesson he took was to keep questioning, both about his own impact and about whether the overall project matters.
He then worked on reviews in Google Maps for about a year before moving to DeepMind. He describes DeepMind then as a special place, with early rumblings around AlphaGo and a focus on the hardest problems. Given his background, he was drawn to it. There he worked on research infrastructure and tooling, a theme he says he carried for almost a decade: building tools and products that make other people more efficient.
A ChatGPT-like bot inside DeepMind, a year early
The host asked about something Sottiaux had posted on X: he had helped build an internal Google chatbot roughly a year before ChatGPT. He explains that DeepMind's main focus was grand challenges, games, and reinforcement learning outside the language setting. Google Brain, separate at the time, had its own language model work. One group inside DeepMind pushed the hypothesis that scaling language models on large text corpora might be enough to reach general intelligence, which he says was hotly debated.
Because Sottiaux was building research tooling, the question arose of how to present these models to researchers so they could inspect inputs and outputs. That naturally ended up as a chat interface. The early models were "a little bit absurd," not very coherent or useful, but fun to tinker with. The application spread "like wildfire" inside DeepMind as people shared conversations, and it began to feel like more than a tool for researchers. There was a desire to launch it externally, but he says DeepMind was not set up for that. Google has a blessed production stack and a well-optimized process for launching products, which he credits for doing things well but describes as a very hard environment for real innovation.
Why he left for OpenAI
Sottiaux says he was comfortable at Google. What pulled him away was the chance to join a mission he believed in, with people who wanted direct impact, instead of a setup where research happens in one place and making it useful is someone else's job. He wanted research and product to be co-designed. When he met people from OpenAI and learned that only about 20 people worked on ChatGPT, he was struck by how empowering that must be at that scale.
He joined just before the reasoning launch. About a month after he arrived, OpenAI released o1-preview, which he calls exhilarating. He says the values he took from that period, which he later applied to Codex, were moving fast, caring about impact, having a community, and running an intense feedback loop with it.
How Codex began
Sottiaux joined OpenAI in 2024 and started, as before, by building research infrastructure. With o1-preview and later models, he says it became clear that OpenAI had to use its own models to go faster. He became focused on what was limiting that. With colleagues in research, he began training models and building small agents that he calls the true precursors of Codex. Those internal models were trained to be very good at OpenAI's Python codebase and to have good taste in architecture and code style. They were Python-only, and the aim was to build infrastructure quickly and help researchers code faster.
According to Sottiaux, Greg and Sam were very supportive, and Greg insisted the work should benefit the world and not only OpenAI, which meant turning it into a product. At that point the research effort merged with the internal "A-SWE" (autonomous software engineer) project. A sprint followed that produced the first cloud version of Codex. Sottiaux says it did not have product-market fit because it involved too much friction. The Codex CLI launched as well, and the team kept iterating on how to make models genuinely helpful.
Why the core agent is written in Rust
The host pointed out an odd choice: Codex's core is written in Rust, although models at the time were stronger in Python and TypeScript, and most competing harnesses used those languages. Sottiaux says the decision followed from treating the product interface and the agent as separate things from the start. The agent core had to be robust, secure, and built for efficiency and scale. From earlier projects that grew from "fun thing" to "scale of the largest data center," he had learned that early decisions matter a great deal, as long as they don't cost too much velocity.
In his account, the team had very strong Rust developers, OpenAI's internal models were "not bad at Rust," and Rust's compile-time static verification is useful for agents too. So it quickly became clear to them that Rust would work well for agents if they invested in it. The main drivers were correctness and efficiency.
He also says Codex could probably have succeeded in TypeScript or even Python and simply been rewritten later. For him, the real value was the clean separation between the agent, which should exist independently of any product, and the product layer. With everything in one codebase and one language, he says, you inevitably get "a little bit sloppy," entangle things, and block later innovation. The Rust boundary helped enforce that separation.
Open source: the benefits and the sting
Codex's CLI, SDK, and app server are open source, which the host noted is unusual among the major labs. Sottiaux gives several reasons. Codex is a coding agent, so the team expected people to point it at its own code and hoped a community would use it to improve it. The team also believed that if they succeeded, open source and the role of code would change, and they wanted to be inside that community rather than separate from it, because problems are hard to solve without seeing them firsthand. And since it was early, they had ideas about what a good harness looks like, co-designed with training and research, but not all the answers. So they published technical deep dives, wanted to learn from other open source projects, and tried to create a level playing field that encouraged tinkering.
When the host asked for an honest look at both sides, Sottiaux listed benefits first. New hires have usually already read the repo and its PRs, and onboarding largely means exploring the repo with Codex. The project gets many good contributions. The team also gets energy from contributing to the community directly, not just saying they care about it, even though it costs effort they didn't have to spend.
He named three downsides. First, the open code is separate from the rest of OpenAI's code, which sometimes forces artificial boundaries and work across multiple repositories. Second, when the team builds something exciting in the open, others sometimes copy it before OpenAI releases it. He says that is "part of the game" under a very permissive license, but it "does sting a little bit." Third, like other maintainers, they face a "tsunami" of random contributions and the tax of handling them, though he says this also pushes them to find solutions.
Supporting other companies' models
Codex also isn't tied to OpenAI models. Sottiaux says coupling an excellent community harness to your own model felt wrong and "quite disappointing." Since the code is open, anyone could fork it and add another provider in about ten lines. Forcing people onto a fork just for that seemed silly, so the team supported other providers directly.
He adds that optionality helps OpenAI as well. A user who wants to try a new model doesn't have to rebuild their setup, and OpenAI gets feedback on what that user liked or disliked about it. The team also tests other models in the same harness. He says companies OpenAI works with care a lot about this optionality, and OpenAI leans into it. He wants to win users with the best and most efficient models and the best product, not through lock-in. He argues that forcing people to use a product because of one constraint would not attract the best people to work on it, and that the experience should feel delightful regardless of which model powers it.
Where the agent runs: sandbox, local machine, and cloud
By default, Codex runs entirely on the user's machine inside a sandbox, and has for more than a year. Any command that needs permissions beyond the sandbox requires the user's approval. Users can instead run tasks in the cloud on a managed VM, the same one used by ChatGPT's work mode, isolated in a Kata container. In that case only the input and the streamed output touch the local machine, which saves local CPU and allows much more scale.
Sottiaux calls this one step. He expects running on cloud machines to become much more seamless, possibly with some execution local and some in the cloud. His reasoning is that more capable models can use far more compute than a laptop has, so limiting them to local execution would eventually become a constraint.
The host objected that local execution is valuable because their tools and a local Postgres database are already there, and asked about cloud development environments, which were popular around 2022–23. Sottiaux says that outside large tech companies, cloud dev boxes never took off because of the high upfront setup cost and ongoing maintenance, which solo developers and small teams don't recoup. Now that agents are capable enough, he argues, that setup is "almost free": the agent can replicate your local database, servers, and MCP configuration on a cloud box and keep them in sync. He predicts a resurgence of fully cloud-orchestrated machines that free people from their laptops. He points to ChatGPT's work mode on mobile, which he says he uses each morning over coffee to dictate tasks, with access to his calendar, email, and Slack. He notes that Codex remote still executes on the laptop, and says it would be wonderful not to have to keep the laptop open.
The harness is "a little bit ahead" of the model
The host recalled that early Codex didn't run their unit tests after making a change, and a few months later it started doing so automatically. They asked whether that came from harness changes or model changes. Sottiaux says the harness is always "a little bit ahead of the model." A model can do certain things, and the harness gives it "crutches" so it does them reliably, efficiently, and the way users expect. The harness supplies guardrails and safety, makes the model more steerable, and writes the developer message injected into context at the start of each turn to shape the agent's behavior.
Test running is his example. At first you might have to remind the model to run tests. Then a better model is trained that reflects more on what the user actually wants, and the reminder is no longer needed. Over time, he says, both the developer message and the harness shrink.
On how the team sets goals, Sottiaux describes co-design between research and the engineering team that builds the core agent harness. For each gap or desired feature, they ask whether it should be a harness change or a model change. If it's a model change, they ask whether it can land in one, three, or six months. Depending on the answer, they may do nothing in the harness and wait for the model. They also use agents to analyze feedback and surface themes. That analysis covers coding and other domains where people now use these agents, including finance, communications, and marketing, along with subcategories. Better pre-training and better overall models lift everything, but some areas get extra attention.
He gave a concrete example later in the conversation. The /goal command, added a few months earlier, was a harness crutch that kept the model focused on a single goal for very long periods, allowing it to run for days or weeks on hard problems. With the new generation of models, he says, /goal is no longer needed: you can tell the model to work for a week and it will.
"Have you asked Codex?": how work gets done on the team
Asked what a new Codex team member is told about how things work, Sottiaux says he introduces them to great people, and the answer they hear most often to any question is "Have you asked Codex?" Inside OpenAI, Codex is connected to Slack, documents, and code by default, and new starters are still surprised that they can ask it almost anything and usually get a good answer, including the state of a project, who is working on something, or why a decision was made. For that reason the team works in public channels and shares documents with broad permissions so that both people and their agents can reason over the information. He mentions additional team-collaboration tools that are not yet released, some of which are planned for DevDay.
His general guidance to newcomers is to care about the user, the coherence of the product, and where the models are heading. If you find yourself writing a 10,000-line crutch to work around a model's flaws, he says, you are probably doing the wrong thing. There is a set of principles, but mostly it is team culture passed on by colleagues.
On shipping, Sottiaux says the process is surprisingly similar for Codex and for ChatGPT, even though ChatGPT reaches about a billion active users. An engineer can make a change and have it shipped the next day or even the same day. People are empowered to make large changes, with a strong expectation of ownership. What's asked for is evidence that the change will be well received, that it's worth adding, and that it's worth maintaining. He adds that maintenance costs have dropped substantially, so the team weighs this differently than two or three years ago. Code review, deploys, and catching regressions are largely automated, so engineers can focus on the idea and its value to users.
He also described the North Star: a delightful, simple personal AGI that knows what it needs to about you, has access to the right resources, and can take sometimes risky actions on your behalf, with a push notification so you can verify them. It would know your schedule and goals, be proactive, and be controllable through natural language and voice, perhaps even understanding your emotions from a camera feed. It should not be "a thing with 10 different buttons and configurations." In his words, "AGI should be simple to use."
How code review is changing
The host noted Google's strict review culture and the traditional benefits of review: knowledge sharing, a second pair of eyes, reducing bus factor, and architecture discussions. They asked where human review still adds value. Sottiaux described one of his early Codex projects: working with research to build a code review model that could catch logic and reasoning errors that would take a human hours to find. These are errors that require digging three or four levels into dependencies, for example discovering that a third-party library's documentation is wrong, that its implementation differs from what you assumed, and that your invariants therefore don't hold. Unless you're an expert in that library, you would miss it.
Those review models were released, and he says their capability is now part of the mainline models, which OpenAI's benchmarks show as superhuman at code review. He says this applies to security as well. Automated security review is now mandatory on all of OpenAI's pull requests, and PRs flagged with a security issue are blocked from merging automatically.
His view is that review always served two purposes: correctness, and a social ritual for sharing information and getting people aligned, sometimes the only place a discussion happened, because merged code runs in production and must be maintained. He expects correctness and security review to be automated. What remains is discussion of intent: what are you trying to do, and should you be doing it? He argues that conversation doesn't need to happen around code. The pull request was a forcing function because it was the last point before code reached production, but intent can be settled in other ways, "and then the code doesn't matter as much."
He described a "box" model. Teams agree on what a component does and which invariants it must meet, with strict guarantees around resource use, data access, and security. Inside those limits, the implementation "could be literally anything" and doesn't need discussion. That agreement deserves a good conversation, perhaps with an agent's help. Once it exists, changes inside the box no longer need your attention. The host agreed that reviews had always been mixed, with some real learning and a lot of chasing reviewers and context switching, and said one upside is not having to spend attention on basic things.
Maintenance and re-architecture get cheap
Sottiaux calls maintenance a tax paid over time to keep things running. It remains necessary, but he says much of it will be automated. His example is upgrading a third-party dependency: with a good changelog and well-documented code, a model can reason through the change and apply it across a codebase in a couple of hours. Teams used to postpone that kind of work even though it matters, especially for security patches. He says a large part of maintenance "just kind of comes for free."
He also highlighted re-architecture. Redesigning a system because of new trade-offs, a better understanding of the workload, or a feature the current design can't handle used to be very expensive, sometimes taking years. He says that is now greatly accelerated, and the cost of mistakes is falling. But he argues the old rules still hold: good abstractions, "boxes" with invariants, and designs that let you change internals quickly without affecting other services.
The host connected this to a conversation with Peter Steinberger, creator of OpenClaw, before he joined OpenAI. Steinberger said he doesn't read the code, yet clearly held the architecture in his head, re-architected often, and designed for modularity so a hundred contributors could work without colliding. The host suggested that this kind of structural thinking, once reserved for architects and staff engineers, now matters for every engineer. Sottiaux agreed and said GPT models are getting better at it too: not only clean code within a file, but whether the architecture reduces maintenance burden and leaves room for future features. He says models are starting to handle this kind of "engineering over time" well.
He finds it fascinating that software now moves through its lifecycle much faster. A project used to grow from a couple of engineers to 50 or 100 over a year or more, giving time to onboard people and write documentation. Now, he says, a project can suddenly have a hundred agents contributing to it "in a weekend."
Does it sting to hand off the craft?
The host asked whether it bothers Sottiaux to watch models take over skills he spent years building. He admits a craft element: he still occasionally opens an editor to write code by hand, and has fond memories of late nights in Vim drinking Coke Zero, thinking only about the problem in front of him. But he says what matters is being in flow and solving problems. If code is a tool for solving problems, you can now solve far more of them. A benchmark you weren't sure about can be started in the background in about 30 seconds and give you real numbers for a better trade-off. At OpenAI, he says, this means running inference more efficiently, getting more effective compute, and deploying it to the world. He says he hasn't yet met anyone who thinks this is bad or not fun.
He also pushed back on romanticizing the old way: there were many late nights three hours into a refactor that turned out to be a dead end, or wondering why something wouldn't compile. The host said they no longer leave work half-finished before bed, because a task ends either done or with evidence that it failed. Sottiaux said he often sends Codex to investigate bigger open questions overnight and looks forward to the results in the morning. Asked whether anyone runs out of problems, he said OpenAI is "not out of problems," with a long road ahead in mathematical and scientific breakthroughs and in building for people "in a deeply human way."
The merge: bringing Codex into ChatGPT
The host said "the merge," Codex appearing inside ChatGPT, looked only moderately interesting from the outside. But people at OpenAI had described a large engineering effort, and Codex usage numbers had grown much faster since. Sottiaux says the core difficulty was two completely different stacks. ChatGPT is fully managed and cloud-based, built traditionally for scale and efficiency. Codex was fully local. The task was to deliver the same capabilities as a local coding agent in a cloud product that could serve tens or hundreds of millions of users efficiently enough to include in the Plus plan.
The result, ChatGPT's work mode, runs the full Codex harness on a cloud computer. He calls it a powerful, permissive machine with internet access. He says users have discovered that with creative prompts they can get ChatGPT to train another model there or install Blender for 3D modeling. He says the team did it quickly, with Codex helping build infrastructure and reconcile many small differences between Codex and ChatGPT, such as merging the plugin architecture and libraries. The goal is a unified product: nothing you can do in Codex but not in ChatGPT, or the reverse, with the same intelligence accessed however you prefer.
He also described Codex acting as a "journalist" throughout the project, documenting the steps, debates, and discussions, which he says were lively: how to do it, what to name it, when and how to introduce it, and what to merge into what, with many permutations considered. Codex produced a full account over time. Internally the project became known as OpenAI's "toggle arc," after the work toggle, which was itself debated before the team came to like it. Sottiaux says the toggle is a temporary state in which work mode offers stronger capabilities, and the plan is to unify further and eventually bring those capabilities to everyone who uses ChatGPT.
How Sottiaux works day to day
The host relayed a question from Peter Steinberger: how does Sottiaux stay cheerful when his calendar looks like Tetris? Sottiaux replied that his calendar is fine and that tools like Codex let him do much more. He has moved much of his work to mobile through ChatGPT's work mode and uses dictation heavily. Instead of writing a question down for later or delegating it, he sends it off and gets a report. He has built up custom skills and instructions so the output (reports, slide decks, code explorations) comes in a form he can absorb efficiently, and he often dictates to his phone between meetings.
Because so much of the team's work lives in Slack, Notion, and Google Docs, he says there is essentially no question he can't have Codex take a first pass at. His examples include public sentiment on a feature, production logs showing how much a feature is used, a list of features to deprecate for lack of traction, and what a particular team is working on. He says he gets answers within about 30 minutes. On weekends he builds prototypes and explores ideas about the product's future, often with different people from the teams. He can take an idea he woke up with and put something in front of colleagues within a day, so they can critique it and ideally be inspired by it. Shipping it is not the point. It gets the idea "out of my system" so he can move on.
Advice for engineers who want to work in AI
Sottiaux named two qualities. The first is deep curiosity about how things work, combined with the ability to understand things quickly. He says people who do extremely well at OpenAI can grasp a system fast and make sense of a new codebase. Agents help with that now, but much of it still comes down to asking good questions and continuing to dig. The second is being in tune with the community or people you're solving problems for, including cases where the direct beneficiary is another group that is itself serving people. That means being crisp about taste, needs, and requirements, and thinking clearly. If you can't explain your intent, lack a connection to a community, or lack taste, he says, it will be much harder to do great work.
The host's reflections
In closing, the host highlighted Sottiaux's candor about open source's downsides: competitors copying features before release, and the burden of low-quality contributions. The host found the idea that the harness is a set of crutches the next model will no longer need a little demotivating as a developer. They suspect the work also involves building tools models will use, and hope future models won't reinvent things like MCP, skills, or plugins on their own. They noted that the merge meant bringing a fully local agent into a managed cloud stack efficiently enough to fit a $20-per-month plan. On Codex serving as the project's journalist by reading every Slack thread and document, the host said it had a "big brother" feel and could become normal at startups, and said they have not yet decided how they feel about it.
A fun story that you recently shared on X as well is how you were part of the team that built this internal Google bot that was ChatGPT but a year before ChatGPT.
That caught up like, you know, wildfires. It felt more than a research project.
For Codex you built it in Rust and at the time the model was not on distribution for Rust.
Turns out it was quite clear that Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it. When someone joins the Codex team, what do you tell them? How do things get done here?
The thing that they hear the most about when they have a question is like, "Have you asked Codex?" It still surprises new starters that you can basically ask it anything.
Building was the fun part, but then maintenance was the painful.
So maintenance is really sort of like a tax that you pay over time just to keep things running. But where I think it changes is like a lot of it is just going to be automated.
So you still have the concept of code review.
The role of code review is changing. And the role of code review now is like I think
Codex is one of the most popular AI coding harnesses today. But how did it all start? Many of you will know today's guest Tibo from his generous and pretty frequent Codex usage resets. He was also there when Codex as a product started and has led the broader Codex team since. Today we cover how Codex started and why it was built in Rust and made open source. How code reviews are changing inside the Codex team and OpenAI. What it means when maintenance and rearchitecting are getting ridiculously cheap. What the merge of Codex into ChatGPT looked like and the many underappreciated engineering challenges of this project. If you want to understand how teams inside of OpenAI plan, review and ship software, this episode is for you.
This episode is presented by Turbopuffer, a ridiculously scalable, fast and cheap hybrid search engine built on top of object storage by an engineering team that I've really grown to like after spending time with them. Turbopuffer is the tool that companies like Anthropic, Notion, Cognition, and Harvey all use to connect their AI products to massive amounts of unstructured data. When I've talked with engineers who use Turbopuffer, the theme that always comes up is reliability and performance at scale. The reasons for this have everything to do with Turbopuffer's architecture. Turbopuffer uses only object storage for state and NVMe SSDs with memory cache for compute.
Data in Turbopuffer is organized into namespaces. You can think of a namespace as a database table or a search index or an S3 prefix depending on the world you come from. When a namespace is not being queried, it stays on cheap object storage with no associated compute cost. When a namespace is active, Turbopuffer pulls it up into hot caching tiers, so queries are very fast. This design fundamentally makes it effortless to scale to hundreds of millions of namespaces. If you're building a multi-tenant AI product, every user and their agent can have their own dedicated search index without any overhead. And each namespace can hold hundreds of millions of documents without any special configuration. You can scale Turbopuffer virtually without limit. And the performance, reliability, and operating model all stay the same. If you need to connect AI to lots of data, Turbopuffer should be your first choice. Check it out at turbopuffer.com/pragmatic.
Tibo, welcome to the podcast. So good to have you here. Thank you for having me. It's so good to see you again.
It's good to do this. Last time we did it in person. Now we're doing our video. First, I wanted to ask you, how did you get into tech? When did you first know that you want to work with computers?
It's a good question. It was a long, long time ago. My parents actually decided to move out of Brussels where I was born and just thought it was great to just buy a small house and refurbish it. But it was in the middle of a village with not much going on. I think there was like roughly 200 people living there. Not many that I felt like I wanted to talk to or, you know, could make friends with.
And so I kind of got stuck. This is like very early, like 8 years old. I kind of got stuck because like, you know, computers and, you know, it's like early days of the, for me, the internet and, you know, that was my way to learn about things. And so the rest is just like, you know, came from that. I sort of like owe it to my parents to, you know, have moved into the middle of nowhere and then, you know, I had no choice but to get interested in computers.
Once you finished high school, like, you went on and went to university, right? Actually studying it properly.
Yes. I studied mathematics, applied mathematics at university. I went there quite early. And so I graduated early as well. Like I thought for a long time that I would actually not make it and that I would drop out. I had like small companies and small consulting business while I was studying. I was working for banks. I was like very interested in supply chain and applied mathematics problems and I sort of like selling that and learning a lot through that.
Eventually ended up in the startup world in Belgium. I did that for a little while and then moved to London to work initially at Google and then DeepMind and then now, you know, moved to be here at OpenAI in California. I love the California weather. We can talk about that. It's been very good.
Right after university you've founded a startup, right? You had the startup bug in you or the entrepreneurial bug.
Yeah. So this startup was all about pharmaceutical supply chain, looking at the supply chain for clinical trials and like try to optimize and decide like, hey, you know, should you produce more medicine, where should you send it, where should you dispatch it, like how do you avoid waste, and through that making clinical trials more efficient. And this was using traditional, like non-ML techniques, like more optimization solving, Monte Carlo simulations, these kinds of things, stochastic multi-stage optimization problems really. And we also applied it on steel industry and we applied it to electrical grid as well in Europe. It's like anything that sort of had the shape of an optimization problem we sort of like got interested in. And, you know, to this day this company still exists and I think they do some of the most interesting work still, but it's changing a lot, you know, with modern AI for sure.
But it's interesting because you kind of said like, oh yeah, that wasn't ML, it was just the traditional stuff, and then you go into like Monte Carlo simulation and optimization and this algorithm. I get a sense that you kind of just went deep, right? That it was like, okay, here's a problem space, like how can I use mathematics, stuff that I learned, stuff that I didn't learn, to just go deeper and deeper. Do I sense that correctly?
Yeah, that's why I was obsessed with applied mathematics. It's just really this idea of: you have theoretical mathematics or you have theoretical science and physics, and there you do it because there's something to be discovered and something beautiful about it, and it's all about patterns and pushing the frontier, but you don't necessarily always know how you're going to apply it. And then there was the real world, right? There's all these cool problems that just lie around, and I was very interested in seeing, you know, how can I make the world better, and so how do I apply sophisticated mathematics to just optimize the world around me. And that was like a lot of the thesis behind that startup.
Yeah, and then after the startup you ended up at Google, first at Google London. It was in 2015, and I remember in 2015 Google was a really, really competitive place to get into, maybe as competitive as OpenAI is today in terms of the industry or in terms of prestige. You worked on Maps initially and then you moved over to DeepMind. Can you talk a little bit about what you worked on, and then why did you move on from an already really interesting space that you clearly loved, you know, like optimization, logistics and all these things?
Yes. I didn't start on Google Maps. I started on a project that was meant to make the web faster and to make websites faster, especially on mobile. At the time, you know, Google was kind of seeing the transition from desktop to mobile and more and more traffic going to mobile phones, and so wanted to get ahead of that. So it funded a number of initiatives and projects. I was working on one of them. This was really, really fun because it was a small group within actually the ads organization. It was meant to sort of offset the ad revenue loss because of the shift of traffic to mobile.
And I worked on it for roughly two years. And then it was cancelled, and although it was the most fun I've had on, you know, solving hard technical challenges, I learned a lot from not having product-market fit, not having the right users, not having the right feedback loop, not trusting your product manager when they say the project is going well when in fact it's not going well at all. And then, you know, one day this VP flew in from California and then it was just like, oh yeah, you know, we're canceling this project. Unfortunately, you only have, you know, hundreds of users and this is clearly not Google scale. And it's unbelievable, but people were surprised. And I think there's a lesson there that I carry with me, of course, is, you know, just always question, always deeply think about the impact that you're having, but also the importance of the overall project that you're contributing to.
And then I moved into Google Maps. Google Maps was super fun. Worked on reviews. And then after roughly a year, I couldn't ignore DeepMind. It was this special place, headquartered in London. So many great things were happening. This was really the early days, you know, with rumblings of things like AlphaGo, and they just seemed to be doing extraordinary things and, you know, really tackling the very, very hardest problems that you can tackle. And with my background I was obviously drawn to that.
Started there. I worked on a lot of the research infrastructure, research tooling. This is a theme that I carried on for almost a decade, and this is very much also how I approach things: how can I build tooling and products that help make others more efficient and bring a lot of utility to them. Initially I was doing this for research and then over time, you know, I got into thinking about things in a much more general and general and general way, you know, eventually ending up where I am now.
Yeah. And a fun story that you recently shared on X as well is how you were part of the team that built this internal Google bot that was, you know, if you want to say, similar to ChatGPT but a year before ChatGPT. Can you talk about that? That is a new story. I haven't heard it before.
This was part of DeepMind. There were multiple efforts as well. There was Brain as well that was separate at the time. They had their own efforts on large language models, but it was definitely something that was being explored. It was not the main thrust of DeepMind. DeepMind was very much worried and busy thinking about grand challenges and games and, you know, thinking about RL, not in the language sense. And so there was this group that was pushing on large language models and, you know, thinking: what if large text corpuses are everything? What if you just pushed language to its maximum and you just scaled language models, would that be enough to get to general intelligence? That was a hot debate at the time, and then one group decided to just really push on that.
And then it felt really natural, you know, as I was building tooling with others for research. Obviously you're like, hey, what can we do with this model, how do we present it to the researcher, how can they debug the inputs and outputs, and eventually you sort of end up with, you know, a chat system. So we built that internally. We had a lot of fun. Initially the models were kind of almost a little bit absurd, not very coherent, not super useful, but it was a lot of fun to sort of tinker with them. That caught up like, you know, wildfires. This application, everyone was kind of sharing little conversations within DeepMind. It felt more than a research project or a project for researchers.
And so then there was this desire over time to launch it as an external product. But DeepMind was just not set up. You know, there was the right way to launch products at Google. There's the whole machinery of how you do that, the whole blessed production stack. Obviously very, very optimized over the years to do things well, but also very, very hard as an environment to truly innovate.
And then I wanted to ask what made you, you know, look around or maybe even consider OpenAI, but I feel you partially answered this question. Just putting myself back into your shoes, like, you know, if it's 2024 or 2023, you're inside of Google, who are publishing amazing papers, doing really good research. You're doing super fun stuff, right, that's pushing the limits of what's been done before. It's inside a company where you already moved. You know, for people who are feeling kind of comfortable or good about where they are right now, which I imagine you must have been, what made you still explore, like, what else might be there?
Yeah, I was very comfortable. It's a good place, but really I had a desire to, you know, meet great people, but also join a mission that I truly believed in, and where I felt like the people were true to the mission and cared deeply about impacting the world in a deeply positive way, but also in a direct way. Not being like, oh yeah, we just do this work over here and then it's the job of someone else to figure out how to make this useful. I wanted to join a group where all the parameters were sort of considered together, where research and product were really co-designing.
OpenAI was just crushing it. I thought ChatGPT was, you know, taking off. I met a couple people from OpenAI and then I was like, wait, what? You only have like 20 people working on ChatGPT? That is an insanely small number. That must be extremely empowering. How does that work? How do you manage to maintain a product with that level of scale and with that level of autonomy with only 20 engineers? And then as I kind of dug and dug and dug, it was just an amazing group of people, amazing mission, super talented, super driven, and it drew me in and then I joined.
Pre-reasoning efforts. Immediately, typical OpenAI fashion, I joined and it was like, oh yeah, you know, there's this thing going on, we're going to launch reasoning models, it's like some new paradigm, and then, you know, start sprinting on that, and like a month later the company launched o1, o1-preview, and that was exhilarating to be part of. I wanted to
be part of like a place that moves fast, cares about impact, would be in tune with the world and, you know, just really listen. And sort of like that's also to me, like, you know, what I've carried with me when building Codex, when building products, is like having a community, listen to the community, just really focus on like a really intense feedback loop, and then building something that is just like, you know, you just really want to care about it and like, you know, care about the utility of it that it provides to the world.
And then of course you started to work pretty quickly on Codex. So you joined in 2024. Can you take us back: what the thinking back there when you joined was about AI or LLMs and code? I know there was this A-SWE effort back then. We talked about it in the deep dive as well that we did in The Pragmatic Engineer, the autonomous software engineer.
A-SWE. Yeah, that's what it was pronounced internally. We don't have an A-SWE effort anymore, like, you know, it's Codex. But really for me it was, I joined, I started building infrastructure for research. A lot of what I did before was large-scale data storage, analysis, and then tools to understand training runs. I did a lot of different things over my years, but it was always about building for others and making them faster and just really caring about, you know, fundamentally doing that well, and then through tooling and infrastructure making new things possible. And so when I joined OpenAI, I was like with the same idea, and then with o1-preview and, like, you know, some of the later models, it was very clear that we had to use the models themselves to help us go faster. And so I just really got obsessed with this idea of what were the limitations, how were we going to use those models for research itself.
So got together with other folks in research. We started training models. We started building little agents. Those were truly the precursor to Codex. And like this was like we were training internal models to be very proficient on the Python codebase of OpenAI and then very proficient with, you know, having like good taste in architecture, good taste in, you know, like code style. It was like Python only, and then the idea was like, you know, we would sort of like use that to build infrastructure very quickly and, you know, help researchers code faster as well. And then, you know, we would move faster, and then over time when you just kind of like push that and simplify it to its core, we found like you could make a lot of progress very quickly and then learn very quickly.
And then Greg and Sam are, you know, they're immensely supportive, and also Greg was very adamant that, you know, we would not just focus on ourselves but we would also focus on benefiting the world. And so he just sort of encouraged that we would be thinking about this not just as a tool for OpenAI itself but also as something that we would actually make into a product. And this is when we merged this research effort with this A-SWE effort and we started building one thing, and then that led to a sprint which was like the initial cloud Codex that we launched, which didn't really have PMF because it was like a little bit too high friction, and then we also launched the Codex CLI and we continued to push. But it was always this idea of, hey, how do we get models to really help here.
You mentioned that first you started to build this model to train on the Python code and actually help build infra better, but then you made this interesting decision where for Codex you built it in Rust, and at the time the model was not on distribution for Rust, right? It wasn't as good in Rust as it was in Python or TypeScript. Why did you make that kind of a decision? Was it kind of like, did you expect that it'll catch up, or you figured that performance is more important? Because it was very counterintuitive. Most of the other harnesses built were actually not built in Rust. They were built on distribution, on TypeScript or Python or something else.
Yes. From first principles, very early on we were thinking about the product interface and the agent as different things. So it was very important to build the core of the agent in a way that was robust, that was secure as well, that was, you know, engineered for efficiency and scale. And having worked through projects over the years that go from "hey, this is a fun thing" to like "hey, we need to scale this to the scale of like the largest data center," the decisions early on really turn out to be quite important, as long as you don't sacrifice too much of the velocity. And so it's a trade-off, but we had very prolific and amazing Rust developers, our internal models were not bad at Rust, and then you get a lot of validation as well at compile time. It's like, you know, statically verified and all these things, and that is great for agents too. So turns out, you know, it was quite clear that Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it. But primarily we were focused on correctness and we were focused on efficiency as well.
Interesting. So you're saying, you know, in your case it was worth thinking ahead of where you want this thing to be, and for example things like a language choice. Obviously with agents you can rewrite a bunch of stuff easier than in the past, but still you can save yourself reworking by putting in the right, I guess, scaffolding, or, well, you know, the baseline of what you're building on, right?
I think we could have been successful if we had written it in TypeScript or, you know, maybe even Python, and then it would have been fine, and then, you know, we would have rewritten it at some point. But having a very clean separation between the agent itself, which can exist irrespective of the product, was a very important principle. And if you write everything in the same codebase, in the same language, inevitably you're going to be a little bit sloppy and you're going to intertwine things more than you should, and then it's going to prevent further innovation after that. And so that was very important. Like, the Rust boundary in a sense was very useful for that.
One interesting decision that you made, which is unique across all of the major labs, is having this built in open source, right? The CLI is open source, the SDK and the app server are all open source. When and why did you decide that? It's not a given, especially, you know, there used to be jokes about OpenAI having things closed, but this is actually the opposite, where this is open, whereas some competitors would ship closed-source harnesses, which again, I think it's very easy to understand why you would want something closed source. Why did you want it open source?
There was something really cool about the idea of having the code open source, because fundamentally what you're building is a coding agent. And so we were sort of like thinking about, well, if you have that, you know, you're obviously going to point it at itself, and, you know, maybe you can build a community of contributors that use it to improve it, and then you can learn a lot from that.
Also, it felt at the time very clear to us that if we were going to be successful, open source itself would change and the role of code itself would change, and so being part of that community seemed important instead of divorced from it. I think, you know, it's hard to solve problems if you don't sort of like witness them yourself.
And then the other thing was, it still feels like early, but it was very early at the time. It felt like we would have some ideas for how to solve things well, and we were co-designing these, you know, with the training and the research, and it's all about expressing the capabilities of the model in the most flexible and the best way. But also we didn't have all the answers, and sort of being very open about, hey, this is what a good harness looks like, this is how we think about it. We did a couple of very technical deep dives and blog posts and we talked about it a lot, and we thought, you know, hey, the world is vast out there, there's, you know, crazy smart people. We're going to get inspired by other open source projects as well. And so let's just make this a level playing field and sort of encourage a lot of tinkering and exploration at this stage.
Now this has been, you know, like a year later, a year and a half later, which is a very long time right now in this AI time frame. But looking back, or taking the experience, what are the benefits you've seen, the kind of engineering benefits, the engineering team's benefits from being open source? And just honestly, what are things that are kind of hard about being open source? Like, there must be downsides. Just trying to get an honest take on both sides.
Yeah, there are definitely downsides. It comes at a cost, right? The benefits are, it's awesome to build in the open. It's awesome to have a small repo as well. Like, whenever we hire someone and they join the Codex team, it's like they've seen the repo before. They've looked at PRs. They're like
Onboarding is done.
You know, it's done. Yeah. Onboarding is just like you use Codex to look at the repo, you know, with you and you ask them questions, but it's not a secret issue that you can get productive right away. We get a lot of good contributions, although we get like, you know, a tsunami of random stuff as well.
Obviously, you and everyone else, right? Open source is changing. I think this is one of the examples.
That's right. And then to me, and to a lot of the team, it just brings a lot of energy to be part of the community and be directly contributing, not just saying that we care about the community but actually doing things that, you know, you can see it's costing us effort, right? We don't have to do it. The downsides are, you know, it's separate from the rest of our code. So sometimes we have to draw artificial boundaries and, you know, work across multiple repos. When we're working on something particularly exciting and we're building it in the open, then at times we find that others copy it before we have the time to release it. And it's just a little bit sad. But also it's part of the game, you know. You're building in the open. That's sort of the contract that you signed, is like, you know, you can copy it. We have a very permissive license as well, but it does sting a little bit when you're working on something. And then the third thing is just like everyone else, you know, we are overwhelmed with random contributions and we have to deal with that additional tax. But then that pushes us to also try and solve for it, right? Which I think is good.
And on top of the open source, one thing that surprised me about Codex, and I didn't even know about it until recently: it's not tied to the OpenAI models. You can use other models with Codex. You know, putting myself in a vendor's shoes, it might not be very obvious, because again, all the other vendors I look at, when they do a CLI, it's kind of "use it with our models." Again, what made you decide to be this permissive about allowing people to use your harness with other models?
It felt quite natural. If you are part of this community and building an excellent coding harness, why would you couple it to your model? That felt quite disappointing to make that decision, so it didn't feel right. And in general, I kind of try to make decisions where I'm like, yes, I can just explain it, you know, it is correct. It's the same reasoning with it being open source in the first place. It would have been trivial for anyone to fork it and then add support for another thing. But then you're just encouraging people to go and use that fork, and now suddenly you have overhead, and the only reason you have a fork is because you wanted to change like 10 lines of code to add support for another model provider. That feels very silly. So why not just support it in the first place?
The other thing is we benefit a lot from being able to just give optionality. So, you know, maybe today you love using OpenAI models and you're super productive with them, but tomorrow there's a new model that comes out and you want to try that. Why force you to go and completely change your setup just to try a new model? And then we benefit from the feedback that we wouldn't otherwise get, which is like, you know, maybe there's something that you liked about that model, maybe it actually didn't work well. But being nice to our users and to the community feels like the right thing to do here. And then we also try other models, right? We try them in the same harness, and it's all good. And then this optionality is also often very important to companies that we work with. This is something that we absolutely lean into.
This last point, I think, you know, as any serious company, you want to have optionality and you want to use a tool that gives you that optionality. But I kind of appreciate it, because it feels to me like it's kind of honest. It forces the whole company to compete to be the best everywhere, in the model layer and the harness layer, with open source, with choosable models, and it kind of doesn't allow you to kick back and say, "All right, we're done. We can hang back for a little bit for now."
Yeah. I want us to win users by having the best models, the most efficient models, the best product, and then, you know, if we do all of these things, we're going to have a good time. If we sort of force you to use the product because of this one thing, then I don't think that will attract the very best people to work on this product either. We're doing our best work here. We care a lot about the experience. It should feel delightful, you know, irrespective of the model that powers it, it should feel delightful.
I love the idea of winning based on merit, not based on lock-in. And this is a perfect time to mention our season sponsor, Entire, who also play by the same rules. Like it or not, Git is becoming a bottleneck for modern agent-heavy software development. Devs are creating more code with agents. These agents are pushing more code. Many devs are running more parallel agents. These are pushing even more code. GitHub is clearly struggling to keep up and has frequent outages. So what's the solution?
Entire was founded by GitHub's last CEO, and he rebuilt Git hosting for the agentic era from scratch. Entire was built to be very fast and to have your repos regionally close to you to reduce latency, allowing for fleets of agents to push in parallel. Some numbers they published: Entire can handle 418 pushes per second. That's up to 89 times faster than every competitor on the market. When GitHub is down, you can still keep working, and you don't even need to migrate from GitHub. You just sign up to Entire and the platform mirrors your repo. And one more neat thing: have you ever wondered what prompt resulted in
this specific code being generated? I find that the prompt and conversation with the agent carries more information than the PR itself, at least for me. Entire captures all the prompt history with your agent right in the repo, easy to check back, and has a pretty innovative UI to show all of this. If you're looking for Git hosting that works even when GitHub is down, head to entire.io/pragmatic, install the CLI, and mirror your repo with a click. I've already done it. Oh, and did I mention that it works with any agent and it's open source?
I'd also like to mention our season sponsor, Antithesis. Tibo talked about how the experience of the software you use should feel delightful. Delightful includes no annoying bugs. But when you're using agents to write your code, how do you avoid shipping bugs? Reviewing every line of code is becoming a challenge with the amount of code that agents generate, which is why Antithesis goes well beyond code review. Antithesis runs your whole system in a hostile simulation. This simulation includes both targeted testing and fuzz testing. By running this simulation, it finds every bug before your users do. And because the simulation is fully deterministic, it doesn't only find bugs, it gives you a perfect reproduction of every issue, which makes it much easier to fix issues. The first thing I thought when I heard about Antithesis is that automated bug discovery and fully deterministic testing sounds like science fiction, but it's actually hardcore engineering under the hood. Jane Street, Fly.io, and the etcd community ship agent-written code with full confidence because they know it's been verified by Antithesis. To see more case studies and details, head to antithesis.com/pragmatic. And with this, let's get back to Tibo and why competition between tools is great.
Yeah. And I think as an engineer, I always see that whenever there's competition, as someone who's using tools, it's always amazing. I remember when Microsoft had with JetBrains the IDE wars, and then there's the clouds battling with each other with all the features, and now of course we have the harnesses, we have the models, and as a user it's great because now we have more choice, they just develop faster, I guess our voice gets heard a bit better. So it's great to hear.
Speaking of the harness, can you tell me how it works today, in the sense of, when I start a Codex task, does it run always on my machine? Does it choose the cloud? Does it use a sandbox? And how do I control this or know this, or how much should I know about this as an engineer?
Yes. So, by default, it runs sandboxed. If there is a command that should run with additional permissions outside of the sandbox, it will ask you as a user for permission, but every tool execution happens within the sandbox by default and it runs entirely on your local machine. And this has been the case for more than a year now.
But it is something that is evolving and shifting, where you can select to run this in the cloud, which then runs in a managed VM, where it's the same VM that you get through chat work, and you can sort of inspect it, but it runs in a Kata container, it's a secure environment. And so everything runs inside of that VM and it doesn't run on your machine, and then the only thing that happens on your machine is your input and then the streaming back of the output. And so that obviously is much nicer on your CPU and your machine, and you can scale much more.
And this is just a step. It's going to be much more seamless in the future to use cloud machines and then maybe have a combination of partial execution on your laptop, partial execution on cloud machines. And really the thing that we're thinking about that is very natural is, as models just get better and more capable, they can leverage so much more compute and many more resources than are available on your local machine, and so it would be a constraint at some point to just limit execution to your local machine.
Yeah. One thing that is great about it running locally, and I think the reason I love it when it runs locally... Of course, it's a pain, because if I'm doing some work, I have several agents, it's eating CPU. If I want to close my laptop, I cannot, I kind of leave it half open, right? When I was in one of the offices of an AI company, I had it half open and they're like, "Are you running agents?" I'm like, "Yeah, I have one running." He's like, "I get it."
But the reason I do it is because I have my local tools, I have my local Postgres database, I have this and that. How are you thinking about... the cloud is amazing, but it doesn't have this setup, or it's just a pain to set it up. Are you thinking about or experimenting with making these setups? I'm kind of reminded of a topic that we talked about pre-AI, which is cloud development environments, and around 2022–23 they were hot, and then we talked about AI more, but...
Yes, I think outside of large tech companies, cloud dev boxes really never took off, because there's a very big upfront cost and then you need to pay a maintenance cost as well, and you just don't benefit from it as a solo developer or a small team. With the level of capabilities that we have in agents now, the setup almost is free, right? So the setup cost and this maintenance cost is... if your agent is capable of doing it, it should just do it for you. So for example, if you're saying, hey, I have my local SQLite or I have a local server and MCPs and whatnot, how hard is it to actually configure exactly the same setup and keep it in sync on a cloud dev box? Well, maybe it's not that hard if the model just does it for you. And so I think we're going to see a resurgence of fully cloud-orchestrated machines, which then frees you from your laptop, right?
One thing that we've seen a ton of success with with chat work is it's just available on your mobile. I start my day just dictating a bunch of tasks into it next to the coffee and it just does it. It has access to my calendar. It has access to my email. It has access to Slack, and it's just so awesome to be able to walk around and get stuff done without having to carry my laptop everywhere. And I think it's the same thing: we shipped Codex remote, where execution is still happening on your laptop, but it would be wonderful if you didn't have to keep your laptop open.
Can you tell me a bit on how, in the past, you improved Codex? Because I remember when I first used Codex, this was one of the early versions, you could talk to it, it did stuff, but for example, I said, all right, make this change, and it did that change, and I had unit tests and it didn't run them. And then a few months later, I don't know exactly when, it just started to run them automatically. Were these things... did you improve the script that runs the instructions, I'm not sure exactly what you call it, the bootstrapping script or whatever that is? Is it improving the model? As a dev, how can I imagine you making each version better between the harness and the model, and what's the connection between the two?
Yeah, this is a good question. So the harness in a sense is always a little bit ahead of the model.
Oh, really? How so?
What I mean by that is that you have the model, it's capable of certain things, but then you set it up with a couple of crutches so that it can actually do the thing to a level of reliability, and in a way that is efficient and also with the behavior that you expect as a user. And that's the role of the harness, right? It's to provide guardrails like safety, make it more efficient, make it more steerable, controllable. And then the harness usually is also responsible for what we call the developer message, which is sort of injected in the context at the start of each turn. And so the purpose of that is obviously to affect the behavior of the agent throughout the turn.
A lot of what you have is the result of the harness and the model. Initially maybe you're like, oh, it doesn't run tests, so you have to remind it to run tests. And then we train a better model that is just capable of better reflecting on what it is that you really want when you ask for something, and then you don't actually have to tell it anymore. So over time what we see is the developer message shrinks and then the harness also shrinks.
Inside of the Codex team, do you have specific goals? Do you say, all right, now Codex as a harness and model combined is not very good at this, or it's doing silly mistakes here? How can I imagine how, as the engineering team, you're working on the next version of Codex? Because the thing that I don't really get as a dev is, okay, there's a model, which to me is this magical thing which will get better, and of course I'm sure you have some feedback channels, but you also have the harness, which is the tools that you're building. That's probably what the team is responsible for. How do you even set your goals? In traditional software you'd be like, we will build this feature, and you build that feature because you know how to do it, but it feels a bit more fuzzy to me, this development process.
Yeah, it is, and it's why we co-design most things, and it's a process where it's a collaboration between research and the engineering team primarily building the core agent harness. It's always a question of, okay, we see today that we are very good at this but we're not very good at this, and we have a desire to do another thing because it would be a very cool product feature. And then we always look at it: should this be a harness change or should this be a model change? And if it's a model change, how soon can we have it? Can we have it in a month? Can we have it in three months, six months? And we sort of work through that, and then depending on how soon we can just fix it in the model, at which level of training, we might decide to not even do something in the harness at all and just wait for the model to solve it.
It's agents all the way, right? So we use agents to analyze a lot of the feedback, to come up with themes, to just help us have these conversations and decide on priorities. But we analyze it across all of coding. We analyze it across all of the other domains like finance, comms, marketing, all the things where our users are using these agents nowadays, and there are subcategories within those, and then we roughly know how well we perform, and then we're always pushing the frontier. And a thing that is interesting is, as our pre-training model gets better, as we make the overall model better, the whole thing lifts up, but then there are sometimes things that we pay a little bit more attention to.
You mentioned it's agents all the way. Can we talk about the software development life cycle on Codex, in the sense of: whenever a new engineer joins a team, any team, it's like, okay, how are things done here? Back pre-AI, you'd join a company like Uber or Google and they would tell you, cool, the way it works is we have an idea or the PM has an idea, we make a plan, we get together, we do some estimations, we break up the work, we code the work, we do tests, we do code reviews, we release, we do feature flags, and then we were on call. That's how it used to be. When someone joins the Codex team, they've clearly been contributing to the open source part, but what do you tell them? How do things get done here if they're a total newbie?
I introduce them to great people, and then the thing that they hear the most about when they have a question is, "Have you asked Codex?" And Codex is just, by default at OpenAI, plugged into everything, so it has access to Slack, it has access to all the documents, access to all the code, and it still surprises new starters that you can basically ask it anything, and it will very often just come up with a really good response. And so the easiest way to understand the state of a project, or who's working on something, or why a decision was made, is that Codex knows about it all internally.
And so you just use all of that. For that reason, we do a lot of work in public channels. We open up documents with fairly broad permissions, so that everyone has access to this information as well, and so that your agent can go through things and reason through things. And then we have a couple of other things that are just really very helpful for team productivity and team collaboration that we haven't released yet but are going to come, some of it at DevDay. All of that just makes you very grounded and in tune with the rest of the team, and allows you to very quickly understand the state of things and produce things yourself.
The general recommendation is just: hey, care about the user, care about the coherence of the product, care about the models and where they're going. If you're doing something and you're building this 10,000-lines-of-code crutch to work around the model's flaws, you're probably doing the wrong thing. So we have a set of principles, but it's really sort of a team culture and ethos at this point, and it very much carries on when people join, through the rest of the team teaching the ropes.
And then when I have an idea, I think it's a good idea, I talk it through with Codex, maybe I talk it through with some of my colleagues: here's a cool new feature I'm going to build as my first contribution, or first major contribution, to Codex. How do I go about that? Obviously I code it with Codex, I obviously test it and make sure that it works. From there on, what's the process? Do you still have the concept of code review or AI code review, of verification, of rolling out, of staged rollouts, those things? Because Codex itself goes out to millions of people, it just crossed a big 20 million active user mark, but if it's ChatGPT, then it goes out to an even bigger number of people.
Yes, but it's surprisingly a similar process whether you ship on Codex or ChatGPT, even though ChatGPT goes out to a billion active users and growing. You can ship a PR, you can make a change and get it shipped the next day or even the same day, and it just goes out to a billion users and it's fine. We just really instill a sense of ownership and care, so people are very empowered to make changes, even large changes. The general thing that is being asked for is evidence that it's going to be well-received, evidence that it's a worthy addition, evidence that it is worth maintaining over time. But also, the cost of maintenance has really gone down significantly as well. So we think about these things slightly differently than, say, two years ago or
3 years ago. The other thing as well is we automate as much as possible. So a lot of the process of code review and deploys and catching regressions, all of that is pretty much automated. And so you get to just focus on really the idea and how it's going to help our users, and you care about the coherence of it all, and so the overall power of the agent, and making things better. We have a long, long list of things that we sort of aspire to do and haven't gotten to yet.
And then there's sort of the north star direction, which is a delightful, simple to use personal AGI that knows everything about you that it needs to know, has access to the right resources, can take sometimes risky actions on your behalf, but then you get the push notification and you can verify that. And it's a thing that you deeply understand as a user, but also it knows about your schedule, it knows about your goals, it can be proactive, and it should be extremely natural. It should be something that you can control through natural language, voice; maybe it should understand your emotions if it has a camera feed. It should be the most natural thing on earth. It should not be a thing with 10 different buttons and configurations. AGI should be simple to use.
You kind of mentioned just briefly the code review, but I wanted to go back to it. You worked at Google on a product used by hundreds of millions, which is Google Maps, and Google is very well known for their culture of very strict code reviews. They have, I think, two layers of code reviews. There's a language correctness review, and I think they've really perfected it across the industry for a long time, and they do believe that it works and they use it. How do you think that part is changing, specifically the human review? Because for a very long time, until maybe a year or two ago, I would have said code review has all these benefits: knowledge sharing, the second pair of eyes, removing the bus factor because now someone else understands and when that person is out they can jump in, conversations are happening about architecture, not just the code. But now there's a lot more code, and what was the value of code review? In what cases? And so on your team, because you guys are so ahead of this, where do you see humans, or developers, being involved in the review stage still valuable? And where did you find it fine to hand it off to an agent?
Yeah, the role of code review is changing. One of the early projects that I did on Codex was working with research on developing a code review model that was going to be at a level where it can spot mistakes in logic and reasoning, to a degree where it would require humans potentially multiple hours to capture the same level of mistake, because it requires really digging three, four levels deep into the dependencies and understanding that maybe the documentation actually was wrong, and the implementation of this third-party dependency is different from what you expected, and so therefore your invariants are not upheld. And these things, unless you're an expert in that library, you wouldn't know, and therefore you have a bug. And so we developed these code review models and we released them, and now the same level of capability and ability to spot these mistakes by doing deep verification is just part of the mainline models. When we benchmark them, they're superhuman in code review.
And this is not just true for correctness. This is also true for security, for example, where you're capable of reasoning across very, very complex things and then coming up with, "Hey, you have a critical security vulnerability here," which is now mandatory across all of OpenAI pull requests. We block pull requests from merging if we flag them with a security issue, and this is all automatic. And the role of code review now is, I think it was always about correctness, it was always about ensuring that things worked, but it was also sort of a little ritual for information exchange and bringing people on the same page and encouraging a discussion, which ideally would have happened before, but sometimes it only happens around the code, because once it's merged it actually runs in production, it's doing stuff.
And then you have to maintain it. So there's this social aspect to it as well. I think all of it is changing. The correctness, the security, I think that will be automated. Really what we see and I see is there's a sort of discussion around the intent that takes place around the pull request. It's like, what are you even trying to do? And is that a right thing to attempt to do? I think you can have that discussion outside of the pull request. It doesn't have to be around code.
So maybe this helps crystallize where a discussion needs to happen versus where we did it because maybe we didn't have the type of tooling that we have right now.
Yeah, I think this is going to change, and it was a forcing function, because you have to have that discussion. It's good to have that discussion before you merge it and it becomes production code, but I think there are other ways to have these discussions and design things together and make sure that the intent is good, and then the code doesn't matter as much.
And it's interesting, because when I think back to all my code reviews, of course I have memories where it was great, we had a good discussion or I learned something really interesting, but a bunch of times, honestly, it was such a pain in the ass. I was trying to get my stuff through. You're pinging, "Hey, could you review my code?" And, "No, right now I'm busy." "No, I really need this to unblock me." And then you context switch. I feel it's always been good and bad, right? So I feel whatever we do there will always be upsides and downsides, but now they're just moving. So I guess one upside is as an engineer you might not have to give your attention to just kind of basic stuff that doesn't need your input per se.
Yes, it saves time, and I think progressively what we're going to see is also you have an agreement on the box and the overall contract of what it's supposed to do, and then what is inside the box, as long as you have strict guarantees in terms of resource utilization, data access, security, these kinds of things, what happens inside the box could be literally anything. You don't really need to care, and really what you need to agree on is what does the box actually do and what are the invariants that must be satisfied. And I think that is then worthy of having a really good conversation on, maybe assisted by your favorite agent. But then once you have that and you have that understanding, changing anything within the box doesn't require discussion, and it just really preserves your attention.
The cost of maintenance has gone down. Maintenance is always such a hot topic whenever we build things inside of all these companies like Google, Uber, even startups. Building was the fun part, but then maintenance was the painful part, and that's when we learned, okay, it was not just building it, etc. Inside of Codex and OpenAI, how do you see maintenance becoming cheaper changing, in terms of what you're building, what the ambition is, the, I guess, custom tooling, those kinds of things?
Maintenance is really sort of a tax that you pay over time just to keep things running, and it's always been necessary. It will continue to be necessary, but where I think it changes is a lot of it is just going to be automated. So it's like, okay, you have these third-party dependencies, you need to upgrade the version number. [laughter] You can fully automate this if you have a good changelog and the code is well documented and the model can just reason through it. You can just blast through your codebase and do it in a couple of hours, and previously you would have punted on it because it's not the most fun thing to do, but it's actually really important for your business. It's really important for your project, especially for security vulnerabilities. You want to stay up to date, right? You want to apply all these patches. I think that's just going to be fully automated. So a large part of maintenance just kind of comes for free, right?
And then I think it's awesome to also think about before, when you wanted to just completely... you had to do a new architecture because you were trying to make space for different kinds of trade-offs, or you had a new understanding of the workload, or you were trying to fit a new feature and suddenly you realized your current system is just very limiting and you need to completely rearchitect it. That was a really, really costly endeavor, right? Sometimes multiple years. And I think this is also super, super accelerated now. So the cost of mistakes, I would say, is going down. But then at the same time, the good old rules, I would say, of software engineering, of having good abstractions, really help. It's going back to this having the box with invariants: if you sort of draw the right shape, you're going to be able to change things much more quickly within the box and not affect the rest of the services or the rest of your infrastructure. And I think it's important to design for very quick iteration.
And I remember when I talked with Peter Steinberger, that was before he joined OpenAI, about OpenClaw and how he thinks about it, he told me that he doesn't read the code, but I could see that he's holding the architecture in his head, and he was telling me how he rearchitects a lot, and he thinks about how to make it modular, how to allow 100 contributors to each build their thing without stepping on each other's toes. So I'm hearing what you're saying: this care, this planning, this structuring has become maybe just a lot more important, which was something that back in the day the architect or the staff engineer or experienced folks were doing, and other engineers around them were kind of building smaller parts. But it sounds like now all engineers need to be aware of this when you're building your software, right, and plan for it.
Yeah, and the GPT models are getting better and better at this as well, of thinking about long-term maintenance and good architecture, and this is a natural next step, right? It's not just about code quality in the sense of, oh, is this code clean within this file, but is the architecture actually correct to reduce maintenance burden over time and make space for future product or feature extensions or changes, and just really this act of engineering over time. That's something that models are starting to become capable of thinking about very well. I think it's just kind of fascinating to understand that the software that we're building is just going through the life cycle much, much faster, right? Before, you were starting it maybe as a small team of yourself and maybe a couple of engineers, and then you would add engineers slowly, and then maybe after a year, if it's very, very successful, you would have 50 engineers on it or 100 engineers on it. You would have time to see it coming. You would have time to see the humans onboard, and you can think about the documentation, all of that stuff. But now it's sort of that explosion of, suddenly you have 100 agents contributing to this thing, and that can happen in a weekend. And so you're just going through it at major, major speed compared to before.
Okay, but how do you and the folks at OpenAI deal with this? Does it not mess with your mind? You know what I mean, in the sense that you've been in this business for quite some time now, like decades or well over, and there was a pace that we kind of got used to, and obviously it's now a lot faster. But how do you get your head around the fact that, A, it's faster, and B, the stuff that you've been doing a year ago you're not doing right now because the model is good at it? How do you kind of reconcile that? Because I'm sure there's stuff that you've been really good at related to software that now you can hand off to the agent. Do you not get a little bit of a sting? We talked about how it's stinging for your features to be implemented in open source. But it can also sting that I've been really good at, I don't know, refactoring, or right now it might be architecture, but maybe the model will be really good at that, and now I'm like, okay, damn, I'm glad, but also it would have been nice for me to do that.
Yeah, I think there's a craft aspect to it, which is why occasionally I still pull up an editor and write some code. And it feels nice, and I have fond memories of late nights sitting [laughter] in Vim, and just like...
Cranking it out.
Drinking Coke Zero, and yeah, just not having to think about anything else other than the problem in front of me. But really, I think it's all about being in the flow and solving problems. And what I find is folks here, and also everyone I talk to, are just adapting very quickly. And I think if you have a mindset where code is a tool to solve problems, you can solve so many more problems. Before, you wanted to benchmark something and you weren't quite sure where you were going to net out. You can just do it. It's going to take you no more than 30 seconds to launch something in the background and get proper numbers and be able to make a better trade-off. It should make you a better engineer if you just really care about the outcome and the system working well. And so what it allows us to do at OpenAI, it allows us to run our inference much more efficiently. It allows us to get much more effective compute and deploy that to the world. And so everyone's just very focused on that and solving important problems at a speed that was not possible before. And I haven't yet encountered someone who's like, "Oh, that's not good. That's not fun."
Do I understand correctly that it sounds like if you have ambitious problems, if you have way more problems than what you can solve today or tomorrow or next week, this is not really a problem, because when you get more efficient somewhere, you keep
going. Which is a lot of startups, right? Like startups are always way more ambitious than what they're able to do.
We're not out of problems for sure, right? And I don't think we will be for a while. We have a long, long road ahead of us in terms of mathematical breakthroughs, scientific breakthroughs, you know, making the world a better place, just really building for humans and solving the most important problems that everyone is facing, and just doing it in a deeply human way. That's what we're here for.
Also, just going back to coding and these late nights, I think it's also maybe a glamorous version of it. I also had a lot of late nights where I was trying to refactor something. And I would be three hours deep into the refactor and then realize, actually, this is a dead end, and I must restart from scratch, and it was very frustrating. And so there are these very, very fun times, but there's also the time where it doesn't compile and you're just like, why is this not compiling yet?
I'm sure you had the time where you go later, it's now super late, you need to go to bed because you need to get some sleep, and then you can't really sleep and you have this thing where you have some task that is halfway and it upsets you sometimes. I remember dreaming about the code as well. And I guess one thing I don't really have these days when I'm working on my software for my business is I don't really have something that is halfway, because I can just tell it do this and then I can leave it at a state where it's kind of done: it's either working or I have proof that it's failed. But it's interesting because everything's sped up, right?
Yeah. Maybe I do have... what a lot of people do and I do myself is I sometimes have bigger questions that I'm asking myself from conversations I've had during the day, or I haven't yet had the time to just look into it, and so I will send off Codex to just look at it overnight, and then I'm very excited to then wake up and look at the results. And so it's always an exciting morning.
Well, I feel there's an art to doing long-running tasks. And of course, you can use /goal, which will go and run. That's also something that was recently added, a few months ago, right? The /goal command to Codex.
Yeah. And back to maybe the harness is a crutch, right? /goal was necessary to keep the model on track on a singular goal for a very long period of time. And it allows the model to literally run for days or weeks if it's a really, really hard problem. But with the new generation of models, what we're seeing is you don't need /goal anymore. You don't need a harness around it. You can just tell the model, hey, go and work for a week, and it will actually do it.
Speaking of hard problems and the fact that you're not out of them. One of the interesting things that you shipped, from the outside, I would say, as an engineer, it was moderately interesting, is what you call the merge, which is Codex appeared inside of ChatGPT. And the reason I say that as engineers it's kind of moderately interesting is because we've been using Codex. Like, yeah, it's there, you can now open it in the ChatGPT app, great. I just went there and I just immediately went to Codex, because I don't really use ChatGPT in the app per se. But I talk with folks at OpenAI and people in your team, and they were telling me there was a lot of preparation going on, a lot of engineering challenges. Can you give a sense of how big this project was, what you needed to do, and why was it difficult to pull off, and how did Codex and other tools help you get it done in ways that would have been hard before? Because since you've launched the merge, the numbers that you keep sharing of how many people use Codex, it's going up way faster than before. So I assume there's a big scale problem you've solved here.
A lot of things were challenging with the merge. First of all, completely different stacks. ChatGPT is fully managed, cloud-based, like you run everything on our systems. We store things; the traditional way of building things. Built for scale, built for efficiency. Codex fully local. And so the merge is just really: how do you get the same benefits and the same capabilities from this local coding agent, and then build a product around it and build it in a way where it can benefit a much, much broader pool of people, which is also why all of us joined OpenAI, to benefit this very, very broad population across the world. And so it was a very exciting journey of figuring out how do we build a cloud version of this that in essence is capable of very much the same things, but is also built in a way where we can serve it to tens and hundreds of millions of users in a way that is still efficient, so that we can include it all the way into the Plus plan.
Work is essentially running the full Codex harness in a cloud, together with a cloud computer. It's a very powerful machine actually. People have sort of picked up on it and showed what you can do. If you are creative with the prompt, you can get ChatGPT to train another model in there.
Wow.
There are some pretty wild things. You can get it to install Blender and do 3D modeling. It's very permissive, it has internet access, it's a powerful machine, and then Codex just works on it. And this is what we ship through ChatGPT Work. A lot of system challenges.
The team did it very quickly. Obviously Codex helped to make it more efficient, and build a lot of the infrastructure, and then help resolve a lot of the little differences as well that had been occurring between Codex and ChatGPT, like merging plugins architecture, merging library, and so really working towards a unified system. Which is really the goal: you shouldn't feel like you can do something in Codex that you can't do in ChatGPT or vice versa. What we're trying to build is one unified product that gives you access to the same intelligence, but in the way that you want to use it.
And so it was very fun as well, because Codex throughout the whole journey also acted as a journalist to sort of document all the steps and the debates and the discussions that the teams were having. And it was very animated debate, of how we should do it and how we should name the thing and when to introduce it, in what way, and what to merge into what. There were many different permutations considered, and so there's a very fun journalistic element to it where we have a full recounting that Codex did over time. And it's kind of become known as well as the toggle arc of OpenAI, where we introduced the work toggle, which there was also a lot of debate around, whether this was the right thing, and then we kind of grew to just really like it. But over time we're going to merge things further. So we're really headed into this direction of full unification, and we kind of view this as a temporary state where you have better, stronger capabilities when you're in work mode, but over time we're bringing this all the way to everyone that uses ChatGPT.
And how do you personally use Codex? What's your working setup in terms of agents, in terms of tasks, in terms of what you manage with it? And related to this, I asked Peter Steinberger what I should ask about you, and he said you need to ask him: how do you deal with the fact that you're involved in all these projects? Your calendar is like Tetris, but usually you show up pretty cheerful.
My calendar is fine. And it's just, I am capable of doing so many more things nowadays because I have the technology like Codex, and I actually shifted a lot of my work onto mobile using ChatGPT Work, where whenever I have something that I want to take note of, I just fire that off. I use dictation a lot. Whenever I have a question, instead of writing it down to look into later or delegating to someone, I just fire it off in ChatGPT Work and I get a report. It has a whole bunch of custom skills and custom instructions where it's now very, very tailored to produce the kinds of reports and slide decks and code explorations in the style that I can consume effectively.
And so every time I'm between meetings, you'll kind of see me, I was just dictating to my phone. As I said before, we do a lot of work in public channels. We have a lot in Slack. We have a lot in Notion and Google Docs as well. And so there's pretty much no question really that I feel I cannot ask that Codex will be able to at least do a first pass of thinking through, whether it is public sentiment on a feature, looking at production logs for how much usage we have on a certain thing, making a list of things that we should deprecate because they're not getting traction, understanding what a certain team is up to. Any question I have, I can get an answer to within 30 minutes. And so that's how I use it. I use it for everything. It's like my personal agent in all the ways.
And then oftentimes on weekends as well, I do some code explorations or I build some prototypes, and I have fun sort of imagining the future of the product in some ways. And I do that with others on the teams. It's not always the same team. And in one day I can build things that I sort of had in my system, right? It's like I woke up one day and I was just like, we should explore what it means to build this. And then I can just sort of express all of that and get something in front of people in a day, so that they can think through it and criticize it and hopefully get inspired by it. It's by no means that we need to ship it, but it's more like, okay, I flush it out of my system and then I go on and do other things. So it's such a magical time and it's so empowering.
And as closing, what would your advice be for a software engineer/AI engineer, someone who builds software, who would want to get the skill set and the experience to have the opportunity to work at a place like the Codex team, like OpenAI, or like an AI startup? So just become this really great builder with these tools. Because the question that comes up is often: should I start with the theory? How important are the basics? Should I just get really good at using the tools?
Yeah. I think there are two things that are important. One is a deep, deep curiosity for how things work, and an ability to train yourself to understand things very quickly. It is the case that things will continue to change, but people that do extremely well at OpenAI are people that are able to grok a system quickly and also dive into a new codebase and make sense of it. But obviously all of that is helped with agents nowadays, right? There's so much information that you need to absorb, and being able to understand and reason through it. And a lot of that is asking good questions really about how do things work, and just going into the five W's, which I think you can just keep digging and digging and digging, and you're learning very, very fast through that.
The other thing is being in tune with the community, or the people that you're trying to solve a problem for. Not everything is solving a direct problem. Sometimes you're solving a problem that will be useful to another group of people in the pursuit of solving a problem for humans. But just being crisp about the taste or the needs or the requirements, and being able to think clearly and exercising this clarity of thought, feels really important to me. If you can't explain what you're trying to achieve, if you can't explain your intent, if you don't have a tie to a community, if you don't have the taste, it's going to be much harder to do great work.
Awesome, Tibo. Well, thanks a bunch for this conversation. This was awesome.
Thanks for having me.
I've always wanted to get together with Tibo, and I'm glad that we finally made it happen. I appreciated how Tibo talked about not just the upsides of open source, but also the downsides, most notably how competitors can copy features you are just working on in the open right now and then ship it right before release, and just how much this stings. Plus, you get a lot of low-quality contributions that you still need to somehow deal with.
Another interesting one was Tibo saying how the harness is always a step ahead of the model. From the inside, the Codex team see their job as building crutches for the model with the harness, the tools, and the setup instructions. And then the next version of the model will be trained to need fewer of these crutches. I'll be honest, as a dev, this sounds a little demotivating: the stuff I build, in the next version of the model, it'll just know, and we can get rid of it. Plus, I do suspect that it's not just about building these crutches, but also building tools that models will use. And it's not like the next version of the model will reinvent an MCP protocol or skills or plugins. At least I hope not.
I also enjoyed hearing what the merge, merging ChatGPT and Codex, looked like from the inside. It was merging a previously fully local coding agent, Codex, into a managed cloud-based stack, and doing it efficiently enough so that it can be included in OpenAI's $20 per month plan, when $20 is not all that much in terms of compute purchase. It was pretty amusing to hear how Codex itself acted as a journalist of the whole project, as it was present in all the Slack conversations and all the documents, and so it could capture all the important debates and decisions. I'm not going to lie, this part felt a little bit of a Big Brother feel to it, where the AI is always watching, but it could well become the new normal in startups in the future. I've not yet decided how I feel about this.
And finally, I appreciated Tibo's advice for engineers to succeed: be curious, understand systems quickly, and be in tune with the group you are building for. It's reassuring to hear from Tibo as well how much the fundamentals still matter.
Do check out the show notes below for deep dives on how Codex, Claude Code, and Cursor were built, and other related topics. If you like what you heard, please hit a rating on the podcast player that you're using. It means a lot to me and to the show. Thanks, and I'll see you in the next
Article published
