Building Codex: Tibo Sottiaux on Rust, Open Source, Harnesses, and Cheap Code Changes

Open on YouTube ↗
Overview

Tibo Sottiaux was on the team when Codex started and has led the broader Codex team since. On the Pragmatic Engineer Podcast, the host asked how OpenAI's coding agent was built and how it keeps changing. They covered why the CLI is written in Rust and released as open source, why it works with other companies' models, how the harness and the model share the work, and what happens to code review, maintenance, and architecture when changing code becomes cheap. Throughout, Sottiaux argues that much of the scaffolding engineers build around models is temporary, and that the durable skills are clear intent, good abstractions, and staying close to users.

28 min read

From a Belgian village to applied mathematics

Sottiaux traces his interest in computers to his parents' decision to leave Brussels for a village of about 200 people. At around eight years old, with few people nearby he wanted to spend time with, he turned to computers and the early internet to learn about things. He credits the move with leaving him "no choice" but to get interested in computers.

He studied applied mathematics, finished early, and ran small consulting work for banks and on supply chain problems while still a student. He says the appeal of applied mathematics was taking sophisticated theory and using it to improve the world, rather than doing mathematics for its own sake. After university he co-founded a startup that optimized pharmaceutical supply chains for clinical trials. It decided how much medicine to produce and where to dispatch it to reduce waste. The company used non-ML techniques: Monte Carlo simulation and stochastic multi-stage optimization. It later applied the same methods to the steel industry and the European electrical grid. He says the company still exists and is now changing a lot because of modern AI.

Google: a canceled project, Maps, and DeepMind

Sottiaux joined Google in London in 2015. His first project was not Maps. It was an effort to make websites faster on mobile, run by a small group inside the ads organization. The goal was to offset ad revenue lost as traffic moved from desktop to mobile. He worked on it for about two years and calls it the most fun he has had solving hard technical problems. Then a VP flew in from California and canceled it: the product had only hundreds of users, which was "clearly not Google scale."

He says people on the team were surprised, and that the surprise was itself the lesson. The project lacked product-market fit, the right users, and the right feedback loop, and the team had trusted a product manager who said things were going well when they were not. The lesson he took was to keep questioning, both about his own impact and about whether the overall project matters.

He then worked on reviews in Google Maps for about a year before moving to DeepMind. He describes DeepMind then as a special place, with early rumblings around AlphaGo and a focus on the hardest problems. Given his background, he was drawn to it. There he worked on research infrastructure and tooling, a theme he says he carried for almost a decade: building tools and products that make other people more efficient.

A ChatGPT-like bot inside DeepMind, a year early

The host asked about something Sottiaux had posted on X: he had helped build an internal Google chatbot roughly a year before ChatGPT. He explains that DeepMind's main focus was grand challenges, games, and reinforcement learning outside the language setting. Google Brain, separate at the time, had its own language model work. One group inside DeepMind pushed the hypothesis that scaling language models on large text corpora might be enough to reach general intelligence, which he says was hotly debated.

Because Sottiaux was building research tooling, the question arose of how to present these models to researchers so they could inspect inputs and outputs. That naturally ended up as a chat interface. The early models were "a little bit absurd," not very coherent or useful, but fun to tinker with. The application spread "like wildfire" inside DeepMind as people shared conversations, and it began to feel like more than a tool for researchers. There was a desire to launch it externally, but he says DeepMind was not set up for that. Google has a blessed production stack and a well-optimized process for launching products, which he credits for doing things well but describes as a very hard environment for real innovation.

Why he left for OpenAI

Sottiaux says he was comfortable at Google. What pulled him away was the chance to join a mission he believed in, with people who wanted direct impact, instead of a setup where research happens in one place and making it useful is someone else's job. He wanted research and product to be co-designed. When he met people from OpenAI and learned that only about 20 people worked on ChatGPT, he was struck by how empowering that must be at that scale.

He joined just before the reasoning launch. About a month after he arrived, OpenAI released o1-preview, which he calls exhilarating. He says the values he took from that period, which he later applied to Codex, were moving fast, caring about impact, having a community, and running an intense feedback loop with it.

How Codex began

Sottiaux joined OpenAI in 2024 and started, as before, by building research infrastructure. With o1-preview and later models, he says it became clear that OpenAI had to use its own models to go faster. He became focused on what was limiting that. With colleagues in research, he began training models and building small agents that he calls the true precursors of Codex. Those internal models were trained to be very good at OpenAI's Python codebase and to have good taste in architecture and code style. They were Python-only, and the aim was to build infrastructure quickly and help researchers code faster.

According to Sottiaux, Greg and Sam were very supportive, and Greg insisted the work should benefit the world and not only OpenAI, which meant turning it into a product. At that point the research effort merged with the internal "A-SWE" (autonomous software engineer) project. A sprint followed that produced the first cloud version of Codex. Sottiaux says it did not have product-market fit because it involved too much friction. The Codex CLI launched as well, and the team kept iterating on how to make models genuinely helpful.

Why the core agent is written in Rust

The host pointed out an odd choice: Codex's core is written in Rust, although models at the time were stronger in Python and TypeScript, and most competing harnesses used those languages. Sottiaux says the decision followed from treating the product interface and the agent as separate things from the start. The agent core had to be robust, secure, and built for efficiency and scale. From earlier projects that grew from "fun thing" to "scale of the largest data center," he had learned that early decisions matter a great deal, as long as they don't cost too much velocity.

In his account, the team had very strong Rust developers, OpenAI's internal models were "not bad at Rust," and Rust's compile-time static verification is useful for agents too. So it quickly became clear to them that Rust would work well for agents if they invested in it. The main drivers were correctness and efficiency.

He also says Codex could probably have succeeded in TypeScript or even Python and simply been rewritten later. For him, the real value was the clean separation between the agent, which should exist independently of any product, and the product layer. With everything in one codebase and one language, he says, you inevitably get "a little bit sloppy," entangle things, and block later innovation. The Rust boundary helped enforce that separation.

Open source: the benefits and the sting

Codex's CLI, SDK, and app server are open source, which the host noted is unusual among the major labs. Sottiaux gives several reasons. Codex is a coding agent, so the team expected people to point it at its own code and hoped a community would use it to improve it. The team also believed that if they succeeded, open source and the role of code would change, and they wanted to be inside that community rather than separate from it, because problems are hard to solve without seeing them firsthand. And since it was early, they had ideas about what a good harness looks like, co-designed with training and research, but not all the answers. So they published technical deep dives, wanted to learn from other open source projects, and tried to create a level playing field that encouraged tinkering.

When the host asked for an honest look at both sides, Sottiaux listed benefits first. New hires have usually already read the repo and its PRs, and onboarding largely means exploring the repo with Codex. The project gets many good contributions. The team also gets energy from contributing to the community directly, not just saying they care about it, even though it costs effort they didn't have to spend.

He named three downsides. First, the open code is separate from the rest of OpenAI's code, which sometimes forces artificial boundaries and work across multiple repositories. Second, when the team builds something exciting in the open, others sometimes copy it before OpenAI releases it. He says that is "part of the game" under a very permissive license, but it "does sting a little bit." Third, like other maintainers, they face a "tsunami" of random contributions and the tax of handling them, though he says this also pushes them to find solutions.

Supporting other companies' models

Codex also isn't tied to OpenAI models. Sottiaux says coupling an excellent community harness to your own model felt wrong and "quite disappointing." Since the code is open, anyone could fork it and add another provider in about ten lines. Forcing people onto a fork just for that seemed silly, so the team supported other providers directly.

He adds that optionality helps OpenAI as well. A user who wants to try a new model doesn't have to rebuild their setup, and OpenAI gets feedback on what that user liked or disliked about it. The team also tests other models in the same harness. He says companies OpenAI works with care a lot about this optionality, and OpenAI leans into it. He wants to win users with the best and most efficient models and the best product, not through lock-in. He argues that forcing people to use a product because of one constraint would not attract the best people to work on it, and that the experience should feel delightful regardless of which model powers it.

Where the agent runs: sandbox, local machine, and cloud

By default, Codex runs entirely on the user's machine inside a sandbox, and has for more than a year. Any command that needs permissions beyond the sandbox requires the user's approval. Users can instead run tasks in the cloud on a managed VM, the same one used by ChatGPT's work mode, isolated in a Kata container. In that case only the input and the streamed output touch the local machine, which saves local CPU and allows much more scale.

Sottiaux calls this one step. He expects running on cloud machines to become much more seamless, possibly with some execution local and some in the cloud. His reasoning is that more capable models can use far more compute than a laptop has, so limiting them to local execution would eventually become a constraint.

The host objected that local execution is valuable because their tools and a local Postgres database are already there, and asked about cloud development environments, which were popular around 2022–23. Sottiaux says that outside large tech companies, cloud dev boxes never took off because of the high upfront setup cost and ongoing maintenance, which solo developers and small teams don't recoup. Now that agents are capable enough, he argues, that setup is "almost free": the agent can replicate your local database, servers, and MCP configuration on a cloud box and keep them in sync. He predicts a resurgence of fully cloud-orchestrated machines that free people from their laptops. He points to ChatGPT's work mode on mobile, which he says he uses each morning over coffee to dictate tasks, with access to his calendar, email, and Slack. He notes that Codex remote still executes on the laptop, and says it would be wonderful not to have to keep the laptop open.

The harness is "a little bit ahead" of the model

The host recalled that early Codex didn't run their unit tests after making a change, and a few months later it started doing so automatically. They asked whether that came from harness changes or model changes. Sottiaux says the harness is always "a little bit ahead of the model." A model can do certain things, and the harness gives it "crutches" so it does them reliably, efficiently, and the way users expect. The harness supplies guardrails and safety, makes the model more steerable, and writes the developer message injected into context at the start of each turn to shape the agent's behavior.

Test running is his example. At first you might have to remind the model to run tests. Then a better model is trained that reflects more on what the user actually wants, and the reminder is no longer needed. Over time, he says, both the developer message and the harness shrink.

On how the team sets goals, Sottiaux describes co-design between research and the engineering team that builds the core agent harness. For each gap or desired feature, they ask whether it should be a harness change or a model change. If it's a model change, they ask whether it can land in one, three, or six months. Depending on the answer, they may do nothing in the harness and wait for the model. They also use agents to analyze feedback and surface themes. That analysis covers coding and other domains where people now use these agents, including finance, communications, and marketing, along with subcategories. Better pre-training and better overall models lift everything, but some areas get extra attention.

He gave a concrete example later in the conversation. The /goal command, added a few months earlier, was a harness crutch that kept the model focused on a single goal for very long periods, allowing it to run for days or weeks on hard problems. With the new generation of models, he says, /goal is no longer needed: you can tell the model to work for a week and it will.

"Have you asked Codex?": how work gets done on the team

Asked what a new Codex team member is told about how things work, Sottiaux says he introduces them to great people, and the answer they hear most often to any question is "Have you asked Codex?" Inside OpenAI, Codex is connected to Slack, documents, and code by default, and new starters are still surprised that they can ask it almost anything and usually get a good answer, including the state of a project, who is working on something, or why a decision was made. For that reason the team works in public channels and shares documents with broad permissions so that both people and their agents can reason over the information. He mentions additional team-collaboration tools that are not yet released, some of which are planned for DevDay.

His general guidance to newcomers is to care about the user, the coherence of the product, and where the models are heading. If you find yourself writing a 10,000-line crutch to work around a model's flaws, he says, you are probably doing the wrong thing. There is a set of principles, but mostly it is team culture passed on by colleagues.

On shipping, Sottiaux says the process is surprisingly similar for Codex and for ChatGPT, even though ChatGPT reaches about a billion active users. An engineer can make a change and have it shipped the next day or even the same day. People are empowered to make large changes, with a strong expectation of ownership. What's asked for is evidence that the change will be well received, that it's worth adding, and that it's worth maintaining. He adds that maintenance costs have dropped substantially, so the team weighs this differently than two or three years ago. Code review, deploys, and catching regressions are largely automated, so engineers can focus on the idea and its value to users.

He also described the North Star: a delightful, simple personal AGI that knows what it needs to about you, has access to the right resources, and can take sometimes risky actions on your behalf, with a push notification so you can verify them. It would know your schedule and goals, be proactive, and be controllable through natural language and voice, perhaps even understanding your emotions from a camera feed. It should not be "a thing with 10 different buttons and configurations." In his words, "AGI should be simple to use."

How code review is changing

The host noted Google's strict review culture and the traditional benefits of review: knowledge sharing, a second pair of eyes, reducing bus factor, and architecture discussions. They asked where human review still adds value. Sottiaux described one of his early Codex projects: working with research to build a code review model that could catch logic and reasoning errors that would take a human hours to find. These are errors that require digging three or four levels into dependencies, for example discovering that a third-party library's documentation is wrong, that its implementation differs from what you assumed, and that your invariants therefore don't hold. Unless you're an expert in that library, you would miss it.

Those review models were released, and he says their capability is now part of the mainline models, which OpenAI's benchmarks show as superhuman at code review. He says this applies to security as well. Automated security review is now mandatory on all of OpenAI's pull requests, and PRs flagged with a security issue are blocked from merging automatically.

His view is that review always served two purposes: correctness, and a social ritual for sharing information and getting people aligned, sometimes the only place a discussion happened, because merged code runs in production and must be maintained. He expects correctness and security review to be automated. What remains is discussion of intent: what are you trying to do, and should you be doing it? He argues that conversation doesn't need to happen around code. The pull request was a forcing function because it was the last point before code reached production, but intent can be settled in other ways, "and then the code doesn't matter as much."

He described a "box" model. Teams agree on what a component does and which invariants it must meet, with strict guarantees around resource use, data access, and security. Inside those limits, the implementation "could be literally anything" and doesn't need discussion. That agreement deserves a good conversation, perhaps with an agent's help. Once it exists, changes inside the box no longer need your attention. The host agreed that reviews had always been mixed, with some real learning and a lot of chasing reviewers and context switching, and said one upside is not having to spend attention on basic things.

Maintenance and re-architecture get cheap

Sottiaux calls maintenance a tax paid over time to keep things running. It remains necessary, but he says much of it will be automated. His example is upgrading a third-party dependency: with a good changelog and well-documented code, a model can reason through the change and apply it across a codebase in a couple of hours. Teams used to postpone that kind of work even though it matters, especially for security patches. He says a large part of maintenance "just kind of comes for free."

He also highlighted re-architecture. Redesigning a system because of new trade-offs, a better understanding of the workload, or a feature the current design can't handle used to be very expensive, sometimes taking years. He says that is now greatly accelerated, and the cost of mistakes is falling. But he argues the old rules still hold: good abstractions, "boxes" with invariants, and designs that let you change internals quickly without affecting other services.

The host connected this to a conversation with Peter Steinberger, creator of OpenClaw, before he joined OpenAI. Steinberger said he doesn't read the code, yet clearly held the architecture in his head, re-architected often, and designed for modularity so a hundred contributors could work without colliding. The host suggested that this kind of structural thinking, once reserved for architects and staff engineers, now matters for every engineer. Sottiaux agreed and said GPT models are getting better at it too: not only clean code within a file, but whether the architecture reduces maintenance burden and leaves room for future features. He says models are starting to handle this kind of "engineering over time" well.

He finds it fascinating that software now moves through its lifecycle much faster. A project used to grow from a couple of engineers to 50 or 100 over a year or more, giving time to onboard people and write documentation. Now, he says, a project can suddenly have a hundred agents contributing to it "in a weekend."

Does it sting to hand off the craft?

The host asked whether it bothers Sottiaux to watch models take over skills he spent years building. He admits a craft element: he still occasionally opens an editor to write code by hand, and has fond memories of late nights in Vim drinking Coke Zero, thinking only about the problem in front of him. But he says what matters is being in flow and solving problems. If code is a tool for solving problems, you can now solve far more of them. A benchmark you weren't sure about can be started in the background in about 30 seconds and give you real numbers for a better trade-off. At OpenAI, he says, this means running inference more efficiently, getting more effective compute, and deploying it to the world. He says he hasn't yet met anyone who thinks this is bad or not fun.

He also pushed back on romanticizing the old way: there were many late nights three hours into a refactor that turned out to be a dead end, or wondering why something wouldn't compile. The host said they no longer leave work half-finished before bed, because a task ends either done or with evidence that it failed. Sottiaux said he often sends Codex to investigate bigger open questions overnight and looks forward to the results in the morning. Asked whether anyone runs out of problems, he said OpenAI is "not out of problems," with a long road ahead in mathematical and scientific breakthroughs and in building for people "in a deeply human way."

The merge: bringing Codex into ChatGPT

The host said "the merge," Codex appearing inside ChatGPT, looked only moderately interesting from the outside. But people at OpenAI had described a large engineering effort, and Codex usage numbers had grown much faster since. Sottiaux says the core difficulty was two completely different stacks. ChatGPT is fully managed and cloud-based, built traditionally for scale and efficiency. Codex was fully local. The task was to deliver the same capabilities as a local coding agent in a cloud product that could serve tens or hundreds of millions of users efficiently enough to include in the Plus plan.

The result, ChatGPT's work mode, runs the full Codex harness on a cloud computer. He calls it a powerful, permissive machine with internet access. He says users have discovered that with creative prompts they can get ChatGPT to train another model there or install Blender for 3D modeling. He says the team did it quickly, with Codex helping build infrastructure and reconcile many small differences between Codex and ChatGPT, such as merging the plugin architecture and libraries. The goal is a unified product: nothing you can do in Codex but not in ChatGPT, or the reverse, with the same intelligence accessed however you prefer.

He also described Codex acting as a "journalist" throughout the project, documenting the steps, debates, and discussions, which he says were lively: how to do it, what to name it, when and how to introduce it, and what to merge into what, with many permutations considered. Codex produced a full account over time. Internally the project became known as OpenAI's "toggle arc," after the work toggle, which was itself debated before the team came to like it. Sottiaux says the toggle is a temporary state in which work mode offers stronger capabilities, and the plan is to unify further and eventually bring those capabilities to everyone who uses ChatGPT.

How Sottiaux works day to day

The host relayed a question from Peter Steinberger: how does Sottiaux stay cheerful when his calendar looks like Tetris? Sottiaux replied that his calendar is fine and that tools like Codex let him do much more. He has moved much of his work to mobile through ChatGPT's work mode and uses dictation heavily. Instead of writing a question down for later or delegating it, he sends it off and gets a report. He has built up custom skills and instructions so the output (reports, slide decks, code explorations) comes in a form he can absorb efficiently, and he often dictates to his phone between meetings.

Because so much of the team's work lives in Slack, Notion, and Google Docs, he says there is essentially no question he can't have Codex take a first pass at. His examples include public sentiment on a feature, production logs showing how much a feature is used, a list of features to deprecate for lack of traction, and what a particular team is working on. He says he gets answers within about 30 minutes. On weekends he builds prototypes and explores ideas about the product's future, often with different people from the teams. He can take an idea he woke up with and put something in front of colleagues within a day, so they can critique it and ideally be inspired by it. Shipping it is not the point. It gets the idea "out of my system" so he can move on.

Advice for engineers who want to work in AI

Sottiaux named two qualities. The first is deep curiosity about how things work, combined with the ability to understand things quickly. He says people who do extremely well at OpenAI can grasp a system fast and make sense of a new codebase. Agents help with that now, but much of it still comes down to asking good questions and continuing to dig. The second is being in tune with the community or people you're solving problems for, including cases where the direct beneficiary is another group that is itself serving people. That means being crisp about taste, needs, and requirements, and thinking clearly. If you can't explain your intent, lack a connection to a community, or lack taste, he says, it will be much harder to do great work.

The host's reflections

In closing, the host highlighted Sottiaux's candor about open source's downsides: competitors copying features before release, and the burden of low-quality contributions. The host found the idea that the harness is a set of crutches the next model will no longer need a little demotivating as a developer. They suspect the work also involves building tools models will use, and hope future models won't reinvent things like MCP, skills, or plugins on their own. They noted that the merge meant bringing a fully local agent into a managed cloud stack efficiently enough to fit a $20-per-month plan. On Codex serving as the project's journalist by reading every Slack thread and document, the host said it had a "big brother" feel and could become normal at startups, and said they have not yet decided how they feel about it.