Inside OpenAI's Codex Team: Agents as Teammates, Moving Bottlenecks, and What Comes Next

Open on YouTube ↗
Overview

At The Pragmatic Summit, Tibo Sottiaux, Head of Engineering for Codex at OpenAI, and Vijaye Raji, OpenAI's CTO of Applications, described how software development is changing inside the company. The host asked what engineering work looks like at OpenAI now, how new graduates fit in, what happens to product managers and designers, how teams without unlimited compute should think about cost, and where things go next. Both speakers argued that building software has changed fundamentally. They also said that foundations, product sense, and human judgment still matter, and that the constraint keeps moving from one part of the process to another.

16 min read

From tool to extension to agent to teammate

Raji, who had been at OpenAI for about six months, said the way the company writes software "has fundamentally changed," and that the change has been visible even within those six months. In his description, Codex went from a tool, to an extension, to an agent, and now to a teammate. He expects engineers to start naming their agents and treating them as teammates.

He gave two examples of scale. On internal leaderboards of Codex usage, some engineers routinely use hundreds of billions of tokens a week. Much of this runs in parallel across many agents rather than one. He also described Codex Box, a tool released internally the week before the talk. It lets engineers reserve dev boxes on servers and send prompts to them, so the work runs remotely while the engineer orchestrates from a laptop. People can close their laptops, go to a meeting, and find the work finished when they return. Raji predicted that within a few months this would be normal in Silicon Valley and then spread further.

The host said that a year earlier this would have sounded like a fairy tale, but he now uses these tools himself. From conversations with OpenAI engineers, he also reported that not everyone there writes all of their code with Codex. Usage has grown across the company but varies by team, and the Codex team is further ahead than the rest.

How the Codex team works: chasing the bottleneck

Sottiaux said the Codex team reinvents its own way of working "almost on a week-to-week basis." The method is to find each bottleneck and remove it, knowing another will appear. The bottleneck was first code generation, then code review. Now it is mostly about understanding user needs faster. That includes triaging tickets and taking in what users say on Twitter, Reddit, and other channels, then turning that into strategy. The team uses agents heavily for all of this.

He told a story that showed how the job is changing. A candidate negotiating to join the Codex team asked how much compute they would get to build products. Sottiaux had never thought about a per-employee compute budget, which he associated with researchers training frontier models. He took the question as a sign that people now see compute as a way to multiply their own output. His view was that someone with good taste, good ideas, and the ability to build software can now do a great deal.

Product engineering still starts with human intuition

Asked how the work of product engineers is changing, Raji said the basics stay the same: "we're still building products for humans to use." Someone still has to imagine the product and keep adjusting it until it is right, and he doesn't expect that to change while software is built for people. He speculated that if software is someday built for agents, agents might become the product engineers and product managers.

What has changed, in his view, is speed, and he said that makes the work more fun. He has been coding with Codex himself. On a recent flight without access to remote dev boxes, he kept his laptop partly open when told to close it so the agent wouldn't stop. The host said everyone now carries their laptops half-closed. Raji's point was that the loop of building, testing, verifying, and going back to Codex is much shorter, and so the reward comes sooner.

New engineering practices: parallel implementations and blurred roles

Sottiaux described two practices that looked odd at first and now make sense. The first concerns technical trade-offs. The old approach was to write a design doc, discuss alternatives, and discard all but one. Now engineers often build several implementations in parallel and pick the one that works best in practice.

The second is that roles are blurring. Sottiaux said Codex team designers now ship more code than engineers shipped six months ago. He attributed this to model quality: the code is good enough that the team would merge it as is.

Raji added an example from Sottiaux's team, which edits video files. Few people remember command-line syntax for tools like ffmpeg. Codex can take a plain description of the task, write the command, and run it.

After coding, the next bottlenecks

Raji said Codex's use has already spread from writing code to code review and security review. He expects a sequence of new bottlenecks. If coding makes every engineer, say, five times more productive, there will be much more code, and review becomes the bottleneck. After that, integration and deployment through CI/CD will be the constraint. He described this cycle of solving the next problem as exciting.

Overnight runs and self-testing

The host asked about something Sottiaux had described privately: overnight runs and self-testing. Sottiaux said it's easy to think of these models as "autocomplete on steroids" that finish a small feature in ten minutes. In his experience the model does much more when given a large task and can run for hours. The team has built an environment and set of skills that let Codex test itself autonomously. They run it overnight as a QA loop that flags regressions.

He also described a researcher on the team who trains models. That researcher said that every time he thought he was more capable than Codex, he found he had prompted it badly or set it up wrong. Sottiaux called this "both exciting and a little bit depressing." According to Sottiaux, Codex now trains a model independently and writes a short PDF report with its findings. The team reads the report, picks the most promising directions, and feeds them back into Codex.

Codex in the meeting room

Sottiaux described two meeting practices. In the weekly analytics review, which covers feature adoption, retention, and the funnel, the team starts by listing questions the dashboards don't answer. The data analyst then starts a Codex thread in the background for each one, and answers come back in about 20 minutes. The team discusses them in the last 10 minutes of the meeting, often for five or six questions per session. Sottiaux compared it to having small consultants working in the background.

The second practice is incident response. When someone is paged, Codex helps work out what went wrong and the fastest path to recovery. Sottiaux said this greatly increases how much information the team can gather and how fast it can fix problems.

New grads and AI-native engineers

The host raised a common industry worry: will junior engineers be squeezed out if seniors can use AI agents? He noted that OpenAI's head of engineering had told him the company is hiring early-career engineers.

Raji confirmed OpenAI is hiring many new graduates and has a strong internship program this year. He believes new engineers will be AI native and able to use these tools from their first day, and said giving them the opportunity matters. This summer brings OpenAI's first batch of new grads, about 100 people, and he wants to keep growing the internship program.

Onboarding on a flat team

Sottiaux said he runs Codex as a very flat organization with 33 direct reports. He does this so he won't be the bottleneck. He thinks leaders are tempted not to change their org structure as fast as their people can now build, and that one person approving every decision "is just like obviously not going to work anymore."

Onboarding starts with Codex. New people ask it questions, use it to navigate the codebase and see what others are working on, and get daily reports. The people responsible for onboarding and team culture are the ones who joined most recently. Sottiaux said a new grad named Ahmed, who joined about six months earlier, is "absolutely crushing it," which surprised him a little. He credited Ahmed's energy and speed and joked that his own brain is probably already in decline.

Playing devil's advocate, the host said experienced engineers have seen how important foundations were for new grads who grew into strong professionals. He asked whether people who start with AI coding and skip years of groundwork will build the right foundations. Sottiaux said foundations remain very important. The team designs the codebase and architecture carefully, still does code review, and does not let Codex write everything unchecked. In his experience, new grads absorb this well. With the right codebase structure and guardrails, they are very productive, so the work is in setting up that environment and planning how the codebase will evolve.

Foundations and the history of abstraction

Asked what a software engineer does day to day now, Raji also said "foundations will never go out of fashion." He drew on 25 years in the industry, including work on developer tools at Microsoft, where he wrote the Visual Studio editor and language services. He remembered how exciting IntelliSense was when first shown. The host recalled developers at the time saying you weren't a real developer if you used IntelliSense.

Raji placed that in a longer pattern. Before, people said you weren't a good engineer if you didn't write assembly, then the same was said about C++, and later people complained about JavaScript. He doesn't think those arguments matter. What matters is strong foundations, product intuition, knowing what you're building, and being able to move up and down the stack to solve problems. He expects that to stay true.

Product managers and designers

On PMs and designers, Raji returned to his earlier point: as long as products are for people, human designers and product managers are needed, and he doesn't see a substitute for product or design sense. He said these roles are becoming more productive. PMs and designers write code, turn designs into prototypes or production, and validate ideas before bringing them to engineers. PMs also use Codex for PowerPoint slides, and OpenAI has Excel plugins, so gains extend beyond engineering.

Spreading new practices: Slack channels, hackathons, and show-and-tell

Asked about internal knowledge sharing, Sottiaux said OpenAI is discovering what the technology can do at about the same time as everyone else. When something starts to work, they ship it, so they have only a short time with "more of the crystal ball." That makes it important for good ideas to spread quickly. The Codex Slack channel and a "hot tips" channel are very active, and the company runs regular hackathons and show-and-tells. He said there's "no one true way" to use these tools and the period is still one of discovery.

His example was Alexander, the single product manager for the whole Codex team. For a one-hour bug bash on features about to ship, Alexander had Codex collect everyone's feedback into a Notion doc, file bug reports and improvement tickets in Linear, assign them, and follow up with each person on progress. Sottiaux described him as becoming a "10x, 50x" program manager this way, and connected it to the bottleneck theme: a team shouldn't let its product manager become the bottleneck.

Raji added that at demo days and hackathons, the demos keep getting deeper. They no longer just show what's possible. Many now handle corner cases and are close to usable products, even when built only to show a capability.

The cost question

The host added a disclaimer: inside OpenAI, everyone has effectively unlimited tokens. Elsewhere, cost matters. Subscriptions run out and people move to paid credits. He asked what teams with budgets should do.

Raji said OpenAI thinks about cost constantly and wants to offer more capable models. He expects thinking to shift toward treating an agent as a teammate that works 24/7, can be assigned Linear or Jira tasks, and is expected to complete them. The question then becomes how much you'd pay that teammate, not how many tokens you use. If each engineer effectively has four or five such teammates, he argued, the spending makes more sense. He also said users should hold OpenAI responsible for making agents capable enough to be treated as teammates.

Sottiaux suggested thinking about how costs shift across a company. Some tasks are now very cheap, such as market research or reviewing an entire feature backlog to find what's easy to implement. He said that work might once have needed about 15 engineers and is now nearly free. He acknowledged that not every company can offer unlimited inference, but said limiting it too early is also a risk, since it is still very early in understanding how much leverage people can get. His recommendation was to give a company's best people generous amounts of inference.

Has anything felt this fast before?

Asked whether anything in his 25 years compared, Raji said, "I don't think I've ever seen anything like this." He listed the dot-com bust during college, Y2K, the mobile revolution, and the social network era, which he worked in. He said this one feels different in both scale and speed, to the point that "some of these charts don't make sense." He called it special and said it's good to be living through it.

Predictions: speed, multi-agent networks, and debugging by symptom

Asked to predict software engineering and engineering management two years out, Sottiaux said two years was far too long a time frame and gave a six-month view. He is confident of another order-of-magnitude gain in speed, which will change things again. He also expects large networks of agents working together on big goals to become practical. He cited Cursor's demonstration of asking agents to build a browser from scratch and getting something like two million lines of code 24 hours later. A codebase that size is nearly impossible for a person to understand. His expectation is that people will set guardrails on what gets built so they no longer need to read the code. Either correctness will be provable in some way, or the system will be constrained enough to be secure and judged by inputs and outputs. Code would then be abstracted away, and the work would center on the system's properties and the real problems it solves.

Raji framed this as a continuation of rising abstraction, which lets people build a lot of product with little code, but said the rate of change is now much faster. He raised one concern. Any sufficiently complex system is harder to debug, so people end up debugging from symptoms. He expects software to reach the point where it has so many layers that engineers, and their tools, will need to get very good at diagnosing problems from symptoms. He sees that as a distinct skill developers will need.

Sottiaux added one more prediction. Instead of monitoring 100 or 200 individual agents, people will have one dedicated personal assistant they can call to check on progress, representing the work of all the agents running in the background. He expects to see this fairly soon, possibly within the year.