Inside OpenAI's Codex Team: Agents as Teammates, Moving Bottlenecks, and What Comes Next
The Pragmatic EngineerAt The Pragmatic Summit, Tibo Sottiaux, Head of Engineering for Codex at OpenAI, and Vijaye Raji, OpenAI's CTO of Applications, described how software development is changing inside the company. The host asked what engineering work looks like at OpenAI now, how new graduates fit in, what happens to product managers and designers, how teams without unlimited compute should think about cost, and where things go next. Both speakers argued that building software has changed fundamentally. They also said that foundations, product sense, and human judgment still matter, and that the constraint keeps moving from one part of the process to another.
From tool to extension to agent to teammate
Raji, who had been at OpenAI for about six months, said the way the company writes software "has fundamentally changed," and that the change has been visible even within those six months. In his description, Codex went from a tool, to an extension, to an agent, and now to a teammate. He expects engineers to start naming their agents and treating them as teammates.
He gave two examples of scale. On internal leaderboards of Codex usage, some engineers routinely use hundreds of billions of tokens a week. Much of this runs in parallel across many agents rather than one. He also described Codex Box, a tool released internally the week before the talk. It lets engineers reserve dev boxes on servers and send prompts to them, so the work runs remotely while the engineer orchestrates from a laptop. People can close their laptops, go to a meeting, and find the work finished when they return. Raji predicted that within a few months this would be normal in Silicon Valley and then spread further.
The host said that a year earlier this would have sounded like a fairy tale, but he now uses these tools himself. From conversations with OpenAI engineers, he also reported that not everyone there writes all of their code with Codex. Usage has grown across the company but varies by team, and the Codex team is further ahead than the rest.
How the Codex team works: chasing the bottleneck
Sottiaux said the Codex team reinvents its own way of working "almost on a week-to-week basis." The method is to find each bottleneck and remove it, knowing another will appear. The bottleneck was first code generation, then code review. Now it is mostly about understanding user needs faster. That includes triaging tickets and taking in what users say on Twitter, Reddit, and other channels, then turning that into strategy. The team uses agents heavily for all of this.
He told a story that showed how the job is changing. A candidate negotiating to join the Codex team asked how much compute they would get to build products. Sottiaux had never thought about a per-employee compute budget, which he associated with researchers training frontier models. He took the question as a sign that people now see compute as a way to multiply their own output. His view was that someone with good taste, good ideas, and the ability to build software can now do a great deal.
Product engineering still starts with human intuition
Asked how the work of product engineers is changing, Raji said the basics stay the same: "we're still building products for humans to use." Someone still has to imagine the product and keep adjusting it until it is right, and he doesn't expect that to change while software is built for people. He speculated that if software is someday built for agents, agents might become the product engineers and product managers.
What has changed, in his view, is speed, and he said that makes the work more fun. He has been coding with Codex himself. On a recent flight without access to remote dev boxes, he kept his laptop partly open when told to close it so the agent wouldn't stop. The host said everyone now carries their laptops half-closed. Raji's point was that the loop of building, testing, verifying, and going back to Codex is much shorter, and so the reward comes sooner.
New engineering practices: parallel implementations and blurred roles
Sottiaux described two practices that looked odd at first and now make sense. The first concerns technical trade-offs. The old approach was to write a design doc, discuss alternatives, and discard all but one. Now engineers often build several implementations in parallel and pick the one that works best in practice.
The second is that roles are blurring. Sottiaux said Codex team designers now ship more code than engineers shipped six months ago. He attributed this to model quality: the code is good enough that the team would merge it as is.
Raji added an example from Sottiaux's team, which edits video files. Few people remember command-line syntax for tools like ffmpeg. Codex can take a plain description of the task, write the command, and run it.
After coding, the next bottlenecks
Raji said Codex's use has already spread from writing code to code review and security review. He expects a sequence of new bottlenecks. If coding makes every engineer, say, five times more productive, there will be much more code, and review becomes the bottleneck. After that, integration and deployment through CI/CD will be the constraint. He described this cycle of solving the next problem as exciting.
Overnight runs and self-testing
The host asked about something Sottiaux had described privately: overnight runs and self-testing. Sottiaux said it's easy to think of these models as "autocomplete on steroids" that finish a small feature in ten minutes. In his experience the model does much more when given a large task and can run for hours. The team has built an environment and set of skills that let Codex test itself autonomously. They run it overnight as a QA loop that flags regressions.
He also described a researcher on the team who trains models. That researcher said that every time he thought he was more capable than Codex, he found he had prompted it badly or set it up wrong. Sottiaux called this "both exciting and a little bit depressing." According to Sottiaux, Codex now trains a model independently and writes a short PDF report with its findings. The team reads the report, picks the most promising directions, and feeds them back into Codex.
Codex in the meeting room
Sottiaux described two meeting practices. In the weekly analytics review, which covers feature adoption, retention, and the funnel, the team starts by listing questions the dashboards don't answer. The data analyst then starts a Codex thread in the background for each one, and answers come back in about 20 minutes. The team discusses them in the last 10 minutes of the meeting, often for five or six questions per session. Sottiaux compared it to having small consultants working in the background.
The second practice is incident response. When someone is paged, Codex helps work out what went wrong and the fastest path to recovery. Sottiaux said this greatly increases how much information the team can gather and how fast it can fix problems.
New grads and AI-native engineers
The host raised a common industry worry: will junior engineers be squeezed out if seniors can use AI agents? He noted that OpenAI's head of engineering had told him the company is hiring early-career engineers.
Raji confirmed OpenAI is hiring many new graduates and has a strong internship program this year. He believes new engineers will be AI native and able to use these tools from their first day, and said giving them the opportunity matters. This summer brings OpenAI's first batch of new grads, about 100 people, and he wants to keep growing the internship program.
Onboarding on a flat team
Sottiaux said he runs Codex as a very flat organization with 33 direct reports. He does this so he won't be the bottleneck. He thinks leaders are tempted not to change their org structure as fast as their people can now build, and that one person approving every decision "is just like obviously not going to work anymore."
Onboarding starts with Codex. New people ask it questions, use it to navigate the codebase and see what others are working on, and get daily reports. The people responsible for onboarding and team culture are the ones who joined most recently. Sottiaux said a new grad named Ahmed, who joined about six months earlier, is "absolutely crushing it," which surprised him a little. He credited Ahmed's energy and speed and joked that his own brain is probably already in decline.
Playing devil's advocate, the host said experienced engineers have seen how important foundations were for new grads who grew into strong professionals. He asked whether people who start with AI coding and skip years of groundwork will build the right foundations. Sottiaux said foundations remain very important. The team designs the codebase and architecture carefully, still does code review, and does not let Codex write everything unchecked. In his experience, new grads absorb this well. With the right codebase structure and guardrails, they are very productive, so the work is in setting up that environment and planning how the codebase will evolve.
Foundations and the history of abstraction
Asked what a software engineer does day to day now, Raji also said "foundations will never go out of fashion." He drew on 25 years in the industry, including work on developer tools at Microsoft, where he wrote the Visual Studio editor and language services. He remembered how exciting IntelliSense was when first shown. The host recalled developers at the time saying you weren't a real developer if you used IntelliSense.
Raji placed that in a longer pattern. Before, people said you weren't a good engineer if you didn't write assembly, then the same was said about C++, and later people complained about JavaScript. He doesn't think those arguments matter. What matters is strong foundations, product intuition, knowing what you're building, and being able to move up and down the stack to solve problems. He expects that to stay true.
Product managers and designers
On PMs and designers, Raji returned to his earlier point: as long as products are for people, human designers and product managers are needed, and he doesn't see a substitute for product or design sense. He said these roles are becoming more productive. PMs and designers write code, turn designs into prototypes or production, and validate ideas before bringing them to engineers. PMs also use Codex for PowerPoint slides, and OpenAI has Excel plugins, so gains extend beyond engineering.
Spreading new practices: Slack channels, hackathons, and show-and-tell
Asked about internal knowledge sharing, Sottiaux said OpenAI is discovering what the technology can do at about the same time as everyone else. When something starts to work, they ship it, so they have only a short time with "more of the crystal ball." That makes it important for good ideas to spread quickly. The Codex Slack channel and a "hot tips" channel are very active, and the company runs regular hackathons and show-and-tells. He said there's "no one true way" to use these tools and the period is still one of discovery.
His example was Alexander, the single product manager for the whole Codex team. For a one-hour bug bash on features about to ship, Alexander had Codex collect everyone's feedback into a Notion doc, file bug reports and improvement tickets in Linear, assign them, and follow up with each person on progress. Sottiaux described him as becoming a "10x, 50x" program manager this way, and connected it to the bottleneck theme: a team shouldn't let its product manager become the bottleneck.
Raji added that at demo days and hackathons, the demos keep getting deeper. They no longer just show what's possible. Many now handle corner cases and are close to usable products, even when built only to show a capability.
The cost question
The host added a disclaimer: inside OpenAI, everyone has effectively unlimited tokens. Elsewhere, cost matters. Subscriptions run out and people move to paid credits. He asked what teams with budgets should do.
Raji said OpenAI thinks about cost constantly and wants to offer more capable models. He expects thinking to shift toward treating an agent as a teammate that works 24/7, can be assigned Linear or Jira tasks, and is expected to complete them. The question then becomes how much you'd pay that teammate, not how many tokens you use. If each engineer effectively has four or five such teammates, he argued, the spending makes more sense. He also said users should hold OpenAI responsible for making agents capable enough to be treated as teammates.
Sottiaux suggested thinking about how costs shift across a company. Some tasks are now very cheap, such as market research or reviewing an entire feature backlog to find what's easy to implement. He said that work might once have needed about 15 engineers and is now nearly free. He acknowledged that not every company can offer unlimited inference, but said limiting it too early is also a risk, since it is still very early in understanding how much leverage people can get. His recommendation was to give a company's best people generous amounts of inference.
Has anything felt this fast before?
Asked whether anything in his 25 years compared, Raji said, "I don't think I've ever seen anything like this." He listed the dot-com bust during college, Y2K, the mobile revolution, and the social network era, which he worked in. He said this one feels different in both scale and speed, to the point that "some of these charts don't make sense." He called it special and said it's good to be living through it.
Predictions: speed, multi-agent networks, and debugging by symptom
Asked to predict software engineering and engineering management two years out, Sottiaux said two years was far too long a time frame and gave a six-month view. He is confident of another order-of-magnitude gain in speed, which will change things again. He also expects large networks of agents working together on big goals to become practical. He cited Cursor's demonstration of asking agents to build a browser from scratch and getting something like two million lines of code 24 hours later. A codebase that size is nearly impossible for a person to understand. His expectation is that people will set guardrails on what gets built so they no longer need to read the code. Either correctness will be provable in some way, or the system will be constrained enough to be secure and judged by inputs and outputs. Code would then be abstracted away, and the work would center on the system's properties and the real problems it solves.
Raji framed this as a continuation of rising abstraction, which lets people build a lot of product with little code, but said the rate of change is now much faster. He raised one concern. Any sufficiently complex system is harder to debug, so people end up debugging from symptoms. He expects software to reach the point where it has so many layers that engineers, and their tools, will need to get very good at diagnosing problems from symptoms. He sees that as a distinct skill developers will need.
Sottiaux added one more prediction. Instead of monitoring 100 or 200 individual agents, people will have one dedicated personal assistant they can call to check on progress, representing the work of all the agents running in the background. He expects to see this fairly soon, possibly within the year.
So what a time to be alive, VJ.
Yes.
Can you tell us, a question that a lot of us are asking, what is happening inside OpenAI right now? More specifically, when it comes to building software, with how engineers are doing stuff and how the whole thing is changing.
I'm glad you clarified that. Lots is happening. I've been there for about six months, and one of the things that I've learned is there is so much to learn from the kind of research that's happening in the company. And just to project out what the possibilities are is just mind-blowing. So I'll tell you this, right: the way we write software has fundamentally changed. It's changed so dramatically, and even in the last six months I've seen us go from Codex as a tool to an extension to an agent, now to a teammate. I fully expect engineers to name their agents now and call them their teammates, and this is happening so fast.
I was looking through some of the leaderboards of the people that are using Codex internally, and some of the engineers routinely hit hundreds of billions of tokens every week. And this is not just one agent. We're talking about, last week we released Codex Box internally, which is a way for us to actually reserve dev boxes on the server and fire off prompts, and it's doing the work, doing the job, while you're on your laptop orchestrating all of this stuff. And then people shut down their laptop, go to a meeting, come back, and all of the work has been done. So this is happening in parallel. This is how fundamentally software has changed internally at OpenAI, and I can't wait for all of this in the center of Silicon Valley and then to expand further and further. In a few months, I think this will be the norm. Everybody is going to be developing software this way, and that's pretty cool.
So if I would just take myself back, you know, six months, even a year, and I would hear you say this, I would think, oh, it's like a magical fairy tale, you're making half of this up. However, actually a lot of us are using it. I'm using it. I'm seeing what's happening. And I've been talking with engineers inside of OpenAI. I love talking with engineers because they have hashtag no filter. This is the secret to part of why The Pragmatic Engineer works: I talk with engineers who don't have this thing called media training or all that, and they just tell me how it is. And inside of OpenAI, one thing that was pretty comforting to me, I'll be honest, is not all engineers are writing 100% of their code with Codex. They're all using it a lot more, but there are levels there. One team that is absolutely on the cutting edge, though, and again I've talked with a bunch of engineers, is the Codex team, and they're even ahead of others inside OpenAI. So Tibo, you're leading the Codex team. Can you tell me how the Codex team works today and what the typical workflow of an engineer is like right now, as of yesterday or this morning?
Right. It's a fast-evolving situation. The thing that's delightful about how the Codex team operates is that they're constantly reinventing how they're working, almost on a week-to-week basis. And the thing that we go after is, we identify every single little bottleneck, and the bottlenecks keep shifting. So it used to be code generation, and then it moved to code review, and now it's very much: how do we understand the user needs faster? How do we triage tickets? How do we figure out what everyone is saying on Twitter, Reddit, all the important surfaces, and synthesize that into a strategy? And everyone is trying to leverage agents for that to the very best effect. And an interesting thing is, the other day it was the first time in a negotiation, someone was trying to join the Codex team, and this person asked me, how much compute am I going to get to build products at OpenAI? I was like, huh, that's an interesting question. I mean, we do have a lot of compute, but I haven't really thought about it, so it's like a compute envelope per employee. Usually that's more reserved for researchers who are actually training really phenomenal models. So I think there's this shift there where people realize you can hyper-leverage yourself in all sorts of novel ways, and if you do have great taste, great ideas, you know how to build software, it's like, what a time to be alive, really. It's just incredible what you can do.
And taking a little bit of a step back, outside of the Codex team: VJ, you have a lot of visibility inside of OpenAI. How is the work of a software engineer, or should I say a product engineer, changing? OpenAI has always hired software engineers who are product engineers, very clearly. How is their work changing? How are things morphing with product, or are they not morphing?
Fundamentally, we're still building products for humans to use. And so there's a lot of product intuition that comes into play. Even when, so I've been messing around with Codex thanks to the new app that makes it even more accessible for everyone to start coding, even in a lot of cases, we have to imagine the product that we have to build and ship, and that's where it starts, and then you have to constantly tweak it to get it to the right place. I don't think that's going to change, as long as we continue to build software for humans. I mean, at some point in the future we may build software for agents, but then maybe the agents will become the product engineers or product managers at that point. But I think the velocity makes it a lot more appealing and compelling and actually more fun, to be honest. I was coding on a plane, and at that time I didn't have access to the dev boxes, so you kind of keep the laptop open. When the flight attendant comes over, like, you have to shut down your laptop. No, no, but I don't want to have the agent stop, so I keep it slightly open and then put it down.
Everyone just runs around with their laptop, you know, half closed right now. It's like
Yeah, what are we doing? I think, you know, I actually think it's more fun now building software, because the gratification cycle is so much shorter, and it's so cool to see the product that you're building, test it, verify it, and then go back to Codex.
And as engineers, what are new, different, or weird engineering practices that you're now starting to see that, you know, start to make sense even as weird?
It used to be that you had difficult technical trade-offs, and you'd do a design doc and discuss it all, and then maybe you're like, oh, what are the other viable alternatives, and then you discard that. I think a delightful thing is that now I see people explore multiple different implementations all in parallel, and then we can actually zoom in on the one that we prove to work better. The other thing is I also see roles blur. So our designers are shipping more code than engineers were shipping six months ago. And that's also because the models have become sufficiently good that the code that they're producing is actually code that we would want to merge just as is.
Do you see anything else in the broader OpenAI? I have noticed, I don't know, do you all remember the command line for every one of the command line tools you use? I don't want to pick on one. I was talking to Tibo,
his team edits video files, and, you know, if you know ffmpeg, I don't think anyone remembers the command lines. Codex is such a great tool for that: okay, well, I want to do this, and then it crafts the command line and goes and executes it. So those are kind of new ways that we're seeing people use Codex specifically. I also think that we've now moved on from just coding to code reviews, security reviews, and then, as Tibo said, we're going to find more bottlenecks. So once you solve coding, for example, now you've just made every engineer five times more productive. What's going to happen is there's going to be more code being written, which means that code reviews will become the bottleneck, and then after code reviews, integrations and deployment, CI/CD, will become the bottleneck. So we're going to have to constantly go solve the next set of problems, which is really exciting actually.
Then Tibo, one really interesting thing when we talked about what you're doing at Codex that I've never heard before is these overnight runs and the self-testing. Can you tell us about that, because that is net new?
Yeah, I think it's easy to get stuck in, oh, this is autocomplete on steroids, and it's just going to implement a little feature, and sure, it will get done in 10 minutes. But what we're seeing is that the model is actually much, much more capable if you give it a very large task. It's capable of running for multiple hours. So we've assembled the environment and the skills so that Codex can fully autonomously test itself. We run this overnight so that it basically performs QA in a loop and flags regressions. The other thing was, I keep talking to this researcher on the team who's actually training the models, and he's like, every time I think I'm more capable than Codex, I figure out I'm wrong and I just didn't prompt it right, or I hadn't set it up in the right way. And this is both exciting and a little bit depressing at the same time. Because he's like, oh, now it's just training a model fully independently and writing a little PDF report at the end with its own insights and findings, and then we just take that and find the most promising things to iterate on and then just put that back into Codex. And so these very, very long-running tasks and achievements, it's incredible to see a model do this independently.
Yeah. And one more thing that we talked about that felt to me a bit like it's from sci-fi: you said that sometimes the Codex team has meetings about Codex and issues that you have, and you told me something interesting, that people get together in a meeting room and then you fire off Codex threads to diagnose stuff with Codex. Can you tell us a little bit about how that's playing out, because that is really like a loop of itself.
Yeah, there are two big things that we do there. So we have this weekly analytics review where we go over feature adoption, retention, we analyze our funnel, and we always start the meeting with questions we have that are just not answered in our dashboards, or that we haven't looked into, where we're just like, oh, this looks interesting. And then our data analyst is just like, okay, let's fire off a little Codex thread in the background, it will come back in 20 minutes, we'll have the answer, and we can talk about it in the last 10 minutes of the meeting. And then we do that for five, six questions that people have in the room, and it's this magical experience where you just have these little consultants working for us in the background. And the other thing is, whenever we get paged on call, Codex is there helping figure out what went wrong, what is the fastest path to recovery. And there it just feels so much accelerated, how much information we can gather and how quickly we can solve for things.
So this is one, and it's absolutely accelerating, right, and we see it elsewhere as well. One big question that keeps coming back across the industry is: what about new grads? What about junior engineers? And when I was talking with the head of engineering at OpenAI, he was saying something interesting, that you are hiring early-career engineers. This is great to hear. Can you talk about how it's going, what you're seeing with them? How founded are the fears that juniors are not great because now seniors can just use AI agents? And how are they getting up to speed?
We are hiring a lot of new grad folks straight from college. We're also having, so this year we have a pretty robust internship program. I actually truly believe that the new software engineers that are being created are going to be AI native. They're going to know these tools in a native way, and they're going to be able to leverage our AI tools from day one. And I think giving them the opportunity is going to be critical and important, and growing them in this kind of environment is going to be amazing. I can't wait to see this. And so this summer is our first batch of new grads that are going to be coming into OpenAI, and I'm really excited for that. It's going to be about 100 people or so. And then I want to continue growing our internship program within OpenAI. So yeah, this is going to be a really, really cool thing to witness in this age.
And then Tibo, how are you onboarding people to the Codex experience specifically? Even within OpenAI, my sense is that the Codex team is maybe a few months or weeks ahead of how others are working. When someone new, either from the outside or even from OpenAI, comes in, how do they get up to speed on how the team works?
So I run the team as a very flat organization. I have 33 direct reports on the team, and they just run around and do cool things, and I don't want to be the bottleneck. I think this is one of the things where, as leads, it's very tempting to not change organizational structure fast enough for how quickly people can actually build, and a single person being the bottleneck on every single decision is just obviously not going to work anymore. But the first thing that people get introduced to, obviously, is Codex itself, right? So Codex is responsible for the onboarding. You just ask Codex questions, you navigate the codebase, understand what other people are doing, you receive daily reports. But then the people who are responsible for the onboarding and the culture and how we build are also the people that most recently onboarded onto the team.
And I find that actually, just talking about the new grads, I have this phenomenal new grad who joined the team six months ago, and he's absolutely crushing it. And that was a little bit of a surprise, but I understood this person has unbounded energy, much more than I do, and is just super, super quick. I think my brain is probably already in decline; this person, Ahmed's, brain is just absolute peak. Just a phenomenal person, and he's been so successful on the team, and that's been really delightful to see.
Now, playing a bit of devil's
advocate, a lot of us more experienced folks who have seen new grads grow into really successful professionals, we have seen that, at least up to now, foundations were so important. And so what do you think will happen if we have new grads whose foundations are using AI coding and they probably skipped the stuff that we did for 10, 20 or more years? Are they building the right foundations, or are we asking the right question here?
Even foundations remain super important, right? So we take great care in designing the overall codebase, just taking care of the overall architecture. We do code review, as you said. We don't fully rely on Codex writing everything and just closing our eyes and being like, this is going to be fine. We have the very best engineers working on this as well. But I find new grads are able to absorb that, and if you have the right structure for your codebase and you set the right guardrails, then they're incredibly productive. And so I think it's just about the environment that you're setting up, and thinking ahead of time of how this codebase is going to evolve.
And how is the role of software engineers changing compared to even six or eight months ago? What does a software engineer do? If you had to explain to a new joiner, and they're going to ask, "Hey, Vijaye, what am I going to do day-to-day?" What are they going to do?
Yeah, I think the idea of foundations, foundations will never go out of fashion. So that is always going to be important no matter what. I think we're all here because we have strong foundations; that's brought us here. And then in terms of the role of a software engineer, it's changed quite a bit. I may be dating myself, 25 years in the industry, I've seen so many paradigm shifts. And I actually worked on developer tools at Microsoft, wrote the editor for Visual Studio and language services. So the first time I saw IntelliSense, that was a really cool moment where you could type, hit the dot, and then the options showed up.
Yeah. But do you remember, I was joining the industry around that time and the devs around me were saying, you're not a developer if you use IntelliSense.
Yes. And I've seen those. I mean, this is probably before my time, when people probably said, okay, if you're not writing assembly, you're not a good software engineer. And then C++, and then the abstractions kept going up and up, and then people used to complain about JavaScript. Remember those days?
I don't think those things actually matter. The point is that as long as you have the strong foundations, as long as you have product intuition, know what you're building, and are able to go up and down the stack to solve problems, those are going to be the more important ones. And I don't think that'll ever go out of fashion. I feel like that is always going to be the case.
We're here among mostly engineers and engineering leaders, but let's just spare a thought for product managers and designers. How do you see their roles changing, especially now that both engineers and they can build features a lot faster? How does that change their roles? Are we getting closer, or do they still have a distinct role from what you see?
I go back to this: as long as we're building products for humans to use, we will need human designers, we will need human product managers. I don't know that there is a substitution for product sense or design sense. Those things will evolve, will get even more productive, even more abstractions, but we will continue to evolve that. They're getting more and more productive, if anything. So product managers are writing code, designers are writing code, they're taking their designs into production, into prototypes, and validating them before they come to engineers. So I think those are already getting a lot more productive. Product managers are also using Codex for building PowerPoint slides, and we have Excel plugins, and so it's all around. It's not just engineers; everyone around is getting more productive.
One cool thing that you're doing inside OpenAI, which I've heard about, is this internal knowledge sharing, this show and tell where teams show what they do. Can you tell us how you came up with it, how you're actually doing the mechanics? And can you tell us some cool things that you've seen teams show and maybe other teams adopt?
Yeah, it's interesting because we're discovering the technology and evolving it, and we're co-evolving with it as well. So just as all of you are discovering, hey, this is what AI can do for me, and this is what it means for the organization, or this is what it means for my project, we're also discovering it pretty much at the same time. As soon as we have something that feels like it's starting to work, we ship it to the world, right? So we have a very small amount of time where we actually are able to have more of the crystal ball than all of you. And it's super important that good ideas diffuse very fast through the organization. So we use Slack, and the Codex Slack channels and hot tips are two channels that are super, super active. And then we organize regular hackathons, show and tell. We just try to diffuse novel ways of working with AI as fast as possible, and it's a highly creative time. So I think there's no one true way to use this stuff. It's very much still in discovery.
And then we have this phenomenal product manager on Codex, Alexander Embiricos, and he's the single product manager for the entire Codex team, and he hyper-leverages himself with the help of Codex. The other day he organized this bug bash. It was an hour; people were going through features that we were about to ship. And then he sent Codex to collect feedback from everyone, this ended up in a Notion doc, and then he dispatched Codex to file bug reports and feature improvement tickets into Linear, and then assign them to everyone, and then follow up with everyone on how it was going. And so he's becoming a 10x, 50x program manager just by leveraging AI as well. And I think it's important, again going back to the bottlenecks, you need to keep going back: your product manager cannot become the bottleneck. So you need to look at it in a principled way.
One thing I'll add is, I've been to these demo days and we've seen a whole bunch of these projects being demoed. I remember going to these hackathons and looking at the demos. One thing I'm noticing is the depth of these demos has been consistently going up. So it's not just a surface-level, here's what is possible. Some of these demos are actually, here's what's possible, but also I've taken care of all of these corner cases, and it's actually a very usable product. So the depth, day by day, of all of these products that people are building, even just to show off some of the capabilities, is definitely going down, going up and getting deeper.
One kind of disclaimer that we need to add is that inside OpenAI everyone has access to unlimited tokens. There's no cost. And people are laughing because it's kind of a big deal, right? In the outside world, if you will, cost is still a problem. You get the Max subscription, and when it runs out you're now on credits. Some people are cool with it, especially founders, but sometimes people ask questions with this in mind: a lot of places are constrained by cost, just for practical purposes. What suggestions and tactics would you have for folks who are inspired by how the team at OpenAI works, but they have these constraints/handcuffs to work with?
Cost is something that we constantly think about. One is, obviously, we want to make our models more and more capable and offer that to our users. And then I also believe that at some point the thinking will shift, because now you should imagine you have a teammate that is working for you 24/7, and you can send instructions to your teammate. You can assign Linear tasks or Jira tasks to your teammate, and you should fully expect your teammate to be capable of taking care of those things. And then the question becomes how much will you pay this teammate, not necessarily how many tokens you are going to use. And so if you start to measure in terms of the productivity of every engineer having a team of four or five of these teammates, then it starts to make a lot more sense. Now, you should hold us responsible to make these agents capable enough to treat them as teammates. And that's what we're working on.
Yeah. I think it's also useful to think about how it displaces costs across the company. There are things that you can do now that are actually very cheap for you to do, like doing marketing research, or going over the entirety of your feature backlog and figuring out which ones are the ones that you can trivially implement. Before that, you would have needed to allocate maybe 15 engineers to go and look through that backlog, and now it's almost free. Obviously not everyone can provide the perk of having unlimited inference to their employees, but I do think limiting it prematurely is a risk as well. We're at very, very early stages of how well-leveraged people can get. And so I would definitely be saying, hey, for the best people at your company, give them very, very comfortable, large amounts of inference.
Reflecting on the pace of change, we know it's fast, and it's getting really, really fast. It feels like that. But taking a step back to your times before OpenAI, and Vijaye, you've been in this business for a long time, more than 25 years. Looking back, what was a time when change also felt fast? And did we see anything somewhat comparable in the past?
I don't think I've ever seen anything like this. I can look back on the 25 years. I've seen the dot-com bubble burst, and that was during my college time. And then I remember Y2K, I remember the mobile revolution, and I was actually part of the social network revolution. And this one feels very different. This one is happening at a massive scale, and also happening very fast. The speed at which this is happening, some of these charts don't make sense. And so I do think this is something very, very special and unique, and it's also cool to be living in this period.
Now, as a closing question: it changes fast, but the two of you have been at OpenAI for quite some time now, so I'm going to ask you to make an honest prediction. In two years' time, what do you think software engineering will look like, and what will engineering management look like, just knowing what you know?
Obviously two years is way too long of a time frame. I think six months from now, the things that I feel very confident saying are that we will get maybe another order of magnitude on speed, and that will change things again. And the other thing that we will get working is large networks of multi-agents that can collaborate together on very, very big goals. For example, it should be within the realm of feasible to say, along the same lines as what Cursor demonstrated, hey, rebuild a browser from scratch, just go, and then 24 hours later you have this thing that was built, two million lines of code. It's pretty much intractable to understand what actually is happening under the hood. And so there, I think what we'll start seeing is we will set guardrails around what is getting built, so that you don't actually have to look at the code anymore, and you can either prove that it's correct in some way, or it is constrained in a way where it is secure and you can just look at the inputs and outputs. And then code will become abstracted away, and it will all become about what the actual challenges are and the properties of the system.
Software has been increasing in abstraction, which makes it easier for us to go build massive amounts of product with very little code. So over the years that abstraction has increased, and I feel like we're in a time frame where that abstraction is increasing, and the rate of change has also increased quite rapidly. At some point, I worry, I'll say this right there, because any sufficiently complex or sophisticated system becomes harder to debug, and so you rely on symptoms to debug these things. And so I think in a few years we'll get to the point where software is so complex, software has gotten so many layers in it, and we get really good at identifying issues by looking at symptoms, and our tools are going to get really good at that too. And so I think that will be a unique ability for software developers to pick up.
Well, Vijaye,
I want to add something to what the future will look like. I think very much you will just be able to call your assistant and check on the work as well, and you will have one dedicated personal assistant that is able to represent the work of all the AI agents that are productively doing things for you behind the scenes, instead of having to monitor and check in with a hundred or 200 individual little agents. I think that's something that we'll see fairly quickly, including this year.
Yeah. Well, thanks so much to Vijaye and Tibo for giving us a peek at what is actually happening inside and how your teams are working, which it feels is either months or weeks or sometimes longer ahead of the curve, but it is happening. And also just what we might or might not see in this really exciting time. Thank you so much.
Thank you. Thank you.
Article published
