Shipping Before It's Perfect: OpenAI's Tara Seshan and Nan Yu on Agents, Planning Horizons, and What Comes After the Chat Box

Open on YouTube ↗
Overview

At the Lenny and Friends Summit, Claire Vo interviewed Tara Seshan and Nan Yu of OpenAI about building AI products while the underlying technology keeps shifting. The session covered shipping features they themselves consider imperfect, how to design agents that people can actually manage, where computer use fits alongside other integrations, how product managers work with researchers, how far ahead a product leader should plan, and which product form factors each panelist expects to matter in 2027. Across these topics, both speakers held that classic product principles still apply, but that empirical testing, speed, and closeness to users matter more than before.

17 min read
0:24

The toggle as a deliberately imperfect solution

Vo opened with what she called "the toggle": in the product the two work on, a switch with one side for one mode and the other side for another. She asked Seshan how the team decides to release something it knows isn't ideal.

Seshan said the shift had been personal as well as professional. She described herself as the child of Asian parents who expected perfect output every time, and said that after roughly seven years at Stripe she was used to a culture where everything was "deeply polished and deeply considered", down to whether an enum had exactly the right name. What changed, in her account, is that urgency now matters a great deal. Teams can theorize at length before shipping, but she said nothing compares to the empirical evidence of watching users try a product and then iterating. She has had to move from trying to get things perfect a priori to shipping and iterating as quickly and effectively as possible.

She said directly that the toggle is not the ideal solution. What mattered was putting an agentic harness, the ability to use an agent to get things done, in front of the more than a billion people who use ChatGPT, without disrupting developers' existing workflows. The toggle was the way to do both at once.

Deprecation and taking users along

Vo noted that she and Yu had talked before about how teams now build and throw away a lot of code and product, and asked how Yu decides what to keep and what to deprecate.

Yu used the toggle as an example. He said there is an obvious next step, which is no toggle, and a story that runs from step A to step B. In Yu's view, users accept changes and reversals much more readily if they are taken along on that journey, especially if they have some transparency into what is happening behind the scenes. As long as the story is coherent and users can follow it with the amount of attention they actually pay, Yu said, a team can change a lot without burning bridges.

Tara Seshan's quality bar

Asked what keeps quality high when things move this fast, Seshan listed several tests. The first is whether a feature is additive: does it unlock real value, is there something genuinely useful in it. The second is an internal bar. Products ship internally first and everyone uses them. She said there is no specific number, but the team looks at whether a product retains well, whether people find something delightful or surprisingly great in it, and whether it unlocks a new use case for the models.

She called the third test probably the most important in this era: aiming about two to three months ahead of where the models will be. Products should be neither too anchored in the present nor so futuristic they are unusable. She described the target, half-jokingly, as being "sufficiently AGI-pilled." In summary, her questions were: Is it valuable? Do users care? Is it surprisingly great? Does it fit model capabilities? And does it get out of the way so the model can do a great job?

5:16

Nan Yu: absorption is the binding constraint

Yu named a different constraint: how much people can understand and absorb. He pointed to the idea of capability overhang, where models can do far more than people take advantage of, and said that gap is the real limit. It doesn't matter what you put into a product, Yu argued, if people can't absorb the changes and the new capability.

5:51

Can enterprises absorb this pace?

Vo, who said she had spent some time in enterprise software, noted a common claim from customer success and sales teams: B2B customers can't absorb change. She asked whether that still holds.

Seshan said it is partly true. The current pace probably feels "beyond breakneck" to most enterprises, a flood of features and advances that may feel overwhelming. But she argued that if you don't ship the frontier to enterprise customers, you get leapfrogged. Her example: most of OpenAI's enterprise customers had been using chat, asking questions and getting answers, while the agent revolution was happening elsewhere. The team felt it had to get agents to those customers as fast as possible, or they might consider alternative products. So even though enterprises have their own pace for absorbing updates, OpenAI shipped what Seshan called a giant, revolutionary change that broke their processes and their assumptions about how updates should arrive, so they could get more value. She tied this to an old product truism: do what users need, not what they say they need. She said it applies now almost more than ever.

One agent or forty?

Vo then raised what she framed as a possible debate: single-identity agents versus many specialized ones. She contrasted her own setup, a Codex agent docked on her computer that she talks to all day, with the 40 bots she said are load-bearing across different parts of her business. Builders, she said, are choosing between one general agent identity that "gets out of the way of the models" and hand-crafted soul.md files for many small, job-specific agents.

Yu said the most successful designs map onto human nature. People have millions of years of evolution that shape how many things they can hold in their heads, and 40 agents is a lot. He compared it to managing 40 direct reports. Maybe someone like Jensen Huang could, he said, but most people would doubt they could track that many threads. So agents need to be grouped into bundles of activity. He observed that when you ask people how they manage their agents, a common answer is that they have a chief-of-staff agent that manages the others. Yu's response was that this is cheating: you aren't managing seven agents, you're managing one that runs your team. In his view, that dynamic shows up as soon as people are given this kind of freedom.

Seshan acknowledged the philosophical side but said the question is mostly tactical. How do data access and permissions work? If there is one entity, does it switch between a service account and the user's account, and how is that shown clearly to the user? Does a single agent need segmented memory? If it joins a private Slack channel, is it effectively a different agent there, "severed" in the sense of the TV show Severance? Should it use each person's credentials or a shared set? For her, the answer depends on the specific use case and on working through edge cases like how memory should behave and how the agent should talk to each person.

The skills agent design demands

Vo said the two answers reflected two sides of product craft. One is getting inside users' mental models. The other is systems and architecture thinking: working through the platform components and edge cases needed to deliver an experience sensibly. She asked whether other hard skills now matter more.

Yu agreed that systems thinking plus user empathy are the classic product management ideas, now applied to a different technology. The work means starting from scratch and rethinking how these products want to be shaped, but he said these are "just the old principles in disguise."

Seshan added a third element. She said she enjoys thinking abstractly, having come from Stripe, which she called a very academic company, and from an academic family. But she said what is also required is relentlessness: trying things again and again, taking a lot of pain along the way, and seeing what works. Yu added the ability to learn from those feedback loops, which he said has always been necessary but now runs at a much faster clock speed.

12:59

Platforms, ecosystems, and computer use as a fallback

Vo said she now loves computer use and would rather point it at a product's UI than bother with an MCP integration, even if it costs more tokens. She asked how the panelists think about the wider ecosystem of work tools.

Seshan described ChatGPT as a platform and said the team considers which functionality belongs natively in the platform and which first- or third-party developers should build in as plugins or additional interfaces. For a third-party meetings app, for example, the question is which hooks to expose so the app works well for someone using Codex or computer use. She rejected a binary between "in our platform" and "in someone's app." Instead she described layers. A plugin might pull in all of a user's meeting data and work out next steps. If a system doesn't interface well, computer use is the next layer, a fallback when pre-built or ecosystem tools aren't available. The goal is to expose hooks for developers, make the resulting experiences composable, and make sure that if the first layer fails, there is a second layer that lets the user finish the task.

Vo then asked how anyone builds a high-quality, well-branded experience from a non-deterministic model, a natural-language interface, a computer that can click anything, and third-party tools the company doesn't control.

Yu said the key is the difference between finishing the whole job and finishing everything except the last mile. He argued the last mile can feel worse than a non-starter. If something fails outright, you do it yourself or try another route. But if it gets 99% of the way and then "barfs," it is a worse kind of unfulfilled promise. What feels great about computer use, in Yu's view, is that it always works. It may be slow or use many tokens, but it gets the job done all the way while the ecosystem catches up with the right hooks and MCPs.

Seshan invoked Charles Eames's idea that a user should feel like a guest in your home, with the host having anticipated what they want and provided it gently. She said users of the app should feel their needs have been anticipated, and that increasingly the model itself is what does that. Vo joked that every time she visits that "home," it tells her it's time to update.

17:28

Working with research

Vo said collaborating with research teams is new for most product people and asked what Seshan had learned. Seshan said it was entirely new to her when she joined OpenAI and offered what she called a novice's view, comparing herself to the Connecticut Yankee in King Arthur's Court. Her first lesson was that working with research is quite different from working with engineering.

Her approach is to bring very specific use cases and a clear picture of what users are trying to accomplish, along with sample readings: the actual session, what it looked like, and why the model didn't do what it needed to. She said learning to write good evals, and writing as many as possible, is the most important thing. The ideal is to show that with a particular prompt and set of skills she got the desired outcome from the model, and then take that to post-training and ask how to train the capability in from the start. She stressed that it begins with recognizing this is a wholly different way of working, then getting good at bringing specific user data, and, once researchers trust you enough, writing the evals that start the iteration loop.

19:27

How far into the future to plan

Asked what time horizon a product leader should work on, Seshan joked that she lives in the now, then said she does have to live a little in the future: ideally two to three months out. Planning in terms of years is a mistake, she said, because those predictions are almost always wrong. She keeps a personal document of things she predicted that didn't happen, as a reminder that she can't forecast years ahead. But building only for today means getting left behind.

Vo polled the audience about who was doing 2027 annual planning and called the gap between that and a 60-to-90-day horizon the fundamental tension. Seshan said it depends on the business. She contrasted Stripe with OpenAI. In her view, the payments market looks largely as it did before, accelerated in some ways, so you know the forces and can model bull, bear, and base cases. She said she can't do that for OpenAI's business, joking that perhaps Sarah Friar could. Her advice was to understand your market's pace and dynamics and choose a planning cadence to match. When Vo concluded that Stripe PMs still have to do annual planning, Seshan agreed and wished them luck.

21:48

What still matters, and what's newly table stakes

For the closing "then and now" segment, Vo asked what persists from classic PM craft and what has become newly essential.

Yu said the emphasis changes, and, given absorption rates and capability overhang, onboarding matters far more than before. He noted that onboarding is often seen as a niche area many engineers don't love working on, but said real success depends on designing the initial experience and continually helping people understand what the product offers.

Seshan pointed to users' understanding of privacy, of how their data is used, and of the product's norms. Before, she said, products could get away with unpredictable behavior because a settings menu or popup could explain it. Now that agents act semi-autonomously, a user needs to be able to predict at any moment what a feature or new primitive will probably do. On PM craft generally, she said her answer might sound boring: all the usual things still apply, including being in the details, holding the high level and low level together, and deeply understanding the design and the product. What has increased is the importance of testing empirically, moving faster, and testing with users.

23:54

The "DMable PM"

Vo observed that OpenAI's product and engineering team is highly accessible in public Slack. She said she regularly sends them feedback IDs with complaints, such as her agent Astra having "an attitude." She asked whether direct user relationships are becoming table stakes.

Seshan said her background in developer and enterprise products meant she had always been reachable by users asking for features. What feels new is offering the same accessibility on a huge consumer product. To her, direct contact with users is simply normal, and it connects to her earlier point: a PM's value to research lies in knowing specifically what users want. She said she had been thinking Vo should send her those Astra outputs, because that is exactly the kind of information that matters.

Yu compared this to dev tools, where engineers are picky and detail-oriented and often don't give useful feedback until the second or third follow-up question. He said many PMs now feel that change, because problems with natural-language agentic experiences are subtle. A complaint that an agent has an attitude needs unpacking: what exactly did it say, what was the context, did the user provoke it? A deeper, more direct relationship with users pays off there in ways that may not have been possible before.

26:14

Predictions for 2027

Vo closed by asking what would be big in products in 2027, noting that new form factors seem to emerge every week or month.

Seshan answered: voice. She said voice has changed how she works and feels much more natural than typing. She said it has greatly reduced the tech support she does for her own family, and that when she was onboarding OpenAI's people team onto a new product, voice "changed the game" because it is so intuitive for so many people. She linked this to Yu's point about human nature: speaking is how people have evolved to interact. She said the models are finally getting good at voice and getting faster, and called herself very bullish on it.

Yu predicted something that looks a lot like self-driving software. He described the "empty input box problem": whether it's a general tool like ChatGPT or an industry-specific one, users get the product and then face the question of what to do with it and how to learn it. Since these products are now, as Yu put it, literally intelligent, they can operate themselves to help users and provide a gentle on-ramp. He expected this to become so normal that people will look back and wonder why it didn't always work that way.

Vo offered her own prediction: hardware that replaces carrying an open laptop, something like "a little Codex in a box." She asked the note-takers in the room to serve as accountability buddies for checking which of the three predictions about 2027's defining product proved right.