Shipping Before It's Perfect: OpenAI's Tara Seshan and Nan Yu on Agents, Planning Horizons, and What Comes After the Chat Box
Lenny's PodcastAt the Lenny and Friends Summit, Claire Vo interviewed Tara Seshan and Nan Yu of OpenAI about building AI products while the underlying technology keeps shifting. The session covered shipping features they themselves consider imperfect, how to design agents that people can actually manage, where computer use fits alongside other integrations, how product managers work with researchers, how far ahead a product leader should plan, and which product form factors each panelist expects to matter in 2027. Across these topics, both speakers held that classic product principles still apply, but that empirical testing, speed, and closeness to users matter more than before.
The toggle as a deliberately imperfect solution
Vo opened with what she called "the toggle": in the product the two work on, a switch with one side for one mode and the other side for another. She asked Seshan how the team decides to release something it knows isn't ideal.
Seshan said the shift had been personal as well as professional. She described herself as the child of Asian parents who expected perfect output every time, and said that after roughly seven years at Stripe she was used to a culture where everything was "deeply polished and deeply considered", down to whether an enum had exactly the right name. What changed, in her account, is that urgency now matters a great deal. Teams can theorize at length before shipping, but she said nothing compares to the empirical evidence of watching users try a product and then iterating. She has had to move from trying to get things perfect a priori to shipping and iterating as quickly and effectively as possible.
She said directly that the toggle is not the ideal solution. What mattered was putting an agentic harness, the ability to use an agent to get things done, in front of the more than a billion people who use ChatGPT, without disrupting developers' existing workflows. The toggle was the way to do both at once.
Deprecation and taking users along
Vo noted that she and Yu had talked before about how teams now build and throw away a lot of code and product, and asked how Yu decides what to keep and what to deprecate.
Yu used the toggle as an example. He said there is an obvious next step, which is no toggle, and a story that runs from step A to step B. In Yu's view, users accept changes and reversals much more readily if they are taken along on that journey, especially if they have some transparency into what is happening behind the scenes. As long as the story is coherent and users can follow it with the amount of attention they actually pay, Yu said, a team can change a lot without burning bridges.
Tara Seshan's quality bar
Asked what keeps quality high when things move this fast, Seshan listed several tests. The first is whether a feature is additive: does it unlock real value, is there something genuinely useful in it. The second is an internal bar. Products ship internally first and everyone uses them. She said there is no specific number, but the team looks at whether a product retains well, whether people find something delightful or surprisingly great in it, and whether it unlocks a new use case for the models.
She called the third test probably the most important in this era: aiming about two to three months ahead of where the models will be. Products should be neither too anchored in the present nor so futuristic they are unusable. She described the target, half-jokingly, as being "sufficiently AGI-pilled." In summary, her questions were: Is it valuable? Do users care? Is it surprisingly great? Does it fit model capabilities? And does it get out of the way so the model can do a great job?
Nan Yu: absorption is the binding constraint
Yu named a different constraint: how much people can understand and absorb. He pointed to the idea of capability overhang, where models can do far more than people take advantage of, and said that gap is the real limit. It doesn't matter what you put into a product, Yu argued, if people can't absorb the changes and the new capability.
Can enterprises absorb this pace?
Vo, who said she had spent some time in enterprise software, noted a common claim from customer success and sales teams: B2B customers can't absorb change. She asked whether that still holds.
Seshan said it is partly true. The current pace probably feels "beyond breakneck" to most enterprises, a flood of features and advances that may feel overwhelming. But she argued that if you don't ship the frontier to enterprise customers, you get leapfrogged. Her example: most of OpenAI's enterprise customers had been using chat, asking questions and getting answers, while the agent revolution was happening elsewhere. The team felt it had to get agents to those customers as fast as possible, or they might consider alternative products. So even though enterprises have their own pace for absorbing updates, OpenAI shipped what Seshan called a giant, revolutionary change that broke their processes and their assumptions about how updates should arrive, so they could get more value. She tied this to an old product truism: do what users need, not what they say they need. She said it applies now almost more than ever.
One agent or forty?
Vo then raised what she framed as a possible debate: single-identity agents versus many specialized ones. She contrasted her own setup, a Codex agent docked on her computer that she talks to all day, with the 40 bots she said are load-bearing across different parts of her business. Builders, she said, are choosing between one general agent identity that "gets out of the way of the models" and hand-crafted soul.md files for many small, job-specific agents.
Yu said the most successful designs map onto human nature. People have millions of years of evolution that shape how many things they can hold in their heads, and 40 agents is a lot. He compared it to managing 40 direct reports. Maybe someone like Jensen Huang could, he said, but most people would doubt they could track that many threads. So agents need to be grouped into bundles of activity. He observed that when you ask people how they manage their agents, a common answer is that they have a chief-of-staff agent that manages the others. Yu's response was that this is cheating: you aren't managing seven agents, you're managing one that runs your team. In his view, that dynamic shows up as soon as people are given this kind of freedom.
Seshan acknowledged the philosophical side but said the question is mostly tactical. How do data access and permissions work? If there is one entity, does it switch between a service account and the user's account, and how is that shown clearly to the user? Does a single agent need segmented memory? If it joins a private Slack channel, is it effectively a different agent there, "severed" in the sense of the TV show Severance? Should it use each person's credentials or a shared set? For her, the answer depends on the specific use case and on working through edge cases like how memory should behave and how the agent should talk to each person.
The skills agent design demands
Vo said the two answers reflected two sides of product craft. One is getting inside users' mental models. The other is systems and architecture thinking: working through the platform components and edge cases needed to deliver an experience sensibly. She asked whether other hard skills now matter more.
Yu agreed that systems thinking plus user empathy are the classic product management ideas, now applied to a different technology. The work means starting from scratch and rethinking how these products want to be shaped, but he said these are "just the old principles in disguise."
Seshan added a third element. She said she enjoys thinking abstractly, having come from Stripe, which she called a very academic company, and from an academic family. But she said what is also required is relentlessness: trying things again and again, taking a lot of pain along the way, and seeing what works. Yu added the ability to learn from those feedback loops, which he said has always been necessary but now runs at a much faster clock speed.
Platforms, ecosystems, and computer use as a fallback
Vo said she now loves computer use and would rather point it at a product's UI than bother with an MCP integration, even if it costs more tokens. She asked how the panelists think about the wider ecosystem of work tools.
Seshan described ChatGPT as a platform and said the team considers which functionality belongs natively in the platform and which first- or third-party developers should build in as plugins or additional interfaces. For a third-party meetings app, for example, the question is which hooks to expose so the app works well for someone using Codex or computer use. She rejected a binary between "in our platform" and "in someone's app." Instead she described layers. A plugin might pull in all of a user's meeting data and work out next steps. If a system doesn't interface well, computer use is the next layer, a fallback when pre-built or ecosystem tools aren't available. The goal is to expose hooks for developers, make the resulting experiences composable, and make sure that if the first layer fails, there is a second layer that lets the user finish the task.
Vo then asked how anyone builds a high-quality, well-branded experience from a non-deterministic model, a natural-language interface, a computer that can click anything, and third-party tools the company doesn't control.
Yu said the key is the difference between finishing the whole job and finishing everything except the last mile. He argued the last mile can feel worse than a non-starter. If something fails outright, you do it yourself or try another route. But if it gets 99% of the way and then "barfs," it is a worse kind of unfulfilled promise. What feels great about computer use, in Yu's view, is that it always works. It may be slow or use many tokens, but it gets the job done all the way while the ecosystem catches up with the right hooks and MCPs.
Seshan invoked Charles Eames's idea that a user should feel like a guest in your home, with the host having anticipated what they want and provided it gently. She said users of the app should feel their needs have been anticipated, and that increasingly the model itself is what does that. Vo joked that every time she visits that "home," it tells her it's time to update.
Working with research
Vo said collaborating with research teams is new for most product people and asked what Seshan had learned. Seshan said it was entirely new to her when she joined OpenAI and offered what she called a novice's view, comparing herself to the Connecticut Yankee in King Arthur's Court. Her first lesson was that working with research is quite different from working with engineering.
Her approach is to bring very specific use cases and a clear picture of what users are trying to accomplish, along with sample readings: the actual session, what it looked like, and why the model didn't do what it needed to. She said learning to write good evals, and writing as many as possible, is the most important thing. The ideal is to show that with a particular prompt and set of skills she got the desired outcome from the model, and then take that to post-training and ask how to train the capability in from the start. She stressed that it begins with recognizing this is a wholly different way of working, then getting good at bringing specific user data, and, once researchers trust you enough, writing the evals that start the iteration loop.
How far into the future to plan
Asked what time horizon a product leader should work on, Seshan joked that she lives in the now, then said she does have to live a little in the future: ideally two to three months out. Planning in terms of years is a mistake, she said, because those predictions are almost always wrong. She keeps a personal document of things she predicted that didn't happen, as a reminder that she can't forecast years ahead. But building only for today means getting left behind.
Vo polled the audience about who was doing 2027 annual planning and called the gap between that and a 60-to-90-day horizon the fundamental tension. Seshan said it depends on the business. She contrasted Stripe with OpenAI. In her view, the payments market looks largely as it did before, accelerated in some ways, so you know the forces and can model bull, bear, and base cases. She said she can't do that for OpenAI's business, joking that perhaps Sarah Friar could. Her advice was to understand your market's pace and dynamics and choose a planning cadence to match. When Vo concluded that Stripe PMs still have to do annual planning, Seshan agreed and wished them luck.
What still matters, and what's newly table stakes
For the closing "then and now" segment, Vo asked what persists from classic PM craft and what has become newly essential.
Yu said the emphasis changes, and, given absorption rates and capability overhang, onboarding matters far more than before. He noted that onboarding is often seen as a niche area many engineers don't love working on, but said real success depends on designing the initial experience and continually helping people understand what the product offers.
Seshan pointed to users' understanding of privacy, of how their data is used, and of the product's norms. Before, she said, products could get away with unpredictable behavior because a settings menu or popup could explain it. Now that agents act semi-autonomously, a user needs to be able to predict at any moment what a feature or new primitive will probably do. On PM craft generally, she said her answer might sound boring: all the usual things still apply, including being in the details, holding the high level and low level together, and deeply understanding the design and the product. What has increased is the importance of testing empirically, moving faster, and testing with users.
The "DMable PM"
Vo observed that OpenAI's product and engineering team is highly accessible in public Slack. She said she regularly sends them feedback IDs with complaints, such as her agent Astra having "an attitude." She asked whether direct user relationships are becoming table stakes.
Seshan said her background in developer and enterprise products meant she had always been reachable by users asking for features. What feels new is offering the same accessibility on a huge consumer product. To her, direct contact with users is simply normal, and it connects to her earlier point: a PM's value to research lies in knowing specifically what users want. She said she had been thinking Vo should send her those Astra outputs, because that is exactly the kind of information that matters.
Yu compared this to dev tools, where engineers are picky and detail-oriented and often don't give useful feedback until the second or third follow-up question. He said many PMs now feel that change, because problems with natural-language agentic experiences are subtle. A complaint that an agent has an attitude needs unpacking: what exactly did it say, what was the context, did the user provoke it? A deeper, more direct relationship with users pays off there in ways that may not have been possible before.
Predictions for 2027
Vo closed by asking what would be big in products in 2027, noting that new form factors seem to emerge every week or month.
Seshan answered: voice. She said voice has changed how she works and feels much more natural than typing. She said it has greatly reduced the tech support she does for her own family, and that when she was onboarding OpenAI's people team onto a new product, voice "changed the game" because it is so intuitive for so many people. She linked this to Yu's point about human nature: speaking is how people have evolved to interact. She said the models are finally getting good at voice and getting faster, and called herself very bullish on it.
Yu predicted something that looks a lot like self-driving software. He described the "empty input box problem": whether it's a general tool like ChatGPT or an industry-specific one, users get the product and then face the question of what to do with it and how to learn it. Since these products are now, as Yu put it, literally intelligent, they can operate themselves to help users and provide a gentle on-ramp. He expected this to become so normal that people will look back and wonder why it didn't always work that way.
Vo offered her own prediction: hardware that replaces carrying an open laptop, something like "a little Codex in a box." She asked the note-takers in the room to serve as accountability buddies for checking which of the three predictions about 2027's defining product proved right.
All right, we're here. And I warned them backstage what we were going to start this talk about because I know there's a lot of questions about OpenAI. There's a lot of questions about the future of product development. But what I want to talk about is—
What do you want to talk about, Claire?
The toggle. I want to talk about the fact that every AI product has two faces. It has left toggle and right toggle. And there's been a lot of conversation today about iterative product development and shipping things that are imperfect and getting things into the wild. And so Tara, I just want you to talk to me about shipping imperfect things instead of a toggle, and how do you make these decisions about putting things in the wild that you know may not be the perfect thing right now but are the right thing ultimately to build the right product?
For sure. I will say this to start: I'm the daughter of Asian parents. I'm not used to doing imperfect things. You have to output something perfect every time. I also worked at Stripe for seven some odd years, and everything at Stripe is deeply polished and deeply considered. And did you think through it 100 times, and did you name this enum the exact correct thing?
And I think what has been a big change about shipping things in this era is actually urgency really matters in terms of getting the functionality out there. And also there's a lot of theorizing that one can do before you ship something, but nothing compares to the actual empirical evidence of seeing users try it, use it, and then iterating from there. So I've really had to shift my mindset from how do you get it perfect a priori versus how do you ship something and iterate on it as quickly and effectively as possible.
So with the toggle, is it the perfect solution? No, it is an imperfect solution. It is certainly not the ideal thing to ship. But what was really important was getting this agentic harness, or the ability to use an agent to get things done, in the hands of the over billion users that use ChatGPT. And so we needed to ship something that allowed us to do that without disrupting developers' workflows.
As such, an imperfect solution. And I do believe that in a year maybe we won't be having the same conversation about this toggle. And Nan, we've had this chat before where we've talked about right now you have to build and throw away a lot of stuff, whether it's code or product. And I'm curious how you think about deciding what to deprecate and what to keep around and keep alive.
The toggle is a good example of this where I think there's an obvious next step, which is no toggle, right? There's a story that evolves from step A to step B. And if you can take users on the journey with you, then they're a lot more accepting of the change and the turn and the product evolution. Especially if they appreciate that they can have an understanding and some transparency over what's actually going on behind the scenes. So I think as long as you have a coherent story that users can follow along with, with the amount of attention that they're paying, there's actually a lot that you can do and not really burn any kind of bridges with them.
Yeah. And I'm curious, the toggle is maybe imperfect, and we may build a lot of stuff that we have to throw away, but you do exercise good judgment and you do exercise caution and reason inside the products that you've worked on or are about to work on. What constraints do you hold really strongly when building product that's moving so fast? What keeps that quality bar high for you?
I think there's a few different elements of it. One is certainly, are we doing something additive for users? Are we unlocking additional value? Is there something genuinely useful in this product?
I think the second is that we have this internal bar of taking a product, shipping it internally, trying it out, having everyone use it. And if there isn't enough uptake internally—and that's not a specific number or exact bar—but there is, is it retaining well? Do people find something delightful about it or something surprisingly great about it? Is it unlocking a new use case of the models? Is it doing something novel? I think that is that sense for this is adding to their experience. This is something that is unlocking a new set of functionality.
And then probably most importantly in this era, is this aiming for two to three months of where the models will be rather than being overly anchored in the present or being overly futuristic and unusable? Is this at that perfect zone of being, I don't know, sufficiently AGI-pilled? That I think is the bar here. So is it valuable? Do users care about it? Is it surprisingly great for users? And is it in line with model capabilities? And is it getting out of the way of the model doing a great job?
Nan, what about you? What constraints do you think help you build great product?
The biggest constraint is honestly people's understanding and ability to absorb what you're giving them, right? We talk a lot about capability overhang, but models can do so much, and people aren't necessarily taking advantage of all those abilities. I think that's where the gap is. So it doesn't matter what you put into it. If the ability for people to absorb those changes and absorb that capability isn't there, then that's the real bounding process that you're dealing with.
And we're talking right now, you two are working on Codex and Work, and there's a lot of business users, including myself, that have built businesses around relying on this tool. And I think I spent a minute and a half in enterprise, and a lot of people say that B2B customers cannot absorb change. Customer success people tell me they can't absorb change. Salespeople tell me salespeople can't absorb change. We cannot absorb the rate of change. Do you feel like that still stands true in the AI era, or do you feel like that statement is broken right now?
Some of it is certainly true in that I think the current pace probably to most enterprises feels beyond breakneck, something beyond breakneck, like a Lollapalooza of features and advancements. It certainly feels maybe to some overwhelming. But at the same time, if you're not shipping the latest and the frontier to these enterprises, you'll end up getting leapfrogged.
So a great example of this is that most of our enterprises were using chat. So they were merely asking questions and getting answers, and whilst that was happening, the agent revolution was going on, and we needed to get agents to these enterprises as fast as possible. Otherwise they might consider alternative products. And so while they have their own, maybe in the day-to-day, pace of absorbing changes, we needed to ship this giant revolutionary change that broke all of their processes and broke all of their conceptions about how enterprises should receive new updates so that they could actually unlock more value.
And so it's one of those things, it's a truism of building product, which is do what your users need, not what your users say they need. And I think that applies here almost more than it ever has before.
I want to talk about agents since you brought them up, because I do think we're in the age of agents, and I have certainly been in the age of agents for a little bit. And I want to have a debate with you.
Oh gosh.
Or I want your opinion. Maybe we won't have a debate. Maybe we'll agree. Single-identity agents versus multi-identity agents. My beloved Codex docked on my computer that I talk to all day versus my 40 Grok bots that are load-bearing across different parts of my business. I do think as product builders, people who are building agents, we're seeing this divergence of do we centralize on a central generic agent identity that can do anything, get out of the way of the models, or do we artisanally craft SOUL.md files and have little micro agents for the job? But I'm just curious, coming in fresh eyes, where do you land, if anywhere, on that spectrum from a design principles perspective?
I think when designing this, the most successful efforts have really tried to map onto human nature, right? And we have millions of years of evolution that help us understand how to hold different things in our heads. And 40 agents is quite a lot, right? And if you asked anybody, would you manage 40 direct employees? Jensen Huang might do it, right? But a lot of people would look at that and be like, I don't know if I could actually keep up with that many threads ongoing.
So I think there's something in there where we need to group these things into bundles of activity. You ask people, hey, how do you manage your agents? Usually what you'll hear is, well, first I have this chief of staff agent that manages all the other ones for me. It's like, okay, well, you've just cheated, right? You're not managing seven, you're managing one that's running your team. So I think that kind of dynamic immediately comes to the front as soon as you let people do this kind of stuff.
What about you? What do you think?
My mind with this, certainly there's a philosophical element to it of do you have this one unified entity? Do you have these smaller entities? But actually there's a lot of really tactical questions that go into it. How do you think about data access and permissions? If you have one entity, is it switching between using a service account or your account? How do you manifest that to the user in an easy way to understand? If you have a single agent, does it have segmented memory? Every time it joins a private Slack channel, is it almost a different agent? Is it severed in the Slack channel in the Severance sense?
And so actually to me, answering this question goes back to really specifically the use case and probably thinking through all of those edge cases of how should the memory behave? How should it talk to each person? Should it be using their credentials, or should it be using a universal set of credentials? There are philosophical things one could have and opine on in this topic, but to me there are just so many practical product considerations. The use case matters so much in deciding this.
What I love about these two answers is I think for product folks, everybody in the audience plus designers, there are these two sides of the craft that we do. There's this very human side where you're trying to—there's Inception and Severance, I guess. There's the Inception of how do I get in the mind of a user, and what are their mental models, and how do they interface with the world, and what is the most natural way for our products to touch their problems and solve them? And then there is systems and architect thinking, which is in order to manifest this into the world in a realistic way, I need to think through all these systems, all these platform components, all these edge cases to make sure that comprehensively we can even deliver this to the user in a sensible way.
Do you think those are the two skill sets you need to design and build great agents in this moment? Are there other hard skills you feel like you're relying on now more than ever as you design or think about agentic experiences?
You're describing systems thinking plus user empathy. These are very classic product management types of ideas, right, that we applied to a different universe of technology in the before times. But now we have to start from scratch and from those principles rethink how do these products want to be shaped, right? So I think those are very important, but ultimately they're just the old principles in disguise.
I would maybe add a thing to that, which is, again, I mentioned I used to work at Stripe, which is a very academic company, and my parents are very academic, and I love thinking about those things in abstract. But I'd say the third thing that you just really need in addition to user empathy and systems thinking is a relentlessness to try it out and keep iterating. And you're going to take a lot of pain while you do that, and there's just an element of relentlessly trying and trying and trying and trying and trying to see what works.
That and the ability to learn maybe from those feedback loops. That certainly has always been necessary, but just the clock speed has gone way—
Yeah, you're throwing a lot of spaghetti against a lot of walls all the time. Actually, your chief of staff agent is telling all the other agents to fling spaghetti at the walls.
Talking about systems thinking and platform, I am really curious because all of us are looking to tools, certainly ones that you build, for productivity and efficiency. And I was telling somebody backstage, I just love computer use now. Don't even give me an MCP. I want to point computer use at your UI. And then I don't care the token. Congratulations on selling me tokens. But I think about what is the intersection between the tools we had and the tools we have now, what you build, what you platform for. I'm curious how you think about the broader ecosystem as you're trying to enable productivity for people at work.
So I think we think a lot about platforms and ChatGPT as a platform, and what is the functionality that natively goes into the platform? What is functionality that we might allow folks to build in, whether that's first party or third party, as plugins or additional interfaces on top of it? And that's how that relates to the ecosystem. So in that platform, what are the hooks that we need to expose to a third-party meetings app so that it can be most effective for you using Codex, using computer use, using all these tools?
But I don't think it's so binary as what should be in our platform or in our productivity system versus in one of these apps. I think we should offer an array of different tools, perhaps a plugin to be able to get all your meetings data and understand what you need to do next. But if that doesn't work and those systems don't interface well, then there's always computer use as a next layer or as a fallback to be like, well, if the user wants to get this task done and the pre-built tools or ecosystem tools aren't available for it, then there's always something else that they can fall back on.
And so I think it's more about structuring these capabilities in the platform in terms of layers, exposing them as a set of hooks that third-party developers and first-party developers can hook into and build cohesive experiences, having the right system for those experiences to be composable with one another. And then ultimately all of this is just serving the user's goals of, as they try to get their task done, if the first layer doesn't work, there's a second layer for them to be able to accomplish their task as you traverse all those capabilities.
What do you think about integrated experiences, like do it all in one platform, let's live in our house, versus expanding into this ecosystem? And how do you build, I mean the thing—
That I was thinking as you were talking is you have a non-deterministic model on top of a natural language interface, functionally, with a computer that can click and do anything, with third-party tools and integrations you don't control, and you're still trying to build a comprehensive, high-quality, good, branded, beloved experience. How do you even climb that mountain as a product leader? What genius are you going to bring in to help us here?
Look, I don't know if this is a new or unique idea, but there is a huge difference between getting all of the job done versus getting everything except for the last mile. And the last mile honestly feels sometimes worse than just it's a non-starter. If it's a non-starter, you can just, look, I'm just going to do it myself, or I'll try another route, or whatever it is. But if it gets 99% there and then it just barfs, then you're like, well, this is actually a worse sort of unfulfilled promise.
And I think what feels great about computer use a lot of times is that it will always work, right? It might be a little bit slow, or it might cost a lot of tokens, like you were talking about, but it's going to get the job done all the way. And as the ecosystem, we wait for it to catch up, right, and to provide the right hooks and the right MCPs and all that kind of stuff, it will 100% do the job for you. And I think ultimately that's what makes it feel super great.
I think the right experience amidst all those complexities and changing things, the ecosystem, the interface, the model itself, the tools like computer use, the unpredictability of all of it, is that the user should almost feel like the Charles Eames quote, that your user should almost feel like a guest in your home. You've anticipated all the things that they want to do and you've provided for them gently. So should our user in our app feel like we've anticipated their needs. And in some ways it's the model increasingly doing that.
Well, I have to tell you, every time I go into your home, it tells me it's time to update.
Yeah. Get the latest build.
That is the welcome mat that I get. It's like I walk in, it's like, you want a Renault? I got a Renault for you.
Well, I'm curious. You're designing this product while it's flying off the shelf. A lot of the things, Nan, as you said, are old things: user empathy, systems thinking. But there are some new things, right? There are some new skills, and I'm curious about working with research because this is something I don't think a lot of folks in the audience have done in their careers, because I think it's very new for product people or product leaders. What have you learned from going into this new role where you really have to collaborate with research?
Yeah, I will say for me, working with research was entirely new coming to OpenAI. So I'm certainly not the expert on this topic. I will give you the novice's view, like A Connecticut Yankee in King Arthur's Court sort of view of this. My take coming in working with research was, oh, working with research is different than working with engineering. That might be obvious to everyone here, but I was like, oh, it is quite different.
Inasmuch as I can add something to working with research, it is for me to come with very, very specific use cases, a really clear idea of what users want to do and their goals, to have done sample readings. So understand literally what was the session and what did it look like and why did the model not do what it needed to do. If I can write evals, I'm there. Learning how to write good evals and writing evals as much as possible is the most important thing. And inasmuch as I can say, with this prompt and these skills, I was able to achieve the outcome with the model, that is the ideal thing I can take to post-training and be like, look, now how can we post-train this capability into the model from the beginning?
But I think it starts with first understanding that this is a wholly different way of working that I really didn't know how to do before, and then getting to the point where you're good enough at bringing the specific user data. And then if they trust you sufficiently, you can then write the eval to get that iteration loop going.
How far in the future do you feel like you live as a product leader? Because I do think there's this time horizon of, I could fix what's right in front of my face, and there's stuff to fix. There's, I know what's coming next, and I can anticipate where the product's going. And then there's, the models next year are going to be out here. Where do you feel like you have to live as a leader building on top of that foundation that's moving?
As soon as you said that, do you live in the future, I was like, I live in the now. I don't live in the future. I live in the now. No, I do. I think I have to live a little bit in the future. To me, the ideal is two to three months. Can't be too big-brained and be like, this is the future in years, because I think inevitably our predictions of that are almost always wrong. I have a personal doc of incorrectly predicted things that I thought were going to happen and didn't happen, just to remind myself of my inability to predict the future years out. And then at the same time, if you build exactly for today, you're going to get left behind. So I constantly repeat this, but it's two to three months out.
I do want to just, let's poll the audience again. It's the end of the day. Let's get interactive. Who right now is doing 2027 annual planning? And who just heard her say you can't go out more than 60 or 90 days? I think this is the fundamental tension. It's so hard to even guess.
But it's not the same for every business. Again, just to use an example I'm familiar with, Stripe's business is very different than OpenAI's business. The dynamics of the payments market look largely similar to how they looked before. Certainly accelerated in other ways, but you know the forces and you can model, oh, here's the bull case, bear case, base case, in a way that for OpenAI's business you just cannot do. Well, maybe Sarah Friar can, and she's the expert, and I'm not as good as her for sure. But it's hard for me to predict OpenAI's business the way I can feel like, oh, I understand the levers in Stripe's. And so maybe it is appropriate to plan many, many months out, but I think it involves understanding the pace of the market you're in and understanding, given the pace and the dynamics of that market, at what cadence should I be thinking and planning in the future.
So what you're hearing is Stripe PMs still have to do their annual planning.
Oh, for sure. Good luck to you all. Yeah, for sure.
All right, let's talk a little bit. Let's close out with then and now. We've all been in product for a minute and a half as fresh 26-year-olds. What still matters from a classic PM craft? And what do you think is newly table stakes?
I think the emphasis changes, right, for a lot of reasons. But we talked about absorption rate and capability overhang. So I think something that probably matters way more now than it's done before is onboarding. I think people have talked about onboarding as an important thing, but it's also something where it's a little bit specific, and a lot of engineers don't love working on it and things like that. But I think really, to get a lot of success, thinking about what is the initial experience, how do you get people to really understand what your offering is constantly, is probably something that matters way more now than it does before.
Tara, what do you think?
Yeah, I think users understanding privacy, their data, how that's used is way more important than it always has been. Users understanding norms of your product. Before, any sort of unpredictable behavior you could get away with because there was some settings menu or some pop-up that could appear. But now if the agent is doing things semi-autonomously, you actually really need to make sure that at any given point a user can predict what is probably going to happen with a feature or a primitive that you're introducing.
And in terms of just product thinking or getting up and running as a PM, in some ways it's going to sound like a really boring answer, but all the usual stuff still applies, except the importance of testing things empirically and moving faster has only increased. So yeah, do all the other stuff you used to do, which is get in the details, have the high level and the low level at the same time, really understand the design, really understand the product, but also do it all faster. And testing with users has just become way, way more important.
I have one more that Nan is going to become a victim of when you are in the Open Slack, which is the OpenAI product and engineering team is highly DMable, highly accessible, highly visible. I mean, they're getting feedback UUIDs from me on the regular, like, I didn't like what Astra said. It had an attitude. And so I'm curious, certainly I have tried to do this in my own life, but do you feel like direct user relationships right now are becoming table stakes for product leaders? This idea of the DMable PM, I think, is one that I see a lot out of OpenAI, and I'm curious, staying close to the ground, if you feel like that's becoming more and more important.
My background is in developers and enterprise. So I've always been a DMable PM in the sense that people could always message me and ask for features. But I think what feels new is that for a large consumer product or a product that everybody uses, to do the same level of accessibility. So in some ways, it doesn't feel normal to me. Of course you should be in direct contact with your users. And as I mentioned earlier, your value to research is knowing specifically what the users want and what their use cases are. And so the more you can be a DMable PM or engineer, the better.
If you didn't like Astra's output, I was actually thinking, Claire should send me those outputs. It would be really interesting to know what's happening there. That stuff is the important thing. In part, that's how you add value. But I'm curious Nan's take.
Yeah. I think you worked in dev tools, right? And engineers are extremely picky. They're extremely detail-oriented. Often you're not going to get good feedback until the second or third follow-up question. And I think that's what a lot of PMs are feeling a change about, right? Because if you're designing an agentic experience with natural language interfaces, the problems that people encounter, like Astra gave me attitude, they're pretty subtle, right? That bears some explaining. What exactly did it say? What was the context? Did you make it mad? All sorts of follow-up questions that you need to have. So I think having a more direct, deeper relationship bears a lot of fruit there, right, that maybe you wouldn't have been able to take advantage of before.
Neither of you have gotten my DMs, so I'm very excited for the next 90 days for you all. We'll put us in a group chat. It'll be really fun.
Okay, last question. So when Dan was up here with the Anthropic team, he asked what in a year about the product craft, about the role, will be different. I want to talk about the stuff. What do you feel is going to be different about products, like what we touch, the buttons on the site, the experiences? Where do you feel like this is all going? Because I feel like every week, every month, we come up with a new form factor, a new thing that matters. Just curious, as product thinkers, what do you think is going to be big in 2027?
I have an answer for you, and it's voice.
Voice.
I love voice. You try out GPT Live. I love voice. We're yappers.
Yappers. Hillary Fidley. The yappers API genius.
Voice has changed the way I work. I just want to talk to a computer. It's so much better. It feels so much more natural in some ways. I'm sure all of you have been tech support for your parents. The existence of voice has reduced the amount of tech support I've had to do for my family to an incredible degree. I was onboarding our own people team at OpenAI onto a new product, and the existence of voice changed the game. It is such an intuitive interface for so many people. It's the natural way they interact.
To the thing Nan said earlier about, you kind of have to think about human nature and the way we've evolved over millennia to interact with others, speaking feels so natural. The models are finally getting good at it. It is getting faster. I'm so bullish on voice.
Voice. What about you?
I think probably something that looks a lot like self-driving is probably going to be the thing that really becomes much more important and prominent in the next year. I think a big problem people have is empty input box problems, right? And it doesn't really matter if you're using a general tool like ChatGPT or something more industry-specific. You have the same, it's like, okay, I have the product. Now what? Well, now you've got to learn how to use it, right? These products are smart now. They're literally intelligent. So having them use themselves and help you out and make that a really, really gentle on-ramp, I think, is probably going to be a very normal thing that we're going to look back on and be like, why didn't it always work that way?
Okay, we need the notetakers to be our accountability buddy. So we got voice, we got self-driving. I'm going to say hardware that replaces carrying around our laptop open. So hardware. Yeah, I want a little baby bot over here, a little Codex in a box where I don't have to do—she's nodding—where I don't have to do this. That's my prediction. So whoever took notes, we're going to see which of us were right on what 2027 hot product of the year. Everybody thank these lovely leaders from OpenAI for coming.
Claire, thank you.
Article published · Updated
