Why Claude Can't Be Your PM (Yet): Anthropic's Product Leaders on Judgment, Agents, and Parallel Bets
Lenny's PodcastAt the Lenny and Friends Summit, Dan Shipper put a question to Anthropic's Ami Vora and Mike Krieger: a year ago people were saying AI might make product managers obsolete, and now it seems everyone is supposed to become one. Which is it? Both argued that the PM role has not gone away. The underlying job is unchanged, but the pace of change raises the bar for adaptability, judgment, and operational discipline. From there the discussion moved to agent-native software, running many overlapping experiments, and folding the ones that work back into a coherent product.
The Human Side Changes Slowly, the Technology Changes Every Two Months
Vora described product as "a bridge between the real problems that people have in the world and the technology that you can use to solve it." That definition hasn't changed. What has changed is the pace. Technology used to shift every five or ten years, and now it feels like every two months, which Vora said is hard for a person to absorb. Human problems, in contrast, stay roughly the same. Her approach is to stay obsessed with the human problem while keeping a "fresh-eyed" view of the technology, which means being willing to throw away most of what she knows about what used to work and try again with everything she knows about the problem.
She compared the current moment to "rewinding the clock" to when she started in product. Back then nobody had clear definitions of the job, and you were a general problem solver who kept going until you hit a wall, then asked someone or figured it out. She sees roles blurring again in the same way. The difference is that there is now a lot of scaffolding available when you do hit a wall.
Mike Krieger's "Before and After" Moment
Krieger told a story from his own work. Earlier in the year he moved into an individual-contributor role and has mostly been building. As his project got close to shipping, a PM lead named Cat, who works with Vora, pulled him aside and told him the project needed a PM. His first reaction was doubt: "Claude's got it." She insisted. Once a PM joined, Krieger saw all the glue and connective tissue that would have been dropped without one. A week later he messaged Cat to say she had been "100% right."
He listed what that work covered. Anthropic serves everyone from prosumers to very large enterprises, so someone has to bring all of those constituencies along. Someone has to make sure the customer success team can talk about changes as questions come up in real time. Someone has to loop in safeguards, which he called very important. Someone has to keep people on track. Nobody has infinite capacity, even with AI assistance. A person in this role lets builders stay heads-down "in full-on Claude mode" instead of switching to the very different activity of making sure everything goes well. Because teams now move faster, Krieger said, the role demands more operational excellence than before. If you are recruiting a team, this hat is "increasingly important."
Why Not Just Have Fable Do It?
Shipper pushed back and asked why Krieger didn't just have Fable, Anthropic's model, do the PM work. Krieger acknowledged that Anthropic runs on a lot of Claude-provided connective tissue. For example, Claude can flag that something relevant is happening in another part of the organization, surfaced through search. But he said Claude is not yet a convener. In a recent talk on the role of humans in a world of very powerful AI, he described several archetypes, and one was the convener: someone who still has to bring the Claudes and the people together to get work done.
He joked that the turning point will come when Claude schedules a meeting for him and, when he asks who set it up, explains that it thought two people should talk. That hasn't happened yet. He added that this is "not because it can't," and suggested it may just be a matter of turning on the setting.
Vora added a second reason the role matters. The tools make it possible to build a lot very quickly, which expands the universe of what could be built. You still need good judgment about what to build, often on limited information. With so many forking paths, you also have to be even more relentless about making sure the right thing actually happens. She summed up the attitude this way: I know the problem, I have a feedback loop with the users, and I'm going to go for it. She called that relentlessness and judgment in the face of ambiguity really important.
Skills That No Longer Matter
Shipper asked Vora about her earlier remark that she spent decades learning to answer certain questions she no longer needs to answer. Her example was the mix of product and design judgment she worked hard to build: imagining a user's situation, where their finger would land on the phone, where a button should go. Reviews used to revolve around exactly those questions because building and shipping was so expensive. She doesn't think anyone will ask her that again. It is now faster to build three versions and try them.
She said this is hard on identity. She compared it to the innovator's dilemma, where companies keep doing what they're good at even when something else might matter more, and said the same thing happens to individuals. Part of the technology changing every two months is that you have to throw away what you knew about yourself in the job. She admitted she has never worked at a model company before and often doesn't know whether she'll be good at the next thing. Every couple of months she has to try something new and see if she can figure it out.
Framing the Chaos, and Giving Someone the Pen
Asked what lets people do this well, Vora named adaptability and a higher tolerance for change, along with judgment and relentlessness. For her team, she stressed talking about these things openly, because otherwise constant adaptation can feel lonely when in fact everyone is going through it. She talks about "framing the chaos" so it feels less taxing. In uncertain times it is tempting to map out exactly which products will be built or exactly how a career will go as a way of exerting control. She thinks that locks you out of trying new things. The goal is to make engaging with chaos feel safe and plausible, and to acknowledge the emotions involved, because that is where much of the "magic and productivity" will be.
Krieger added an organizational counterweight: chaos needs clear ownership. In Anthropic's Labs group, exploratory efforts are called "bets," and each has a bet lead who is the directly responsible individual (DRI). That person decides whether to double down or wind something down, and whether a team needs more or fewer people. Because things move so fast, he said this role matters even more. "We're all lost together," but someone holds the pen on the next most important thing to de-risk, understand, or learn.
What Agent-Native Software Looks Like
Shipper then turned to what gets built, now that both humans and agents use software, sometimes through delegation and sometimes collaboratively. Krieger sketched a rough timeline. First, AI sat in a sidebar or mini window, disconnected from the product and perhaps useful for support questions. Then came deeper integrations with whole AI-powered features. Now there is a move toward being agent-native. He credited Every, Shipper's company, with a write-up on agent-native architectures. He fed it to Claude to create a skill while he worked on the topic this year. The core idea he took from it: everything a human can do, an agent should be able to do too. Very few products get this right, he said, including many of Anthropic's own. When a product does, emergent behavior appears. The agent can combine things in new ways or proactively suggest different approaches.
The next stage he is thinking about is interfaces that are themselves malleable by the agent. People have talked about malleable software for years, and he feels it is now becoming real. He gave an internal example. Anthropic is shipping a complex project with at least four independent workstreams. Claude monitors progress and also built the UI through which first the TPM and then the whole team see what's happening. If you don't like the display, you aren't stuck with what an internal team or third party built. You can change it with Claude. Krieger called this one of the biggest shifts in how Anthropic has worked over the past year: much of the software people interact with internally is built, maintained, and iterated on by Claude, and it changes all the time.
That raises open questions. Who can update the data: only humans clicking through, or Claude in the background too? How do you track data provenance? He still wants building blocks, predictability, and design systems, so things don't become "absolutely insane." But he is most inspired by the idea of tailoring software to a given task, project, or even user so it feels personal and extremely useful.
Advice for an Established SaaS Company
For a company with a scaled SaaS app deciding whether to add its own agent or open up to agents like Claude, Krieger recommended building the right primitives as an infrastructural layer. When he tries new products, he often asks the agent to do something a human could do, and he can tell whether everything runs through shared plumbing that serves both agents and the product's REST API. He expressed empathy for companies with 20-year-old applications, for whom bolting this on is difficult. Once the primitives are right, the product can evolve gradually. You might start with a side panel, then try a more malleable personal landing page on the same primitives, without blowing up the whole UI or inventing new infrastructure. The failure mode is AI that feels "super bolted on" rather than native.
Many Shots on Goal Versus a Stable Experience
Shipper asked Vora how Anthropic balances pushing the frontier with serving users who just want to chat, or enterprises whose interfaces can't be overhauled overnight. She called it a really hard problem. Anthropic's approach is to meet people where they are, because it's still very early. Nothing is stable enough to simply build "the right answer," and both the models and the ways people use them keep changing. During exploration it's fine to try many things. Once something works, it has to be made sensible for people who want stability and completeness.
She acknowledged tension with her own background, which valued simplicity, reliability, and predictability because they help users build muscle memory and intuition. Still, she encourages the team to take many "shots on goal" and to accept overlap and duplication, as long as they consolidate after reaching product-market fit. In her view, the failure mode is over-constraining a product from the start based on a theory. She would rather have "five overlapping products that all work in different ways" than one overly constrained product that never got a chance to meet the market.
Keeping Parallel Teams Motivated
Shipper asked how to run parallel experiments without them feeling incoherent or demoralizing, for example for the third team working on a similar idea. Krieger described a discussion from that same morning about two possible product directions. One had a team with high conviction and the other didn't. Assigning a team to pursue an idea nobody is calling for feels manufactured and creates a "B team" feeling. He has never succeeded by tasking a team with an area they weren't excited about, because the product ends up bad. When two or three teams each have exciting directions, he lets them explore.
He said the underlying infrastructure has improved a lot recently. At one point, by his estimate three to six months ago, Chat and Cowork had different memory systems, different MCP setups, and different file storage. Any new experiment built on top would have been disconnected and disadvantaged from the start, for example by not sharing memory. A foundations team has since worked on making memory available wherever you are across Anthropic's products. That sounds simple but turned out to be complicated "for these eight reasons." Solving it in a separate team from the Labs or frontier product teams lets overlapping experiments feel complementary.
Vora added that simply naming the situation matters. If there are three bets, it can feel demoralizing, or like one team isn't set up to succeed. She tells teams that this is the world we're in, that nobody knows the answer, that everyone is on the same team, and that they will learn from whatever works. Saying this out loud, she said, is important because everyone is "trying to figure everything out in the dark."
On how to tell whether something is working, Vora pointed to ordinary product-market-fit signals. Do people use it, like it, come back, talk about it, and get value? Does it solve a problem in a way people can use easily and feel good about?
Parking Projects and Avoiding "Capability Blindness"
Krieger raised a complication: something may look like it doesn't work, and then a new model makes it work. You don't want to kill it too early or be too far ahead. Labs sometimes parks projects. His example was the first computer-use product Anthropic built internally in 2024, which he said was "so bad because the models just weren't there yet." The first thesis was that even if it couldn't automate much, it might help with education, such as showing how to do something complex in Photoshop. In practice it would flail through trial and error ("did that work? Nope") and eventually succeed 20 minutes later, leaving the user having learned nothing. It worked neither for automation nor education.
The team parked it and kept it running in an eval harness, retrying with each new model. With the 3.7 model, they saw its performance suddenly jump. Reading the transcripts, they realized it was "succeeding more often than it's not," which they recognized as a real moment. Krieger said Labs counts it as a win when a project only shows the models aren't ready yet, as long as the work is turned into an eval or connected to the research team, so it can be revisited months later. You have to build early to develop that intuition.
Shipper called the opposite failure "capability blindness": trying something once, never retrying or writing an eval, and assuming models will never do it, until a model three months later improves in unexpected ways. He polled the audience. Many had used Fable, and fewer had used Astra. He said he often meets people who say a model can't do something and then admit they last tried an older model such as Sonnet 4.5. Vora tied this back to her theme: you have to be willing to throw everything away, or you'll assume the last thing that happened stays true forever.
Turning Experiments Into a Coherent Product
Shipper noted that Anthropic, OpenAI, and his own company all face the problem of merging a successful experiment back into the main product. Options include adding a tab, which leads to many tabs that each work slightly differently, or keeping a separate app that never merges. Nobody seems to have settled the answer.
Vora joked, "I don't know what you're talking about," then admitted they haven't fully figured it out. Part of her answer returned to primitives. Users expect to walk up to something without having to think too hard. It's easy to shift cognitive load onto them by handing over a pile of tools and letting them choose. The aim is primitives that let the system know the user, and a UI that always feels accessible, refined iteratively as people react.
Krieger drew on his Instagram experience, where tab usage followed a strong power law. The main feed accounted for about 80% of usage, and improving Explore might move it from 10% to 15%. He expects every product to face the same question of what the main tab is and how to make it great, with that experience serving as an entry point to other features. He described half the job as saying no to another sidebar item or tab, and working out whether something can be folded in once proven, without killing it too early. That, he said, is "the entire art of product development" when things move this fast.
Looking Ahead a Year
Asked what will be different if they return next year, Vora hoped people will feel they can do much more: build however they want, get answers when they need them, and start businesses. She hoped for more small teams and solo builders, and said builders magnify impact for many people who will never use the tools directly.
Krieger focused on the gap between what models can do and what most people use them for. He said that gap is not users' fault but Anthropic's responsibility to close through better products. He said it was hard in 2024, harder last year, and is getting harder still. Closing it would be democratizing. The fully working multi-agent setup would no longer belong only to the most "Claude-pilled" software engineer. It would also reach someone running research in parallel or someone managing their business. The challenge is helping them understand the moving parts, check in at the right moments, and do long-horizon work "not in a way that feels disempowering." He called that "the art," and said success a year from now would mean the gap is closer to closed.
Welcome. How's it going? Good.
Doing great. How are you all?
So nice to see you all.
Isn't this a great event? I feel like it's so happy.
Yeah. It's good vibes.
Yeah. So, I want to start this conversation off with the question that I think is on everybody's minds, which is: a year ago we were told that AI might make PMs obsolete, and now I feel like the vibes are that everyone needs to become a PM. One of the things I love about getting to talk to both of you is just by being at Anthropic and working with the models in the way that you do, you have a little bit of a sense of where the future is going, especially for different job roles. And I'd really love for both of you to talk to us about what does a great PM at Anthropic look like. Is the PM role going away? Is it changing? And if so, how?
I think the job of product is always to be a bridge between the real problems that people have in the world and the technology that you can use to solve it. And that's always been true. And I think a thing that is weird right now is that it used to be that the technology would change every 10 years or 5 years, and now it feels like it changes every two months. And that's really hard. As a human, I feel like it's really hard to learn how to deal with that sort of change.
And so what I always think about is the technology side is changing, but the human side doesn't change that fast. We have the same problems. And so what I always think about is how can I be obsessed with solving the human problems and then stay super fresh-eyed about how the technology works? How can I basically throw away most of what I know about what used to work to solve this problem and then try again, knowing everything I do about the problem?
And so I feel like that part hasn't changed. What I think has changed is it's almost like rewinding the clock. Back when the dinosaurs were in the earth and I started in product, we didn't know, we didn't have definitions. We just didn't know what the job was. And so you were just a general problem solver and you'd walk around and try to solve a problem until you ran into a wall, and then you would ask somebody or figure it out. And I feel like that's what we can do now, is go back to a bunch of the roles blurring, but you just keep going and try to figure out what's going to work. And then when you hit a wall, you have all these different sorts of scaffolding.
Yeah, that makes sense. Mike, anything from you?
I think for me, I had a real realization moment recently. Earlier, beginning of the year, shifted to an IC role, so I'm mostly building. And we were getting close to shipping something recently, and one of the PM leads who works with Ami pulled me aside and was like, "Mike, I know you love to build. This project's going well. It's looking really exciting. We really need a PM on this project." I was like, "I don't know. Do we need a PM? Claude's got it. We've got a lot going on." And she's like, "No, we really need a PM."
And it was such a reminder, once they joined, all of the things and all the glue and all of the connective tissue and all of the things that were about to get dropped if I had not done that, that would have happened. And it was just such a cool, eye-opening before and after. So much so that the week later I pinged her and I was like, "Cat, you were 100% right, and thank you."
And I think about all the things that that encompassed. It was actually bringing along—we're a company that serves everybody from prosumers all the way to really large enterprises. Have you brought all of that contingency along? Right? We need to go and figure out how to enable our customer success team to talk about the changes as questions come up in real time. Has somebody thought about that? Has somebody looped in safeguards? It's obviously very, very important. Are people getting kept on track? All of these things.
And nobody has infinite capacity, even with the assistance of these tools. So to be able to have, call it what you will, but somebody playing that role of making sure that everything is going to go well, everything is going to get connected, we've kept the end user needs in mind, to Ami's point, and that the builders who are heads down and are in full-on Claude mode can keep doing that without having to then pop out and do the very different activity, which is how do you actually make sure all these things go well, is, I think, very essential.
And I think all that's happened now is that because we can move faster, you have to be operationally excellent in that role in a way even beyond what you had to be before. But it definitely has not gone away. In fact, I have felt this renewed, yes, if you're recruiting for a team and you need people to wear these hats, this hat is increasingly important.
Can I ask you a serious question, though?
Only serious questions. Serious. This is serious.
Why didn't you just have Fable do it? And presumably there's some Fable 7 you've got access to that we don't have access to yet. So what is that person doing that you couldn't just have Fable do?
I mean, it is interesting. As you can imagine, we run on Claude, and we have a lot of Claude-provided connective tissue too. And it does catch things where you say, "Hey, there's another thing going on in this other part of the organization that you should be aware of. We found it via the slash search." So there is that sort of piece as well.
I don't think Claude is a convener yet. I gave a talk a couple weeks ago on what's the role of the human in the post-super-powerful-AI age, and there's a bunch of archetypes I was thinking about, but one of it was the convener. Somebody still needs to bring the Claudes and the people together to get the work done. And so I think that is the role that I cla organizational pull or the scheduling ability.
There's going to be a moment, Dan, when Claude schedules a meeting with me or for me where I'm like, "Who scheduled this meeting?" It's like, "Oh, I did. I thought you two should chat." That has not happened yet, but not because it can't. I guess we just turned on that setting, you know.
Stay tuned.
I think we're also living in a world where the tools are really good, and you can build a lot of things, and you can build them really fast. In some ways that just expands the universe of what is possible to build, and you still need to have really good judgment about what to build based on sometimes pretty limited information, because it feels like everything's moving really fast. So you have to figure out what to build, and then I think in some ways you have to be even more relentless about making sure it happens because there's so many forking paths that you could take. Yeah.
I have an idea of what we need to do, and I have the feedback loop to the people who are going to use it, and I know what the human problems are, and I'm just going to go for it. And I think that sort of relentlessness and judgment in the face of ambiguity is really important.
That is definitely something that I have found. Every time there's a shift, every time new models come out, it's sort of like you're looking out on the ocean. You're like, what's beyond the horizon? And then you get to fast-forward beyond the horizon, and what do you see? Well, there's more stuff. There's a whole new landscape. It's not a utopia, and it's not horrible. There's just lots of new things to discover and navigate. And I think the same is true for just models in general.
When we were talking, Ami, in planning for this conversation, one of the things that you said—and you've had a long career in product management prior to Anthropic—is that you spent decades learning how to answer specific kinds of questions as a product leader that you now no longer need to answer. And so I'm curious about, okay, what are the skills that you spent such a long time building that you actually don't feel like are that useful anymore? And what are the things that you're leaning on now more? And then how does that filter down into the team and your expectations for people that you hire?
I think it's really true. I spent a really long time trying to understand how to put myself in the mindset of a user and how something would fit into their lives and what we should build, because it was so expensive to build and ship something. Literally, this would happen where people would have reviews, and what we'd talk about is, "Ami, where should we put this button?" And we would talk about the interaction pattern and where someone's finger would go on their phone and what situation they would be in where they pulled their phone out of their pocket and what they wanted. And I feel like I tried really hard to get good at that mix of product and design. And I don't think anyone's ever going to ask me that question again.
That's a Fable. That's a Fable one.
That's in the past. Now it's actually faster to just build three versions and try them out and be like, "Oh yeah, this works for this one." So a thing that I prided myself on is not actually that useful anymore. And I do think that's really hard just in terms of identity.
We talk about the innovator's dilemma when it comes to companies, where you're so good at one thing, you keep on trying to do it even when there's this other thing that maybe you should try doing, but you can't because you're already so good at one thing. And I feel like that's also true on a personal level, where I sometimes feel like, oh, I know what I'm good at. I'm just going to keep on trying to do that even when there might be a different way to do it that might be more important in the future.
And so that's what I just encourage everyone to think about. When you think about that product bridge, part of the technology changing every two months is you have to throw away what you knew about yourself in the job using that technology and just try something different. And most of the time I don't know if I'm going to be any good at it. I've never worked at a model company. This is not a place where I have expertise, and I feel like every couple of months I have to constantly be like, "All right, well, let's try this other thing and see if I can figure it out."
What are the characteristics of someone that is able to do that well? And what have you had to develop in order to do that well?
I think it comes back to the sorts of things we were talking about. I think raising your ceiling on tolerance for change and being willing to embrace change is super important. So I think adaptability—I'm sure this is not new—adaptability is just such a key thing given how rapidly the world is changing. I think that sense of judgment and relentlessness, as we talked about, those are the things that I keep coming back to.
And really what I try to think about for the team is I think it's important to talk about those things, because otherwise it can feel really lonely, like you're the one who has to constantly adapt. But actually everybody has to constantly adapt, and I think it's hard on all of us. But I also talk about trying to frame the chaos so it doesn't feel as taxing.
I think when there's times of uncertainty, sometimes it's tempting to try to map out, here's exactly what's going to happen. Here's exactly the products we're going to build. Here's exactly how my career is going to go. And that's a way to exert control. And I think that locks you out of trying all the new stuff. And so what we have to do is frame the chaos so it feels safe and it feels plausible to engage with it. But that's where a bunch of the magic and productivity is going to be. So just talking about it, making it safe, acknowledging that there's also emotions around all of this and it's not easy, but it is a way to just participate and understand and work with the future. I think that's something we talk about a lot.
I think there's an organizational leadership question that I think, Ami, you do really well, which is you can have the chaos, but you also need very clear leadership and DRI-ship within some areas. So you can have a bunch of people exploring something very frontier, but having a clear, this is the person who's ultimately going to make the call whether this is the thing to pursue or not.
In Labs we call them leads, or we call them bets within Labs. So the bet lead is a really important role, and they are the DRI for saying, "I think we should double down on this," or, "Actually we should wind this down. This team needs more people, this team needs less." And that role is even more important because we're moving so quickly, because things are shifting really rapidly, to still have that, we're all lost together, but somebody's got the pen on what's the next most important thing we can do to de-risk, to understand, to learn, to move things forward too.
We've talked a lot already so far about the psychology and the skills of being a product leader. One other place that's changing a lot is what you actually build and how you think about the features and the functionality, in particular because we're moving from a world where you'd expect only humans to use your software to a world where humans and agents are using software. Sometimes it's delegated, sometimes it's fully collaborative. And I know, Mike, you've sort of been at the center of figuring out what does agent-native type software look like? And I think you have a lot of thoughts on where that is going. So can you talk to us about how you think about what to build, what goes into a product versus an agent, and what's the overlap?
I love this question. Maybe there's a timeline to look at, where V1 or era one was, great, we've got this AI thing. Maybe it belongs in a sidebar or a little mini window, and you can ask questions, and it's kind of disconnected. And maybe it helps you with some customer support questions. And then sort of greater integrations, where maybe there are whole features that are AI-powered, and then this movement towards being more agent-native.
And I think Every literally wrote the book on it. For that whole year where I was working a lot in this idea of agent-native architectures, I would basically feed their write-up on this and then create a skill out of it, because it was a really good encapsulation of it. It was the idea of everything a human could do, an agent should be able to do as well, and very few products actually get this right. I think a lot of our products don't quite get this right yet.
But when you do that, all of a sudden emergent behavior gets unlocked, where the agent can piece together things that weren't pieced together before, or can proactively offer different ways of doing things. I think we're still on the journey of making that good, but I think the next stage that I've been thinking about a lot is when the interfaces actually become something that is malleable by the agent too. And we've talked about malleable software for years. I feel like it's actually now really manifesting and becoming real.
And maybe back to the role of product and connecting to this question, one of the things that we have, we have
A pretty complex thing that we're trying to ship right now. And it's got at least four independent work streams. And the way we've operationalized it is we've had Claude both keep an eye on how everything is going, but also create the UI by which first the TPM and then the entire team understands what is happening in that project.
And what's really cool about that is that if you didn't like that particular display, you're not bound to however it was determined to be built, either by some internal acceleration team or some third party. You can change it, and you can change it along with Claude, and it's just been a very cool thing to see internally.
One of the biggest shifts in how we've worked in the last year, I think, at Anthropic is so much of the software we interact with is built, maintained, and iterated on by Claude. Of course, reacting to all the things that are happening in the organization, but very much changing all the time. And so that leads to all sorts of other interesting questions, like who can update the data, right? Is it just the humans clicking through? Is it also Claude in the background? How do you have provenance of data? It opens up some interesting questions there as well.
But I feel like that is the piece, both the trend and the thing I'm most inspired by right now, is how do we make that the way almost all of our software operates? And of course, you still want building blocks and predictability and design systems, and you want to frame the cats. You want to make it not absolutely insane, but for the given task, for that given project, for even that given user potentially, what is the way in which we pull in the right ways and make that really feel both personal and extremely useful because of that?
And how does someone think about maybe you're not in Anthropic and you have a SaaS app that's scaled, and you're thinking about, okay, do I put an agent in it? Do I make it open and agent native so that Claude can be in it? How do you think about those choices and what a company should care about versus leave to a model company like Anthropic to plug in?
Yeah, I think a lot of it is really building the right primitives into the product, almost like an infrastructural layer. And you can tell. I always poke at, whenever I'm using new products, I'll ask it to do something that a human could do, but might require somebody to think through or architect the thing where everything goes through some shared plumbing that both agents and whatever sort of REST API you've built can go through as well. And you can tell when they've been thought through that way.
I have a lot of empathy because if you've been building an application for 20 years, just bolting that on is difficult. So I think it's a journey that a lot of companies are going on. But once you get it right, I think that question becomes, you can just evolve it over time, right? You can start with, okay, at first I'm just going to have maybe more of the side pane thing. I'm not ready to blow up my whole UI. I'm not ready to create fully generative UIs. But maybe you start experimenting with, all right, there's also maybe a personal landing page that can be more malleable and is using the same primitives. You're not inventing a whole new infrastructure underneath. So I think that you can evolve your product along the way, but you can tell when things feel super bolted on and feeling like it's not a native part of the product.
So Ami, I'm curious. We're hearing we can go all the way to agent native, malleable software, but at Anthropic you have to serve this whole spectrum of people. There's this challenge of pushing the frontier, and also some people just want to use Claude to chat, or you have big enterprise contracts where you can't just rip out their whole interface and then say, hey, it's malleable, in one day. Or maybe you can, I don't know, but I assume you can't. So how do you balance, in terms of thinking about product, pushing the frontier and creating a consistent experience for users that they can understand, and bringing along the people who are maybe not going to try all the new things on day one?
I think it's a really hard problem. Our approach is to just try to meet everyone where they are. I think we're just so early in understanding what works. It's not like stuff is stable and we can just build the right answer. We're all trying to explore together. And part of everything changing every couple months is the models change, but also the way people use products and the models changes, and we constantly have to update for that.
And so a thing that I think about is when we're just exploring a product like Mike's talking about, it's okay to just try stuff. We just have to try a bunch of things because we haven't actually explored the frontier. And then when we figure out what works, then we have to make it make sense for all the people who want something that feels more stable and complete.
I think this is a really tough trade-off because, raised in the product world that I was, I was all about simplicity and reliability and total predictability, because I think that helps people build a muscle memory and an intuition. But I also think the reality is we just don't know yet what a lot of things should look like. And so I actually encourage the team to just take a lot of shots on goal. It's okay to have things that feel somewhat overlapping or duplicative as long as, when we figure them out, when we hit product-market fit, we go back and we reincorporate into something that works for everyone.
Because I feel like the fail state is you decide from the beginning, based on a theory of exactly what needs to happen, and you overconstrain a product. And I would much rather have five overlapping products that all work in different ways than one overly constrained one that we never even gave a chance to actually meet the market.
And how do you—and I definitely see this, we do this at Every, but I see this all around the industry—just doing lots of parallel experiments and seeing which one works, and then starting to consolidate as you start to get signal. How do you do that in a way that doesn't feel confusing or incoherent? Or if I'm working on the third team doing this thing, it can be maybe demotivating. How do you do that well?
I think there's a couple of things. One, we were just having this conversation this morning internally where there were two potential product directions, and one has a team with very high conviction going after it, and the other one doesn't really. And it feels very manufactured to be like, we should explore this other idea because we think it's good, but nobody's really calling for it. Then that's that kind of B-team feeling. I've never been successful tasking a team to go take a product area that they're not excited about. The product is going to suck at the end of that process, right? So you want the excitement. But if there are two, three different teams that have exciting directions, then yeah, let that explore.
I think we've made a lot of progress on, even in the last six months, having the underlying infrastructure to support all of that. So we inventoried this. It feels like six months ago. It's probably only three, but there was a point where chat and Cowork had different memory systems and different MCP and different ways of storing files. And it was like, okay, for anybody to experiment on top of the surface, they're going to have a really hard time, because if they built a third thing that then doesn't share memory with any of those, for example, it's going to feel disconnected and almost disadvantaged from the beginning.
So a lot of the work—we have a foundations team that has been just thinking about, you should be able to get your memory wherever you are within our products. And it sounds like a simple thing, but then it's very clearly, you pull back the thing, oh, it's complicated for these eight reasons. But it's worth solving that, and solving it in a different team than your labsy innovation or frontier product teams, because then those experiments can feel complementary even if they are overlapping too.
I also think this sounds so simple, but I also think it's just really important to talk about it and just name, hey, if there's three different bets, sometimes it can feel demoralizing or like one of you is not set up for success. That's okay. That's the world we're living in. And so what we're asking you to do is come along this journey with us. We don't know the answer. We don't think anyone knows the answer. We're all going to discover it. And we're all fundamentally on the same team. And so we will just learn from whatever works. And just saying stuff like that out loud, I think, is really important because everybody is trying to figure everything out in the dark.
What do you look for to know that something's starting to work?
I think a bunch of that is the normal product-market fit. Do people use it? Do they like it? Do they come back? Do they talk about it? Do they get value? Again, going back to the goal of product has remained the same. It's to solve problems. Thinking about, does this actually solve a problem in a way that people can use and use easily and feel good about using? It comes back to the basics for that.
One of the interesting things that feels challenging to me about that is sometimes a thing looks like it doesn't work, and the new model comes out, and then it really starts working. And you don't want to throw it away too early, but you also don't want to be too far ahead. So how do you do that?
One of the things we do in labs sometimes is just park a project. When we built the first computer use product inside Anthropic in 2024, it was so bad because the models just weren't there yet. Our first thesis was, well, it's not going to be good at a bunch of automation, but it might be good about education. So you'd be like, "All right, how do I change this thing in some complex software? How do I do this really complex thing in Photoshop?" And it would get there, but it would get there in the very early agent Claude way, which was like, "This, did that work? Nope, that didn't work. I got to close this." It was like 20 minutes later, "I did it." You're like, "I learned nothing." So it was neither good for automation nor education.
But we parked it, and then we would basically just try it with every new model. And the way we would try it is we actually just had it running in an eval harness. And the way we actually learned that 3.7 was a leap in computer use was just seeing it suddenly took over. And then we actually looked at the transcript. We're like, wow, it's succeeding more often than it's not. This is a real moment.
So it is really valuable. We see it as a win in labs if all your project did is show that the models aren't ready for it yet, but you're able to now externalize that as an eval or create some connectivity to the research team. That's great because it means that two, three, four, five, six months from now we'll be able to revisit, and hopefully we've learned something along the way as well. But I think you need to build early so that you can even have that intuition.
I definitely agree. I've been calling this—you can get capability blind if you try something and then you don't ever try it again, or you don't make an eval, and you're like, oh, the models are just never going to do this. And then it just turns out a model comes out in three months and it just completely changes. It improves in ways you never would have thought it would.
I'm actually curious, if you're here, how many of you have used Fable? Raise your hand. Oh, a lot of people are using—okay, this is great. What about Astra? Fewer. Interesting. Okay, so this is great. You guys are on the edge. This is very helpful because one of the things I find a lot is I talk to someone and they're like, "Oh, it doesn't do this." And I'm like, "Have you used the latest model?" And they're like, "No, I used Sonnet 4.5," or whatever. And that really changes things.
I think that's also part of why it's so important to be willing to throw away everything you know.
Yeah.
Even though it's hard, because then otherwise you just think that the last thing that happened is going to be true forever.
One big question that I see you guys working on from afar, and I see this happening at OpenAI too, and we run into it a little bit, is when you're doing all these parallel experiments and one starts to really work, then you have to push it back into your actual product. And there are several different ways to do it. It's like, okay, you can make a tab, but then you have a bunch of different tabs, and all the things only work in slightly different ways, and that's kind of ugly. Or you can have it just be a separate app and you never merge them. There's all these different things that I don't think anyone has really come to, this is the answer, this is how you do it. But it seems like a critical problem is how to take those parallel experiments and turn it into a coherent product. What have you learned, and how do you think about it?
I don't know what you're talking about.
You're right. I don't think we've fully figured it out. But I think some of it comes back to what Mike was talking about, of making sure that we're building primitives so that our products feel like the way they should feel when you're using them. What does a user expect? The user expects to walk up to something and hopefully not have to think about it too much. And right now it's easy to just transfer a bunch of cognitive load to the user and be like, here's some tools, why don't you choose? And so one of the things that Mike's talking about is how do we make sure that we build primitives so the right system knows you? So a UI that feels really accessible all the time. And I think all of that is a bit iterative as we see how people react.
I think overall, if I had to bet, there's some lessons you'll learn and some that are real. The thing that we always saw at Instagram, there's a very strong power law in tab usage. It makes sense, every single product, right? Main feed was 80% of the use, right? And you can make Explore better, and maybe you would go from 10 to 15%, but it was main feed. That is where you land. That is your big thing. So I think that will always be the journey that we're all on, which is what is the main tab? What is that experience? How do you still make that great? And then luckily it can be an entry point to a lot of other downstream interesting experiences.
But yeah, I think half of the job is saying no to another thing in the sidebar or another tab or another thing, and figuring out, can this be rolled in once we've proven it out, but not killing it so early? So this is the entire art of product development, I think, as things are moving this quickly.
All right, we've got time for one last question, which I'd love to hear from both of you, which is if we are on this stage here next year and Lenny invites us back. I don't know if he'll invite me back. I've been cursing too
Much, but he'll invite you two back. What do you expect will be different, and how will being a PM have changed?
I hope that people feel like they can just do much more. They can just go and build in whatever way they want. They can get answers whenever they want. They can start businesses, and we just see a lot more small teams and solo empowerment, and we start to get the uplift of what that means for everyone else. I think builders magnify impact for a bunch more people who will never use the tools directly but who can benefit from people who do build things. And so I hope that's just way easier.
Yeah. I think the biggest one for me is, right now, I often talk about the gap between what the models are capable of and what most folks are using them for. Not through any fault of their own. It's on us. We got to build the products to make that possible. And that was a hard thing in '24. It was even harder last year. It's getting even harder now.
And so the biggest thing I want to do is close—I think it's very democratizing if you can close that gap, because it means it's not just your most Claude-pilled software engineer that has this fully working multi-agentic setup. It's also the person that is doing a bunch of research in parallel. It's also the person that's trying to manage their business, but they feel like they've unlocked it, and we've brought them along the way in terms of how do you understand the moving parts, check in at the right places, do long-horizon work, but not in a way that feels disempowering. That's the art. If we succeed a year from now, that'll be closer.
Amazing. Thank you so much, and Mike, to see you.
Article published · Updated
