Why Claude Can't Be Your PM (Yet): Anthropic's Product Leaders on Judgment, Agents, and Parallel Bets

Open on YouTube ↗
Overview

At the Lenny and Friends Summit, Dan Shipper put a question to Anthropic's Ami Vora and Mike Krieger: a year ago people were saying AI might make product managers obsolete, and now it seems everyone is supposed to become one. Which is it? Both argued that the PM role has not gone away. The underlying job is unchanged, but the pace of change raises the bar for adaptability, judgment, and operational discipline. From there the discussion moved to agent-native software, running many overlapping experiments, and folding the ones that work back into a coherent product.

17 min read
1:10

The Human Side Changes Slowly, the Technology Changes Every Two Months

Vora described product as "a bridge between the real problems that people have in the world and the technology that you can use to solve it." That definition hasn't changed. What has changed is the pace. Technology used to shift every five or ten years, and now it feels like every two months, which Vora said is hard for a person to absorb. Human problems, in contrast, stay roughly the same. Her approach is to stay obsessed with the human problem while keeping a "fresh-eyed" view of the technology, which means being willing to throw away most of what she knows about what used to work and try again with everything she knows about the problem.

She compared the current moment to "rewinding the clock" to when she started in product. Back then nobody had clear definitions of the job, and you were a general problem solver who kept going until you hit a wall, then asked someone or figured it out. She sees roles blurring again in the same way. The difference is that there is now a lot of scaffolding available when you do hit a wall.

2:44

Mike Krieger's "Before and After" Moment

Krieger told a story from his own work. Earlier in the year he moved into an individual-contributor role and has mostly been building. As his project got close to shipping, a PM lead named Cat, who works with Vora, pulled him aside and told him the project needed a PM. His first reaction was doubt: "Claude's got it." She insisted. Once a PM joined, Krieger saw all the glue and connective tissue that would have been dropped without one. A week later he messaged Cat to say she had been "100% right."

He listed what that work covered. Anthropic serves everyone from prosumers to very large enterprises, so someone has to bring all of those constituencies along. Someone has to make sure the customer success team can talk about changes as questions come up in real time. Someone has to loop in safeguards, which he called very important. Someone has to keep people on track. Nobody has infinite capacity, even with AI assistance. A person in this role lets builders stay heads-down "in full-on Claude mode" instead of switching to the very different activity of making sure everything goes well. Because teams now move faster, Krieger said, the role demands more operational excellence than before. If you are recruiting a team, this hat is "increasingly important."

Why Not Just Have Fable Do It?

Shipper pushed back and asked why Krieger didn't just have Fable, Anthropic's model, do the PM work. Krieger acknowledged that Anthropic runs on a lot of Claude-provided connective tissue. For example, Claude can flag that something relevant is happening in another part of the organization, surfaced through search. But he said Claude is not yet a convener. In a recent talk on the role of humans in a world of very powerful AI, he described several archetypes, and one was the convener: someone who still has to bring the Claudes and the people together to get work done.

He joked that the turning point will come when Claude schedules a meeting for him and, when he asks who set it up, explains that it thought two people should talk. That hasn't happened yet. He added that this is "not because it can't," and suggested it may just be a matter of turning on the setting.

Vora added a second reason the role matters. The tools make it possible to build a lot very quickly, which expands the universe of what could be built. You still need good judgment about what to build, often on limited information. With so many forking paths, you also have to be even more relentless about making sure the right thing actually happens. She summed up the attitude this way: I know the problem, I have a feedback loop with the users, and I'm going to go for it. She called that relentlessness and judgment in the face of ambiguity really important.

7:15

Skills That No Longer Matter

Shipper asked Vora about her earlier remark that she spent decades learning to answer certain questions she no longer needs to answer. Her example was the mix of product and design judgment she worked hard to build: imagining a user's situation, where their finger would land on the phone, where a button should go. Reviews used to revolve around exactly those questions because building and shipping was so expensive. She doesn't think anyone will ask her that again. It is now faster to build three versions and try them.

She said this is hard on identity. She compared it to the innovator's dilemma, where companies keep doing what they're good at even when something else might matter more, and said the same thing happens to individuals. Part of the technology changing every two months is that you have to throw away what you knew about yourself in the job. She admitted she has never worked at a model company before and often doesn't know whether she'll be good at the next thing. Every couple of months she has to try something new and see if she can figure it out.

9:55

Framing the Chaos, and Giving Someone the Pen

Asked what lets people do this well, Vora named adaptability and a higher tolerance for change, along with judgment and relentlessness. For her team, she stressed talking about these things openly, because otherwise constant adaptation can feel lonely when in fact everyone is going through it. She talks about "framing the chaos" so it feels less taxing. In uncertain times it is tempting to map out exactly which products will be built or exactly how a career will go as a way of exerting control. She thinks that locks you out of trying new things. The goal is to make engaging with chaos feel safe and plausible, and to acknowledge the emotions involved, because that is where much of the "magic and productivity" will be.

Krieger added an organizational counterweight: chaos needs clear ownership. In Anthropic's Labs group, exploratory efforts are called "bets," and each has a bet lead who is the directly responsible individual (DRI). That person decides whether to double down or wind something down, and whether a team needs more or fewer people. Because things move so fast, he said this role matters even more. "We're all lost together," but someone holds the pen on the next most important thing to de-risk, understand, or learn.

12:11

What Agent-Native Software Looks Like

Shipper then turned to what gets built, now that both humans and agents use software, sometimes through delegation and sometimes collaboratively. Krieger sketched a rough timeline. First, AI sat in a sidebar or mini window, disconnected from the product and perhaps useful for support questions. Then came deeper integrations with whole AI-powered features. Now there is a move toward being agent-native. He credited Every, Shipper's company, with a write-up on agent-native architectures. He fed it to Claude to create a skill while he worked on the topic this year. The core idea he took from it: everything a human can do, an agent should be able to do too. Very few products get this right, he said, including many of Anthropic's own. When a product does, emergent behavior appears. The agent can combine things in new ways or proactively suggest different approaches.

The next stage he is thinking about is interfaces that are themselves malleable by the agent. People have talked about malleable software for years, and he feels it is now becoming real. He gave an internal example. Anthropic is shipping a complex project with at least four independent workstreams. Claude monitors progress and also built the UI through which first the TPM and then the whole team see what's happening. If you don't like the display, you aren't stuck with what an internal team or third party built. You can change it with Claude. Krieger called this one of the biggest shifts in how Anthropic has worked over the past year: much of the software people interact with internally is built, maintained, and iterated on by Claude, and it changes all the time.

That raises open questions. Who can update the data: only humans clicking through, or Claude in the background too? How do you track data provenance? He still wants building blocks, predictability, and design systems, so things don't become "absolutely insane." But he is most inspired by the idea of tailoring software to a given task, project, or even user so it feels personal and extremely useful.

Advice for an Established SaaS Company

For a company with a scaled SaaS app deciding whether to add its own agent or open up to agents like Claude, Krieger recommended building the right primitives as an infrastructural layer. When he tries new products, he often asks the agent to do something a human could do, and he can tell whether everything runs through shared plumbing that serves both agents and the product's REST API. He expressed empathy for companies with 20-year-old applications, for whom bolting this on is difficult. Once the primitives are right, the product can evolve gradually. You might start with a side panel, then try a more malleable personal landing page on the same primitives, without blowing up the whole UI or inventing new infrastructure. The failure mode is AI that feels "super bolted on" rather than native.

17:09

Many Shots on Goal Versus a Stable Experience

Shipper asked Vora how Anthropic balances pushing the frontier with serving users who just want to chat, or enterprises whose interfaces can't be overhauled overnight. She called it a really hard problem. Anthropic's approach is to meet people where they are, because it's still very early. Nothing is stable enough to simply build "the right answer," and both the models and the ways people use them keep changing. During exploration it's fine to try many things. Once something works, it has to be made sensible for people who want stability and completeness.

She acknowledged tension with her own background, which valued simplicity, reliability, and predictability because they help users build muscle memory and intuition. Still, she encourages the team to take many "shots on goal" and to accept overlap and duplication, as long as they consolidate after reaching product-market fit. In her view, the failure mode is over-constraining a product from the start based on a theory. She would rather have "five overlapping products that all work in different ways" than one overly constrained product that never got a chance to meet the market.

19:34

Keeping Parallel Teams Motivated

Shipper asked how to run parallel experiments without them feeling incoherent or demoralizing, for example for the third team working on a similar idea. Krieger described a discussion from that same morning about two possible product directions. One had a team with high conviction and the other didn't. Assigning a team to pursue an idea nobody is calling for feels manufactured and creates a "B team" feeling. He has never succeeded by tasking a team with an area they weren't excited about, because the product ends up bad. When two or three teams each have exciting directions, he lets them explore.

He said the underlying infrastructure has improved a lot recently. At one point, by his estimate three to six months ago, Chat and Cowork had different memory systems, different MCP setups, and different file storage. Any new experiment built on top would have been disconnected and disadvantaged from the start, for example by not sharing memory. A foundations team has since worked on making memory available wherever you are across Anthropic's products. That sounds simple but turned out to be complicated "for these eight reasons." Solving it in a separate team from the Labs or frontier product teams lets overlapping experiments feel complementary.

Vora added that simply naming the situation matters. If there are three bets, it can feel demoralizing, or like one team isn't set up to succeed. She tells teams that this is the world we're in, that nobody knows the answer, that everyone is on the same team, and that they will learn from whatever works. Saying this out loud, she said, is important because everyone is "trying to figure everything out in the dark."

On how to tell whether something is working, Vora pointed to ordinary product-market-fit signals. Do people use it, like it, come back, talk about it, and get value? Does it solve a problem in a way people can use easily and feel good about?

22:21

Parking Projects and Avoiding "Capability Blindness"

Krieger raised a complication: something may look like it doesn't work, and then a new model makes it work. You don't want to kill it too early or be too far ahead. Labs sometimes parks projects. His example was the first computer-use product Anthropic built internally in 2024, which he said was "so bad because the models just weren't there yet." The first thesis was that even if it couldn't automate much, it might help with education, such as showing how to do something complex in Photoshop. In practice it would flail through trial and error ("did that work? Nope") and eventually succeed 20 minutes later, leaving the user having learned nothing. It worked neither for automation nor education.

The team parked it and kept it running in an eval harness, retrying with each new model. With the 3.7 model, they saw its performance suddenly jump. Reading the transcripts, they realized it was "succeeding more often than it's not," which they recognized as a real moment. Krieger said Labs counts it as a win when a project only shows the models aren't ready yet, as long as the work is turned into an eval or connected to the research team, so it can be revisited months later. You have to build early to develop that intuition.

Shipper called the opposite failure "capability blindness": trying something once, never retrying or writing an eval, and assuming models will never do it, until a model three months later improves in unexpected ways. He polled the audience. Many had used Fable, and fewer had used Astra. He said he often meets people who say a model can't do something and then admit they last tried an older model such as Sonnet 4.5. Vora tied this back to her theme: you have to be willing to throw everything away, or you'll assume the last thing that happened stays true forever.

24:47

Turning Experiments Into a Coherent Product

Shipper noted that Anthropic, OpenAI, and his own company all face the problem of merging a successful experiment back into the main product. Options include adding a tab, which leads to many tabs that each work slightly differently, or keeping a separate app that never merges. Nobody seems to have settled the answer.

Vora joked, "I don't know what you're talking about," then admitted they haven't fully figured it out. Part of her answer returned to primitives. Users expect to walk up to something without having to think too hard. It's easy to shift cognitive load onto them by handing over a pile of tools and letting them choose. The aim is primitives that let the system know the user, and a UI that always feels accessible, refined iteratively as people react.

Krieger drew on his Instagram experience, where tab usage followed a strong power law. The main feed accounted for about 80% of usage, and improving Explore might move it from 10% to 15%. He expects every product to face the same question of what the main tab is and how to make it great, with that experience serving as an entry point to other features. He described half the job as saying no to another sidebar item or tab, and working out whether something can be folded in once proven, without killing it too early. That, he said, is "the entire art of product development" when things move this fast.

27:16

Looking Ahead a Year

Asked what will be different if they return next year, Vora hoped people will feel they can do much more: build however they want, get answers when they need them, and start businesses. She hoped for more small teams and solo builders, and said builders magnify impact for many people who will never use the tools directly.

Krieger focused on the gap between what models can do and what most people use them for. He said that gap is not users' fault but Anthropic's responsibility to close through better products. He said it was hard in 2024, harder last year, and is getting harder still. Closing it would be democratizing. The fully working multi-agent setup would no longer belong only to the most "Claude-pilled" software engineer. It would also reach someone running research in parallel or someone managing their business. The challenge is helping them understand the moving parts, check in at the right moments, and do long-horizon work "not in a way that feels disempowering." He called that "the art," and said success a year from now would mean the gap is closer to closed.