Sergey Levine on the Robot Flywheel: Why General-Purpose Robots May Arrive Within Five Years

Open on YouTube ↗
Overview

Sergey Levine is a co-founder of Physical Intelligence, a company building robotics foundation models, and a professor at UC Berkeley. In this conversation with podcast host Dwarkesh Patel, the central question is when robots will be widely deployed and doing real work, and what has to happen technically and industrially to get there. Levine's position is that the key date is not when robotics is "done" but when a self-sustaining data flywheel starts. He thinks that could begin within a year or two, and he gives five years as a reasonable median for robots capable enough to run a household. He also stresses how much remains uncertain about which data, methods and hardware will get there.

31 min read

Where Physical Intelligence Is Today

Levine describes the company's goal as building general-purpose models that could in principle control any robot to perform any task. He sees robotics as encompassing essentially all of AI: a truly general robot could do a large share of what people can do.

About a year after founding, Levine says the company has "built out a lot of the basics," and that these basics work well. Their robots can fold laundry, go into a new home and try to clean the kitchen, and, as the host saw on a visit, fold a cardboard box with simple grippers. Levine cautions against reading too much into these demos. When results are compressed into a three-minute video, viewers may think folding a T-shirt is the goal. For Levine, those tasks mainly confirm the initial hypothesis that the foundations are solid.

The real target is much larger. Levine describes the ideal "prompt" as a standing instruction: make dinner at 6 p.m., I leave for work at 7 a.m., have laundry ready on Saturday, check in every Monday about the shopping list. The robot would then carry this out for six months or a year. Asked to make a particular salad, it should work out what that involves, look up the recipe and buy the ingredients.

Meeting that bar requires common sense, recognising edge cases that call for more careful thought, continuous improvement, an understanding of safety, reliability when it matters, and the ability to fix its own mistakes. Levine says the underlying principles are leveraging prior knowledge and having the right representations.

Timelines: Dating the Flywheel, Not the Finish Line

Asked for a percentile estimate, Levine says he does not expect everything to be developed in a lab and then shipped as "a robot in a box" sometime in the 2030s. He expects robotics to follow the path of AI assistants. Once robots reach a basic level of usefulness they will be deployed, and deployment will generate experience that makes them better. So he thinks about timelines as the date when the flywheel starts.

That start could be soon, he says, but it involves a tradeoff: the more narrowly a product is scoped, the earlier it can be deployed. Physical Intelligence is already exploring which real tasks could start the flywheel. For something people would actually care about, he calls "single-digit years" very realistic and hopes it will be "more like one or two," while stressing it is hard to say. "Out there" means a robot doing a task people actually want done, competently enough to do it for real.

For the fully capable household robot, Levine again says single-digit years rather than double-digit. The uncertainty comes from open research questions. He does not think these require profoundly new ideas, but they do require the right synthesis of what is already known, and he notes synthesis can be as hard as inventing something new. If the work goes as planned and "we're a bit lucky," single digits is reasonable. When the host narrowed it down, Levine said "five is a good median."

Why the Flywheel Hasn't Obviously Spun for LLMs, and Why Robots Might Have an Edge

The host pointed out that LLMs are already widely deployed without producing an obvious flywheel in which models learn every job in the economy. Levine thinks this is close to working and is "100% certain" that many organisations are working on it. He argues there already is a flywheel with humans in the loop, since everyone deploying an LLM watches what it does and adjusts it. Automating this is hard because it requires deriving the right supervision signals and grounding them in the system's behaviour, and the details around algorithms and stability get "pretty gnarly." He does not consider it a profoundly impossible problem.

He sees no deep reason robotics should be fundamentally different, but some small differences help. When a robot works alongside people, there are natural sources of supervision, and people have strong incentives to help it succeed. Physical tasks also involve frequent, visible mistakes that can be recovered from and learned from. If an assistant answers a question wrongly, the user may never notice. If a robot messes up folding a T-shirt, "it's pretty obvious."

Scope, Productivity and the Economic Question

Levine compares the gap between a first deployable robot and a fully autonomous housekeeper to the history of coding assistants. Early tools completed functions and got about half right. Today the best ones can put together most of a pull request for fairly formulaic work. Robots will be given growing scope in the same way: first making the coffee, eventually running the coffee shop.

The host pressed on economic impact. If a robot can run a house, it could presumably do most blue-collar work. He also noted that AI company revenues are around $20–30 billion a year, compared with $30–40 trillion of knowledge work, so LLMs capture roughly a thousandth of it by revenue. Would robots be in a similar place in five years?

Levine said he was not prepared to estimate what share of physical labour robots could do, because he does not have a sufficient understanding of that broad a cross-section of work. He also pushed back on the framing that a switch flips and workers are replaced. With coding tools, the biggest productivity gains go to experts who use them. He expects the same with robots: robot plus human will outperform either alone. This arrangement also makes bootstrapping easier, because people can help and give hints, not just label data.

He illustrated this with a story from the π0.5 project, whose paper was released the previous April. The team initially controlled robots by teleoperation. Once the model was good enough, they realised they could make significant progress by instructing it in language, standing there and saying "pick up the cup, put the cup in the sink." Those words gave the robot information it could learn from. Levine expects future systems to also learn by observing people and from the natural feedback of working together. Understanding that kind of interaction is where prior knowledge from large pretrained models becomes especially valuable.

Why This Won't Be Another Self-Driving Saga

The host raised self-driving cars: Google's effort started around 2009, and only now are autonomous vehicles actually deployed. Why wouldn't a cool robotics demo in five years be followed by another decade of waiting?

Levine's first answer concerns 2025 versus 2009, not robotics versus driving. Perception in 2009 was poor. It is the kind of problem where an engineered system can produce an excellent demo and then hit a brick wall when you try to generalise. Today there is much better technology for generalisable, robust perception and world understanding. He notes that in machine learning, "scalable" really means "generalisable."

His second answer concerns the problem itself. Robotic manipulation is in some ways much harder than driving, but it is easier to get started with limited scope. Nobody would let a teenager learn to drive alone, let alone a five-year-old, yet most people would let a child try the dishes without supervision even though dishes can break. Many manipulation tasks let you make a mistake, correct it, complete the task, and learn to avoid the mistake next time. In driving, mistakes have serious consequences, so that loop is hard to run.

For manipulation tasks that are safety-critical, Levine points to common sense: making reasonable guesses about what might happen without having to experience the mistake first. He says this was essentially unknown five years ago, but VLMs can now answer questions like what happens if you walk over a floor marked slippery, which no 2009 autonomous car could have done. Common sense plus the ability to make and correct mistakes resembles how people learn. It doesn't make manipulation easy, but it lets robotics start small and grow.

What Changed Relative to Earlier Industrial Efforts

Asked why companies like Google and Meta hit roadblocks despite years of transformers and large datasets, Levine first corrected the premise. They made a lot of progress, and much of Physical Intelligence's work builds on it; many of the founders were at Google and took part. What was missing was framing. Those efforts were fundamental research. Making robotic foundation models work also requires industrial-scale building, "more like the Apollo program than … a science experiment": deploying robots, collecting representative data at scale, building the systems, and focusing on the foundation model for its own sake rather than for papers or a research lab.

Why not just hire 100 times more operators and collect 100 times more data? Levine says the challenge is knowing which axes of scale drive which capabilities. Going from 10 tasks to 100 can be solved by scaling horizontally. Practical usefulness also requires high robustness, speed and efficiency, and intelligent handling of edge cases. Those can be scaled too, but only once the team knows what data to collect, in what settings, and with what methods. He says they don't fully know this yet, expect to figure it out "pretty soon," and want to get it right so scaling translates directly into practically relevant capability.

On data volume, Levine says comparison with internet text is difficult because robot time steps are highly correlated: the raw bytes are enormous but information density is low. Compared with datasets used for multimodal training, the last count he recalls put them within one to two orders of magnitude. Whether robotics ultimately needs as much experience as language is unknown. He finds it more useful to ask how much data is needed to start a self-sustaining, ever-growing data collection process in which acquiring the data is itself useful work. He would ideally like that to be reinforcement learning with the robot acting autonomously, but says mixed autonomy is plausible, given the wide range between fully teleoperated and fully autonomous robots.

How the π0 Model Works

Levine describes the current model as a vision-language model adapted for motor control. Using a "fanciful brain analogy," a VLM is an LLM with a small pseudo visual cortex (a vision encoder) grafted on. Physical Intelligence's models add an action expert, essentially an action decoder, which serves as a small motor cortex.

The model reads the robot's sensory input and does internal processing that can include intermediate steps. Told to "clean up the kitchen," it might reason that it needs to pick up the dish and the sponge, working through a chain of thought down to the action expert, which outputs continuous actions. The action expert is a separate module because actions are continuous, high-frequency and in a different format from text, but the whole system is still an end-to-end transformer, technically resembling a mixture-of-experts architecture. The host described the stream as alternating text, images and actions. Levine agreed, with the caveat that actions are not discrete tokens: they are produced with flow matching and diffusion, because dexterous control needs precision.

The host noted that the model builds on Google's open-source Gemma, so robotics and language models share not just techniques but literally the same weights. Levine said what makes these building blocks valuable is that the AI community has become much better at leveraging prior knowledge. Pretrained LLMs and VLMs supply somewhat abstract knowledge about the world, such as identifying objects and roughly locating them in images. In one sentence, the main benefit recent AI brings to robotics is the ability to use that prior knowledge.

Video Models, Representations, and the Value of Having a Purpose

The host relayed an argument from Sander, a researcher at Google DeepMind who works on video and audio models: there is little transfer between modalities because text is represented at a high semantic level, while images and video are essentially compressed pixels.

Levine said he had bad news and good news. The bad news is that this touches a long-running problem. Getting intelligence from video prediction is an older idea than getting it from text prediction, yet text became practically useful first. Video generation is impressive, but it hasn't produced systems with deep world understanding that can do much beyond generating images and video. He explained why with a thought experiment. Point a camera out of the building and try to predict everything: you could study crowd psychology to predict pedestrians, or the physics of water molecules and ice particles to predict clouds, and you could spend decades going down to the subatomic level without ever getting to the rest. Text has already been abstracted into the parts humans care about.

The good news is that a robot has a purpose, and its perception serves that purpose. Levine points to psychology experiments showing people have striking tunnel vision, failing to see things directly in front of them that aren't relevant to their goal. He reasons that such a strong focusing mechanism must be very important for achieving goals, and that robots will have it because they are trying to achieve something.

On whether robots can simply learn from all of YouTube, Levine used an analogy. If you watched sports recordings for a year and were then told to play tennis, that would be "pretty dumb." If you were told first that you would play tennis and then studied, you would know what to look for. He does not want to understate the difficulty and says embodiment is not a silver bullet, but embodied models that learn from interaction may be better at absorbing other data sources. They have already seen that including web data in robot training helps generalisation, and he suspects in the long run embodiment will make hard-to-use data sources easier to exploit.

Emergence Through Composition

The host worried that because robot data is collected deliberately, robots won't get the unexpected emergent abilities LLMs get from internet text, and every subtask of a barista's or waiter's job might need thousands of manually collected episodes.

Levine replied that emergent capabilities do not just come from the breadth of internet data. They also come from generalisation becoming compositional past a certain level. His example, one a student liked to use, is the International Phonetic Alphabet, which is used almost exclusively to write the pronunciations of individual words in dictionaries. An LLM can nonetheless write a whole recipe in IPA, something it has almost certainly never seen, by composing familiar pieces in a new way. With enough diversity of behaviours, a robot model should likewise work out how to combine them as situations demand.

He says they have already seen small instances of this. While experimenting with laundry-folding policies, a robot accidentally picked up two T-shirts. It began folding one, and when the other got in the way it picked it up and threw it back in the bin. Nobody had expected this, and on further testing it did so every time. Objects dropped on the table were put back; a shopping bag that tipped over was set upright. No one had collected data specifically for these behaviours, though Levine assumes someone at some point picked up a bag, accidentally or intentionally. He expects this kind of compositionality, combined with language and chain-of-thought reasoning, to give the model much more room to recombine skills. He added that five years from now these examples will probably look tiny.

The host added his own example: during his visit he turned a pair of shorts inside out, and the robot, with just two finger-like grippers, first turned them right side out and then folded them correctly.

One Second of Context and Moravec's Paradox

The host found this surprising because the model sees only about one second of context, whereas language models consider hundreds of thousands of tokens. How can one second support a minute-long task?

Levine said less memory isn't a virtue, and adding memory, longer context and higher-resolution images would make the model better. The reason it isn't the first priority for these skills is Moravec's paradox, which he calls the one thing to know about robotics: in AI, the easy things are hard and the hard things are easy. Picking up objects and perceiving the world are the hard problems, while chess and calculus are often easier.

He argues the memory question is "Moravec's paradox in disguise." Tasks that feel cognitively demanding, like solving a hard math problem or holding a technical conversation, are the ones that require keeping many pieces in mind. A well-practised skill, like an Olympic swimmer's stroke, is performed "in the moment," baked into the neural network through practice. Matching human dexterity means getting those foundations right first and then moving up the stack into reasoning, context and planning, which will also matter.

The Trilemma of Speed, Context and Model Size

The host framed a trilemma. Inference speed, context length and model size all consume inference compute, and the current model (roughly 100 ms inference, about a second of context, a couple of billion parameters) is orders of magnitude short of the human brain on at least two of these. How can all three improve together?

Levine called the representation of context a very interesting technical problem likely to see much innovation. People keep some things symbolically: his shopping list is just "milk," not an image of the milk shelf. Other things are spatial and visual, like what he expected the street and doorway to look like on the way to the studio. Representing context in a form that keeps what matters for the goal and discards the rest is essential. He thinks multimodality involves much more than image plus text, including potentially learned modalities suited to a task, both for representing the past and for plans and intermediate reasoning.

On whether the brain's efficiency comes from hardware or algorithms, Levine said he doesn't know and is not well-versed in neuroscience. His guess is that the brain is extremely parallel, even more so than a GPU. A multimodal model reads images, then text, then generates one token at a time, whereas an embodied system seems better suited to parallel processes. He notes transformers are fundamentally parallelisable, made sequential by position embeddings, so a system that handles perception, proprioception and planning simultaneously need not look mathematically very different from a transformer, though its implementation would differ. It might run long-term memory, short-term spatial information, semantics, current perception and planning in parallel at different rates, with complex processes running slower and reactive ones faster.

Off-Board Inference and Planning Ahead

Asked what will have happened in five years to make human-level systems physically possible, Levine said he is not a systems expert. He imagines affordable robots may externalise part of their thinking: with a poor internet connection a robot might fall back to a "dumber reactive mode," and with a good one it could be smarter. On the algorithmic side, because sensory streams are highly temporally correlated, each new frame adds much less information than the whole frame, so much more compressed representations should be possible. He said he hasn't tackled the systems problem yet, since it makes sense to build the system once the shape of the ML solution is known.

On whether robots will rely on centralised, batched inference like LLMs, Levine guessed both models will exist: low-cost systems with off-board inference, and costlier, more reliable systems with onboard inference where connectivity can't be counted on, such as outdoor robots.

He added that real-time control may require surprisingly little thinking per time step. Recordings from monkey brains show neural correlates of planning before a movement, and the shape of the movement correlates with that earlier activity. Planning sets initial conditions for a process that then unrolls, so processing is front-loaded. It is not open-loop playback: you still react, but at a more basic level of abstraction. The design question is which representations suffice for planning ahead and which need a tight feedback loop. A driver might track lane markers closely while assessing traffic at a lower frequency.

Why Imitation Learning Comes Before RL

The host noted that Levine's earlier lectures argued RL is often better than imitation learning in robotics, yet current models are trained by imitation. Levine said the key is prior knowledge. Learning effectively from your own experience requires already knowing something about the task; otherwise it takes far too long, as with a child learning to write. Supervised training now builds the foundation that lets models learn much faster later. He points to LLMs, which began with next-token prediction that later enabled synthetic data generation and RL, and expects any foundation model effort to follow that path: build the foundation "in a somewhat brute-force way," then improve it with more accessible training.

Asked whether the best knowledge-work model in 10 years will also be a robotics model, Levine said he "really hope[s]" they will be the same, acknowledging he is "extremely biased." He hopes robotics will make everything else better, for two reasons. First, the focus that comes from doing a task structures perception so other signals can be used more fruitfully. Second, deep physical understanding beyond what language can articulate helps with abstract problems: people say a company "has momentum" or "my computer hates me," using embodied and social experience as a hammer for abstract nails. The host suggested inference and size constraints might differ between the two uses, but the same model might be served in different ways.

Simulation, Synthetic Data and Counterfactuals

Why doesn't simulation work better, when pilots and F1 drivers learn in simulators? Levine said pilots are extremely goal-directed: their aim is to fly the real plane, and they know passengers' lives will depend on it. A model trained on several domains just sees multiple things to master. A better analogy is someone who plays a flight video game to master the game and is later put in a real cockpit.

The host raised meta-learning, referencing a 2017 paper by Levine, and suggested making real-world performance the loss for time spent in simulation. Levine said all such ideas depend on being able to train for the real thing. He suspects this may not need to be explicit: meta-learning emerges, as LLMs show through in-context learning. Large models trained on the right objective with real data become much better at using everything else. LLMs trained for hard problems use a lot of synthetic data, but they can do so because they start from models trained on lots of real data. "Perhaps ironically," the key to using simulation is to get very good at using real data first.

On whether future AGIs could rehearse entirely new skills in simulation, Levine said that at a fundamental level, self-generated synthetic experience doesn't teach you more about the world. It enables rehearsal and counterfactuals, but information about the world has to come in from somewhere. Classical robotics treated simulation as a way to inject human knowledge through hand-written equations. Experience from video generation and LLM synthetic data suggests the most powerful synthetic experience comes from a very good learned model, which may know more fine-grained detail than a person, but got that knowledge from the world. Viewed as a black box, information goes in and capability comes out, whether internally via simulation or model-free methods.

Asked about a human analogue, Levine noted that the sleeping brain looks much like the waking brain, replaying experience or generating statistically similar experience, so learned simulation plausibly helps the brain work out counterfactuals. More fundamentally, optimal decision-making always requires answering "if I did this instead of that, would it be better?", whether by a learned simulator, a value function or a reward model. The key is not good simulation but answering counterfactuals.

Can Robots Help Build the AI Buildout?

The host cited estimates of 100–300 gigawatts of AI capacity by 2030, implying $2–4 trillion of capex per year for data centres, chip fabs and solar factories, and asked whether robots would be mature enough to help or whether human labour would be the bottleneck.

Levine said that in principle robots could help a lot, but warned against thinking of them as mechanical people. A better analogy is a car or a bulldozer: much lower maintenance, deployable in odd places, of any shape or size, 100 feet tall or tiny. Intelligence able to drive very heterogeneous robots could do better than mechanical people. Speaking as a non-expert on data centres, he gave the example that robots could build data centres in remote locations without caring whether there's a shopping centre nearby.

On hardware costs, he gave his own history. In 2014 he used a PR2 research robot that cost $400,000. For his Berkeley lab he bought arms at $30,000. Physical Intelligence's arms cost about $3,000 each, and the company thinks they can be made for a small fraction of that. He attributes this to economies of scale, better actuator technology, and software: smarter AI lowers hardware requirements, since factory robots need highly repeatable precision that cheap visual feedback can make unnecessary. On whether arms will cost hundreds of dollars by decade's end, he deferred to co-founder Adnan Esmail, but said the cost declines have surprised him year after year. He didn't know how many suitable arms exist; asked if it was under 100,000, he said "probably," noting that car-factory robots aren't really the relevant kind. On scaling production, he observed that economies are good at meeting demand, citing how few iPhones existed in 2001, while acknowledging the challenge.

He sees his job as finding the "minimal package." Robots shouldn't break constantly, but some questions remain open, such as how many fingers are needed, and extreme precision is probably unnecessary because feedback compensates. He expects no single ultimate robot, but rather a few requirements all good robots share, like touchscreens on smartphones, plus optional features by niche and price point. Once capable AI can plug into any robot, many people can innovate on hardware. There is no "Nvidia of robotics" now, and he would like to see a heterogeneous ecosystem. His main hardware concerns are reliability and cost, mainly because cost determines robot count and therefore data volume. Current AI isn't pushing hardware to its limits, and clearer answers will come once it does.

If Hardware Is the Bottleneck, Does China Win by Default?

The host observed that much of the supply chain behind the AI buildout, from solar wafers to robot arms, is manufactured in China. If hardware is the bottleneck, why doesn't China win by default?

Levine started broadly. An economy that advances through a highly educated, highly productive workforce benefits greatly from automation, which multiplies each person's output, as LLM coding tools do for software engineers. Getting there involves complex decisions about making the transition appealing to society, handling geopolitics, and investing in a balanced robotics ecosystem supporting both software and hardware. He doesn't see these as insurmountable, and is optimistic because the end state of highly productive, well-educated people doing high-value work seems very compatible with automation, which he calls the light at the end of the tunnel pointing in the right direction.

On how hundreds of millions of robots would actually be manufactured in the US or allied countries, Levine said he was probably not the most qualified to answer. One ingredient is that producing robots is itself physical work, so good robotics should help build robots. That circularity needs bootstrapping, but he considers it easier than with digital devices, where computers and phones don't help build themselves. The host countered that the feedback loop might be stronger in China, where the supply chain already exists. Levine returned to balance: AI is exciting but not the only thing to get right. Physical Intelligence takes hardware seriously, builds many of its own components and keeps a hardware roadmap alongside its AI roadmap, but the US, and arguably civilisation, must think about these problems holistically. He said it is easy to be distracted by rapid progress in one area, and he wished for more holistic conversations.

Planning for Full Automation, or for the Journey

The host proposed that society should plan for full automation. Humans contribute through body and mind with "no secret third thing," so after a boom period in which human labour is especially valuable, the end state is full automation in a much wealthier society with some form of redistribution.

Levine called this a reasonable way to look at things directionally, but said technology rarely evolves as expected, and that makes planning for an end state very difficult. The journey matters as much as the destination, and automation will probably show up first in unexpected places. The constant he emphasised is education, which he called the best buffer against the negative effects of change and the single most important lever society can pull.

The host questioned this: by Moravec's paradox, what education gives humans may be the easiest thing to automate, since a model can absorb eight years of grad-school textbooks in an afternoon. Levine replied that education provides flexibility, which is less about the facts you know than about your ability to acquire skills and understanding, adding that it has to be a good education.