Rethinking the Sutton Interview: Why Imitation Learning and RL May Not Be Opposites

Open on YouTube ↗
Overview

After a widely discussed interview with Richard Sutton, Dwarkesh returned to the conversation to lay out what they now understand Sutton's worldview to be, and where they still disagree. Their central claim is that the distinctions Sutton draws between LLMs and "true" intelligence are less mutually exclusive than they appear. Imitation learning and reinforcement learning, models of humans and world models, and in-context learning and continual learning may be continuous with one another. Dwarkesh nonetheless ends by granting that Sutton's critique identifies real gaps in today's models.

12 min read

The Steelman of Sutton's Position

Dwarkesh begins by reconstructing Sutton's argument as charitably as possible, and apologizes in advance for any remaining misunderstandings. The starting point is Sutton's essay The Bitter Lesson. On Dwarkesh's reading, the essay does not argue for throwing as much compute as possible at a problem. It argues for techniques that leverage compute most effectively and scalably.

By that standard, LLMs look poor. Most of the compute spent on an LLM goes into running it during deployment, yet the model learns nothing during that period. It learns only during a special phase called training, which Dwarkesh, voicing Sutton's view, calls "obviously not an effective use of compute." Training is itself highly inefficient, since models typically consume the equivalent of tens of thousands of years of human experience.

All of that learning also comes from human data. This is obvious for pretraining, but Dwarkesh argues it is also roughly true of reinforcement learning from verifiable rewards (RLVR). The RL environments are "human furnished playgrounds" built to teach LLMs skills that people have prescribed for them. The agent is not learning in any substantial way from organic, self-directed engagement with the world. Human data is an inelastic, hard-to-scale resource, so relying on it is not a scalable use of compute.

A further objection is that LLMs do not learn a true world model, meaning a model of how the environment changes in response to the agent's actions. They learn a model of what a human would say next, which leads them to depend on human-derived concepts. Dwarkesh offers an illustration: an LLM trained on all data up to 1900 probably would not come up with relativity from scratch.

The most fundamental reason, in this reconstruction, to expect the paradigm to be superseded is that LLMs cannot learn on the job. Continual learning will need a new architecture. Once it exists, no separate training phase will be needed, because the agent will learn on the fly as humans and all animals do. That paradigm would make the current LLM approach, with its sample-inefficient training phase, obsolete.

The Core Disagreement

Dwarkesh summarizes their position in three claims. First, imitation learning is continuous with and complementary to RL. Second, models of humans can provide a prior that makes it easier to learn "true" world models. Third, they would not be surprised if some future form of test-time fine-tuning could replicate continual learning, since in-context learning already achieves something like it to a degree.

Pretraining as Fossil Fuel

During the interview, Dwarkesh repeatedly asked Sutton whether pretrained LLMs could serve as a prior on which experiential learning, or RL, could be accumulated to reach AGI. To make the case now, they draw on a talk by Ilya Sutskever that compared pretraining data to fossil fuels, an analogy Dwarkesh says has "remarkable reach."

Fossil fuels are not renewable, but using them did not put civilization on a dead-end track. They were crucial. Dwarkesh argues there was no way to go directly from the water wheels of 1800 to solar panels and fusion power plants. A cheap, convenient, plentiful intermediary was needed to get to the next step. The implication is that pretraining data could play the same role for AI even if it eventually runs out or is superseded.

AlphaGo, AlphaZero, and the Role of Human Data

Dwarkesh then turns to Go. AlphaGo was conditioned on human games, and AlphaZero was bootstrapped from scratch. Both were superhuman, and AlphaZero was better. Dwarkesh asks two questions. Will we, or the first AGIs, eventually develop a general learning technique that needs no initial knowledge and bootstraps itself? And will it outperform the best AIs trained up to that point? They think the answer to both is probably yes.

That does not mean, in their view, that imitation learning can play no role in building the first AGI or even the first ASI. AlphaGo was still superhuman despite being "initially shepherded" by human data. Dwarkesh's reading is that human data is not actively harmful. At sufficient scale it simply stops helping much. They also note that AlphaZero used much more compute than AlphaGo.

Human Cultural Learning as an Analogy

Dwarkesh argues that accumulating knowledge over tens of thousands of years has been essential to humanity's success. In any field, thousands and probably millions of earlier people built up understanding and passed it on. We did not invent the language we speak or the legal system we use, and most of the technology in our phones was not invented by people alive today. Dwarkesh says this process looks more like imitation learning than RL from scratch.

They do not claim humans literally predict the next token. Human imitation differs from the supervised learning used to pretrain LLMs. But humans are also not "running around trying to collect some well defined scalar reward." No ML regime perfectly describes human learning, and humans do things analogous to both RL and supervised learning. Dwarkesh offers an analogy: "What planes are to birds, supervised learning might end up being to human cultural learning."

Imitation Learning as Short-Horizon RL

Dwarkesh also argues that the two techniques are not categorically different. On this framing, imitation learning is short-horizon RL in which each episode is one token long. The LLM conjectures the next token based on its understanding of the world and of how the pieces of information in the sequence relate. It then receives reward in proportion to how well it predicted.

They anticipate the objection that this is not ground truth, only a record of what a human was likely to say, and they agree. But they think a different question matters more for scalability: can imitation learning help models learn better from ground truth? Dwarkesh says the answer is "obviously yes." RL applied to pretrained base models has produced gold-medal performance at the IMO and the ability to build entire working applications from scratch. These are ground-truth tests. Either the model solves the unseen olympiad problem or it doesn't, and either the application matches the feature request or it doesn't.

Dwarkesh stresses that these results could not have been reached by RL from scratch, "or at least we don't know how to do that yet." A reasonable prior over human data was needed to kick-start the RL process. Whether that prior counts as a proper "world model" or just a model of humans strikes Dwarkesh as largely a semantic debate. What matters is whether the model of humans helps the system start learning from ground truth and so become a "true" world model.

Their analogy: telling an LLM builder to drop pretraining is like telling someone pasteurizing milk, "Hey stop boiling that milk because we eventually want to serve it cold!" The goal may be cold milk, but boiling is an intermediate step toward it.

Do LLMs Have a World Model?

Dwarkesh argues that LLMs are clearly developing a deep representation of the world, because training incentivizes them to. As evidence, they cite their own experience using LLMs to learn about biology, AI, and history, which they say the models teach with "remarkable flexibility and coherence."

They concede that LLMs are not specifically trained to model how their own actions affect the world. But if their representations cannot be called a "world model," Dwarkesh argues, then the term is being defined by the process assumed necessary to build one, rather than by the capabilities the concept implies.

Continual Learning and the Information Bottleneck

Dwarkesh then returns to continual learning, calling it their "hobby horse" and joking that they are like a comedian with only one good bit who plans to milk it. An LLM trained with RL on outcome-based rewards learns on the order of one bit per episode, and an episode can be tens of thousands of tokens long.

Animals and humans clearly extract far more information from their environment than an end-of-episode reward signal. Dwarkesh's conceptual picture is that we learn to model the world through observation. An outer RL loop incentivizes another learning system to extract as much signal from the environment as possible. They note that in Sutton's OaK architecture, this component is called the transition model.

Fitting this into modern LLMs would mean fine-tuning on all observed tokens. But Dwarkesh reports hearing from researcher friends that the most naive version of this does not work well in practice. They consider high-throughput continual learning from the environment necessary for true AGI and say it clearly does not exist in LLMs trained with RLVR.

Could Continual Learning Be Added to LLMs?

Even so, Dwarkesh thinks there may be relatively simple ways to "shoehorn" continual learning onto LLMs. One idea is to make supervised fine-tuning (SFT) a tool call available to the model. The outer RL loop would then reward the model for teaching itself effectively through supervised learning, so it can solve problems that don't fit in its context window.

Dwarkesh says they are "genuinely agnostic" about how well such techniques will work and notes that they are not an AI researcher. Still, they would not be surprised if these techniques basically replicate continual learning. Their reason is that models already show something resembling human continual learning within their context windows. In-context learning emerged spontaneously from the training incentive to process long sequences. That suggests to Dwarkesh that if information could flow across windows longer than the current context limit, models might meta-learn the same flexibility they already show in context.

Conclusion: Opposite Paths, and Sutton's Enduring Critique

Dwarkesh closes with a contrast. Evolution performs meta-RL to produce an RL agent, and that agent can selectively do imitation learning. LLM development runs the other way. A base model that does pure imitation learning comes first, followed by the hope that enough RL will turn it into a coherent agent with goals and self-awareness. "Maybe this won't work!" they acknowledge.

Even so, Dwarkesh does not think first-principles arguments, such as the claim that LLMs lack a true world model, prove much. They also doubt those arguments are strictly accurate for current models, which undergo substantial RL on "ground truth."

Dwarkesh still credits Sutton's critique. Even if Sutton's "Platonic ideal" is not the path to the first AGI, Dwarkesh says the critique identifies genuine basic gaps: no continual learning, abysmal sample efficiency, and dependence on exhaustible human data. These gaps are so pervasive in the current paradigm that people barely notice them, and Dwarkesh suggests they are obvious to Sutton because of his decades-long perspective. Their final prediction is that LLMs will probably reach AGI first, but the successor systems those AGIs build will "almost certainly be based on Richard's vision."