Rethinking the Sutton Interview: Why Imitation Learning and RL May Not Be Opposites
Dwarkesh PatelAfter a widely discussed interview with Richard Sutton, Dwarkesh returned to the conversation to lay out what they now understand Sutton's worldview to be, and where they still disagree. Their central claim is that the distinctions Sutton draws between LLMs and "true" intelligence are less mutually exclusive than they appear. Imitation learning and reinforcement learning, models of humans and world models, and in-context learning and continual learning may be continuous with one another. Dwarkesh nonetheless ends by granting that Sutton's critique identifies real gaps in today's models.
The Steelman of Sutton's Position
Dwarkesh begins by reconstructing Sutton's argument as charitably as possible, and apologizes in advance for any remaining misunderstandings. The starting point is Sutton's essay The Bitter Lesson. On Dwarkesh's reading, the essay does not argue for throwing as much compute as possible at a problem. It argues for techniques that leverage compute most effectively and scalably.
By that standard, LLMs look poor. Most of the compute spent on an LLM goes into running it during deployment, yet the model learns nothing during that period. It learns only during a special phase called training, which Dwarkesh, voicing Sutton's view, calls "obviously not an effective use of compute." Training is itself highly inefficient, since models typically consume the equivalent of tens of thousands of years of human experience.
All of that learning also comes from human data. This is obvious for pretraining, but Dwarkesh argues it is also roughly true of reinforcement learning from verifiable rewards (RLVR). The RL environments are "human furnished playgrounds" built to teach LLMs skills that people have prescribed for them. The agent is not learning in any substantial way from organic, self-directed engagement with the world. Human data is an inelastic, hard-to-scale resource, so relying on it is not a scalable use of compute.
A further objection is that LLMs do not learn a true world model, meaning a model of how the environment changes in response to the agent's actions. They learn a model of what a human would say next, which leads them to depend on human-derived concepts. Dwarkesh offers an illustration: an LLM trained on all data up to 1900 probably would not come up with relativity from scratch.
The most fundamental reason, in this reconstruction, to expect the paradigm to be superseded is that LLMs cannot learn on the job. Continual learning will need a new architecture. Once it exists, no separate training phase will be needed, because the agent will learn on the fly as humans and all animals do. That paradigm would make the current LLM approach, with its sample-inefficient training phase, obsolete.
The Core Disagreement
Dwarkesh summarizes their position in three claims. First, imitation learning is continuous with and complementary to RL. Second, models of humans can provide a prior that makes it easier to learn "true" world models. Third, they would not be surprised if some future form of test-time fine-tuning could replicate continual learning, since in-context learning already achieves something like it to a degree.
Pretraining as Fossil Fuel
During the interview, Dwarkesh repeatedly asked Sutton whether pretrained LLMs could serve as a prior on which experiential learning, or RL, could be accumulated to reach AGI. To make the case now, they draw on a talk by Ilya Sutskever that compared pretraining data to fossil fuels, an analogy Dwarkesh says has "remarkable reach."
Fossil fuels are not renewable, but using them did not put civilization on a dead-end track. They were crucial. Dwarkesh argues there was no way to go directly from the water wheels of 1800 to solar panels and fusion power plants. A cheap, convenient, plentiful intermediary was needed to get to the next step. The implication is that pretraining data could play the same role for AI even if it eventually runs out or is superseded.
AlphaGo, AlphaZero, and the Role of Human Data
Dwarkesh then turns to Go. AlphaGo was conditioned on human games, and AlphaZero was bootstrapped from scratch. Both were superhuman, and AlphaZero was better. Dwarkesh asks two questions. Will we, or the first AGIs, eventually develop a general learning technique that needs no initial knowledge and bootstraps itself? And will it outperform the best AIs trained up to that point? They think the answer to both is probably yes.
That does not mean, in their view, that imitation learning can play no role in building the first AGI or even the first ASI. AlphaGo was still superhuman despite being "initially shepherded" by human data. Dwarkesh's reading is that human data is not actively harmful. At sufficient scale it simply stops helping much. They also note that AlphaZero used much more compute than AlphaGo.
Human Cultural Learning as an Analogy
Dwarkesh argues that accumulating knowledge over tens of thousands of years has been essential to humanity's success. In any field, thousands and probably millions of earlier people built up understanding and passed it on. We did not invent the language we speak or the legal system we use, and most of the technology in our phones was not invented by people alive today. Dwarkesh says this process looks more like imitation learning than RL from scratch.
They do not claim humans literally predict the next token. Human imitation differs from the supervised learning used to pretrain LLMs. But humans are also not "running around trying to collect some well defined scalar reward." No ML regime perfectly describes human learning, and humans do things analogous to both RL and supervised learning. Dwarkesh offers an analogy: "What planes are to birds, supervised learning might end up being to human cultural learning."
Imitation Learning as Short-Horizon RL
Dwarkesh also argues that the two techniques are not categorically different. On this framing, imitation learning is short-horizon RL in which each episode is one token long. The LLM conjectures the next token based on its understanding of the world and of how the pieces of information in the sequence relate. It then receives reward in proportion to how well it predicted.
They anticipate the objection that this is not ground truth, only a record of what a human was likely to say, and they agree. But they think a different question matters more for scalability: can imitation learning help models learn better from ground truth? Dwarkesh says the answer is "obviously yes." RL applied to pretrained base models has produced gold-medal performance at the IMO and the ability to build entire working applications from scratch. These are ground-truth tests. Either the model solves the unseen olympiad problem or it doesn't, and either the application matches the feature request or it doesn't.
Dwarkesh stresses that these results could not have been reached by RL from scratch, "or at least we don't know how to do that yet." A reasonable prior over human data was needed to kick-start the RL process. Whether that prior counts as a proper "world model" or just a model of humans strikes Dwarkesh as largely a semantic debate. What matters is whether the model of humans helps the system start learning from ground truth and so become a "true" world model.
Their analogy: telling an LLM builder to drop pretraining is like telling someone pasteurizing milk, "Hey stop boiling that milk because we eventually want to serve it cold!" The goal may be cold milk, but boiling is an intermediate step toward it.
Do LLMs Have a World Model?
Dwarkesh argues that LLMs are clearly developing a deep representation of the world, because training incentivizes them to. As evidence, they cite their own experience using LLMs to learn about biology, AI, and history, which they say the models teach with "remarkable flexibility and coherence."
They concede that LLMs are not specifically trained to model how their own actions affect the world. But if their representations cannot be called a "world model," Dwarkesh argues, then the term is being defined by the process assumed necessary to build one, rather than by the capabilities the concept implies.
Continual Learning and the Information Bottleneck
Dwarkesh then returns to continual learning, calling it their "hobby horse" and joking that they are like a comedian with only one good bit who plans to milk it. An LLM trained with RL on outcome-based rewards learns on the order of one bit per episode, and an episode can be tens of thousands of tokens long.
Animals and humans clearly extract far more information from their environment than an end-of-episode reward signal. Dwarkesh's conceptual picture is that we learn to model the world through observation. An outer RL loop incentivizes another learning system to extract as much signal from the environment as possible. They note that in Sutton's OaK architecture, this component is called the transition model.
Fitting this into modern LLMs would mean fine-tuning on all observed tokens. But Dwarkesh reports hearing from researcher friends that the most naive version of this does not work well in practice. They consider high-throughput continual learning from the environment necessary for true AGI and say it clearly does not exist in LLMs trained with RLVR.
Could Continual Learning Be Added to LLMs?
Even so, Dwarkesh thinks there may be relatively simple ways to "shoehorn" continual learning onto LLMs. One idea is to make supervised fine-tuning (SFT) a tool call available to the model. The outer RL loop would then reward the model for teaching itself effectively through supervised learning, so it can solve problems that don't fit in its context window.
Dwarkesh says they are "genuinely agnostic" about how well such techniques will work and notes that they are not an AI researcher. Still, they would not be surprised if these techniques basically replicate continual learning. Their reason is that models already show something resembling human continual learning within their context windows. In-context learning emerged spontaneously from the training incentive to process long sequences. That suggests to Dwarkesh that if information could flow across windows longer than the current context limit, models might meta-learn the same flexibility they already show in context.
Conclusion: Opposite Paths, and Sutton's Enduring Critique
Dwarkesh closes with a contrast. Evolution performs meta-RL to produce an RL agent, and that agent can selectively do imitation learning. LLM development runs the other way. A base model that does pure imitation learning comes first, followed by the hope that enough RL will turn it into a coherent agent with goals and self-awareness. "Maybe this won't work!" they acknowledge.
Even so, Dwarkesh does not think first-principles arguments, such as the claim that LLMs lack a true world model, prove much. They also doubt those arguments are strictly accurate for current models, which undergo substantial RL on "ground truth."
Dwarkesh still credits Sutton's critique. Even if Sutton's "Platonic ideal" is not the path to the first AGI, Dwarkesh says the critique identifies genuine basic gaps: no continual learning, abysmal sample efficiency, and dependence on exhaustible human data. These gaps are so pervasive in the current paradigm that people barely notice them, and Dwarkesh suggests they are obvious to Sutton because of his decades-long perspective. Their final prediction is that LLMs will probably reach AGI first, but the successor systems those AGIs build will "almost certainly be based on Richard's vision."
Boy do you guys have a lot of thoughts about the Sutton interview. I've been thinking about it myself and I think I have a much better understanding now of Sutton's perspective than I did during the interview itself. So I wanted to reflect on how I understand his worldview now. Richard, apologies if there's still any errors or misunderstandings. It's been very productive to learn from your thoughts.
Here's my understanding of the steelman of Richard's position. Obviously he wrote this famous essay, The Bitter Lesson. What is this essay about? It's not saying that you just want to throw away as much compute as you possibly can. The bitter lesson says that you want to come up with techniques which most effectively and scalably leverage compute.
Most of the compute that's spent on an LLM is used in running it during deployment. And yet it's not learning anything during this entire period. It's only learning during this special phase we call training. That is obviously not an effective use of compute. What's even worse, this training period by itself is highly inefficient, these models are usually trained on the equivalent of tens of thousands of years of human experience.
What's more, during this training phase, all of their learning is coming straight from human data. This is an obvious point in the case of pretraining data. But it's even kind of true for the RLVR that we do with these LLMs: these RL environments are human furnished playgrounds to teach LLMs the specific skills we have prescribed for them. The agent is in no substantial way learning from organic and self-directed engagement with the world. Having to learn only from human data, which is an inelastic and hard-to-scale resource, is not a scalable way to use compute.
Furthermore, what these LLMs learn from training is not a true world model, which would tell you how the environment changes in response to different actions that you take. Rather, they are building a model of what a human would say next. And this leads them to rely on human-derived concepts. A way to think about this would be, suppose you trained an LLM on all the data up to the year 1900. That LLM probably wouldn't be able to come up with relativity from scratch.
And here's a more fundamental reason to think this whole paradigm will eventually be superseded. LLMs aren't capable of learning on-the-job, so we'll need some new architecture to enable this kind of continual learning. And once we do have this architecture, we won't need a special training phase — the agent will just be able to learn on-the-fly, like all humans, and in fact, like all animals are able to do. And this new paradigm will render our current approach with LLMs — and their special training phase that's super sample inefficient — totally obsolete.
That's my understanding of Richard's position. My main difference with Rich is just that I don't think the concepts he's using to distinguish LLMs from true intelligence are actually that mutually exclusive or dichotomous. For example, I think imitation learning is continuous with and complementary to RL. Relatedly, models of humans can give you a prior which facilitates learning "true" world models. I also wouldn't be surprised if some future version of test-time fine-tuning could replicate continual learning, given that we've already managed to accomplish this somewhat with in-context learning.
Let's start with my claim that imitation learning is continuous with and complementary to RL. I tried to ask Richard a couple of times whether pretrained LLMs can serve as a good prior on which we can accumulate the experiential learning (aka do the RL) which will lead to AGI.
Ilya Sutskever gave a talk a couple of months ago that I thought was super interesting, and he compared pretraining data to fossil fuels. I think this analogy has remarkable reach. Just because fossil fuels are not a renewable resource does not mean that our civilization ended up on a dead-end track by using them. In fact they were absolutely crucial. You simply couldn't have transitioned from the water wheels of 1800 to solar panels and fusion power plants. We had to use this cheap, convenient and plentiful intermediary to get to the next step.
AlphaGo (which was conditioned on human games) and AlphaZero (which was bootstrapped from scratch) were both superhuman Go players. Of course AlphaZero was better. So you can ask the question, will we, or will the first AGIs, eventually come up with a general learning technique that requires no initialization of knowledge and that just bootstraps itself from the very start? And will it outperform the very best AIs that have been trained to that date? I think the answer to both these questions is probably yes.
But does this mean that imitation learning must not play any role whatsoever in developing the first AGI, or even the first ASI? No. AlphaGo was still superhuman, despite being initially shepherded by human player data. The human data isn't necessarily actively detrimental. It's just that at enough scale it just isn't significantly helpful. AlphaZero also used much more compute than AlphaGo.
The accumulation of knowledge over tens of thousands of years has clearly been essential to humanity's success. In any field of knowledge, thousands (and probably millions) of previous people were involved in building up our understanding and passing it on to the next generation. We obviously didn't invent the language we speak, nor the legal system we use. Also, most of the technologies in our phone were not directly invented by the people who are alive today. This process is more analogous to imitation learning than it is to RL from scratch.
Now, of course, are we literally predicting the next token, like an LLM would, in order to do this cultural learning? No, of course not. Even the imitation learning that humans are doing is not like the supervised learning that we do for pretraining LLMs. But neither are we running around trying to collect some well defined scalar reward. No ML learning regime perfectly describes human learning. We're doing things that are both analogous to RL and to supervised learning. What planes are to birds, supervised learning might end up being to human cultural learning.
I also don't think these learning techniques are categorically different. Imitation learning is just short horizon RL. The episode is a token long. The LLM is making a conjecture about the next token based on its understanding of the world and how the different pieces of information in the sequence relate to each other. And it receives reward in proportion to how well it predicted the next token.
Now, I already hear people saying: "No no, that's not ground truth! It's just learning what a human was likely to say." And I agree. But there's a different question which I think is more relevant to understanding the scalability of these models: can we leverage this imitation learning to help models learn better from ground truth?
And I think the answer is, obviously yes? After RLing the pre-trained base models we've gotten them to win Gold in IMO competitions and to code up entire working applications from scratch. These are "ground truth" examinations. Can you solve this unseen math olympiad question? Can you build this application to match a specific feature request? But you couldn't have RLed a model to accomplish these tasks from scratch. Or at least we don't know how to do that yet. You needed a reasonable prior over human data in order to kick start this RL process.
Whether you want to call this prior a proper "world model", or just a model of humans, I don't think is that important and honestly seems like a semantic debate. Because what you really care about is whether this model of humans helps you start learning from ground truth – AKA become a "true" world model. It's a bit like saying to someone pasteurizing milk, "Hey stop boiling that milk because we eventually want to serve it cold!" Of course. But this is an intermediate step to facilitate the final output.
By the way, LLMs are clearly developing a deep representation of the world, because their training process is incentivizing them to develop one. I use LLMs to teach me about everything from biology to AI to history, and they are able to do so with remarkable flexibility and coherence. Now, are LLMs specifically trained to model how their actions will affect the world? No, they're not. But if we're not allowed to call their representations a "world model," then we're defining the term "world model" by the process we think is necessary to build one, rather than by the obvious capabilities the concept implies.
Continual learning. Sorry to bring up my hobby horse again. I'm like a comedian who's only come up with one good bit, but I'm gonna milk it for all it's worth. An LLM being RLed on outcome-based rewards learns on the order of 1 bit per episode, and an episode may be tens of thousands of tokens long.
Obviously, animals and humans are clearly extracting more information from interacting with our environment than just the reward signal at the end of each episode. Conceptually, how should we think about what is happening with animals? I think we're learning to model the world through observations. This outer loop RL is incentivizing some other learning system to pick up maximum signal from the environment. In Richard's OaK architecture, he calls this the transition model.
If we were trying to pigeonhole this feature spec into modern LLMs, what you'd do is to fine tune on all your observed tokens. From what I hear from my researcher friends, in practice the most naive way of doing this actually doesn't work well. Being able to continuously learn from the environment in a high throughput way is obviously necessary for true AGI. And it clearly doesn't exist with LLMs trained on RLVR.
But there might be some relatively straightforward ways to shoehorn continual learning atop LLMs. For example, one could imagine making SFT a tool call for the model. So the outer loop RL is incentivizing the model to teach itself effectively using supervised learning, in order to solve problems that don't fit in the context window. I'm genuinely agnostic about how well techniques like this will work—I'm not an AI researcher. But I wouldn't be surprised if they basically replicate continual learning.
Models are already demonstrating something resembling human continual learning within their context windows. The fact that in-context learning emerged spontaneously from the training incentive to process long sequences makes me think that if information could flow across windows longer than the current context limit, models could meta-learn the same flexibility that they already show in-context.
Some concluding thoughts. Evolution does meta-RL to make an RL agent. That agent can selectively do imitation learning. With LLMs, we're going the opposite way. We first made a base model that does pure imitation learning. And we're hoping that we do enough RL on it to make a coherent agent with goals and self-awareness. Maybe this won't work!
But I don't think these super first-principle arguments (for example, about how these LLMs don't have a true world model) are actually proving much. I also don't think they're strictly accurate for the models we have today, which are undergoing a lot of RL on "ground truth".
Even if Sutton's Platonic ideal doesn't end up being the path to first AGI, his first principles critique is identifying some genuine basic gaps these models have. We don't even notice because they are so pervasive in the current paradigm, but because he has this decades-long perspective they're obvious to him. It's the lack of continual learning, it's the abysmal sample efficiency of these models, it's their dependence on exhaustible human data. If the LLMs do get to AGI first, which is what I expect to happen, the successor systems that they build will almost certainly be based on Richard's vision.
Article published
