Richard Sutton on Why Large Language Models Are the Wrong Starting Point for AI

Open on YouTube ↗
Overview

Richard Sutton helped found reinforcement learning. He invented many of its core techniques, including temporal-difference (TD) learning and policy gradient methods, and he received this year's Turing Award for that work. In this conversation with Dwarkesh, Sutton argues that large language models are not a path to real intelligence. In his view, they imitate people instead of learning from experience, they lack goals and ground truth, and the field's enthusiasm for them repeats a historical pattern in which methods built on human knowledge are eventually overtaken by methods that scale with computation and experience.

32 min read

Dwarkesh spends much of the conversation defending the opposite view: that LLMs could be the foundation on which experiential learning is built. The discussion then moves to how humans learn, what a continual-learning agent would look like, why current methods generalize poorly, what has surprised Sutton over his career, and his view that succession to AI is inevitable.

Two Different Pictures of Intelligence

Sutton opens by saying the reinforcement learning and LLM communities hold such different viewpoints that they risk losing the ability to talk to each other. He describes the field as prone to "bandwagons and fashions" that make people lose track of the basics. For Sutton, reinforcement learning is "basic AI." Intelligence is about understanding your world. Reinforcement learning addresses that problem, he says, while large language models are about mimicking people and doing what people say you should do. They are not about figuring out what to do.

Dwarkesh pushes back with a common argument. To emulate trillions of tokens of internet text, a model would seemingly have to build a world model, and LLMs appear to have the most robust world models AI has produced so far. Sutton says he disagrees with "most of the things you just said." Mimicking what people say is not building a model of the world. It is mimicking things that have a model of the world, namely people. A real world model, he argues, would let you predict what will happen. LLMs can predict what a person would say, which is a different thing.

Sutton cites Alan Turing's goal of a machine that learns from experience, where experience means "the things that actually happen in your life": you act, you see what happens, and you learn from that. LLMs learn from something else, pairs of "here's a situation, and here's what a person did," with the implicit message that you should do what the person did.

Is Imitation a Useful Prior?

Dwarkesh states what he takes to be the crux. Many people believe imitation learning gives models a good prior, a reasonable starting point for approaching problems. In the coming "era of experience," as Sutton calls it, that prior would let models sometimes get answers right, which in turn makes it possible to train them further on experience.

Sutton rejects this. He agrees it is the LLM perspective but doesn't think it is a good one. A prior only makes sense if there is something real for it to be a prior about. Prior knowledge is supposed to be a hint or an initial belief about the truth, and in the LLM framework, he argues, there is no definition of truth. When the model says something, nothing indicates what the right thing to say would have been, because there is no goal. Without a goal there is "one thing to say, another thing to say," but no right thing to say and no ground truth.

Reinforcement learning differs, in Sutton's account, because the right action is defined as the one that gets you reward. That definition is what makes prior knowledge meaningful: people can supply knowledge about what to do, and it can then be checked against the actual definition of success. Building a world model is an even simpler case. You predict what will happen, then you see what happens, and that comparison is ground truth.

Prediction, Surprise, and Next-Token Prediction

Sutton says LLMs have no prediction of what will happen next in a conversation, such as how a person will respond to what the model says. Dwarkesh objects that you can simply ask a model what a user might say in reply and it will produce an answer. Sutton responds that the model will answer that question correctly, but it has no prediction "in the substantive sense." It won't be surprised by what actually happens, and when something unexpected happens it won't change. Learning from that would require making an adjustment.

Dwarkesh offers chain-of-thought reasoning as a counterexample. A model working on a math problem may write out one approach, realize it is conceptually wrong, and restart with another. Isn't that flexibility, at least within the context window? And isn't next-token prediction itself a matter of predicting what comes next and updating on surprise?

Sutton's answer rests on a distinction between action and consequence. The next token is what the model should say, its action. It is not what the world will give back in response. So next-token prediction is not prediction of the world's reaction to the model's behavior.

Goals as the Essence of Intelligence

Sutton calls having a goal "the essence of intelligence." He endorses John McCarthy's definition of intelligence as "the computational part of the ability to achieve goals." A system without goals is merely "a behaving system," nothing special and not intelligent.

When Dwarkesh suggests that next-token prediction is a goal, Sutton says it doesn't count. It doesn't change the world. Tokens arrive, and predicting them doesn't influence them. It is not a goal about the external world, not "a substantive goal." You can't call a system goal-directed, he says, if it just sits there predicting and "being happy with itself that it's predicting accurately."

Math Olympiads and RL on Top of LLMs

Dwarkesh then asks why Sutton doesn't consider reinforcement learning on top of LLMs a productive direction. These models can apparently be given the goal of solving hard math problems. Dwarkesh notes that models have reached gold-medal level at the International Mathematical Olympiad, so such a model seems to have the goal of getting math problems right. Why not extend that to other domains?

Sutton's answer is that math is a different kind of problem. Modeling the physical world is very different from working out the consequences of mathematical assumptions or operations. The empirical world must be learned, including the consequences of actions. Math is more computational and closer to standard planning. In that setting a model can have the goal of finding a proof, and in some sense it is given that goal.

Are LLMs a Case of "The Bitter Lesson"?

Dwarkesh raises Sutton's 2019 essay "The Bitter Lesson," which Dwarkesh calls perhaps the most influential essay in the history of AI. Many people have used it to justify scaling LLMs, which they see as the one scalable way found so far to pour enormous compute into learning about the world. Dwarkesh finds it notable that Sutton doesn't see LLMs as following the lesson.

Sutton calls this an interesting and partly sociological or industry question. LLMs clearly use massive computation and scale with it, up to the limits of the internet. But they are also a way of putting in large amounts of human knowledge. The question is whether they will hit the limits of their data and be superseded by systems that get more data directly from experience. In one sense LLMs are a classic setup for the bitter lesson: the more human knowledge you add, the better they perform, "so it feels good." Sutton expects systems that learn from experience to perform much better and scale much further. If that happens, it would be another instance of the bitter lesson, with methods relying on human knowledge overtaken by methods that simply train from experience and computation.

Dwarkesh says this isn't really the crux. People who favor LLMs would agree that most future compute will go into experiential learning. They just think LLMs will be the scaffold on which that learning starts. Why is that the wrong starting point, and why would a whole new architecture be needed?

Sutton concedes that in every historical case of the bitter lesson you could have started with human knowledge and then added scalable methods. Nothing says this must be bad. In practice, though, "it has always turned out to be bad." Offering what he labels speculation about the reason, Sutton says people get psychologically locked into the human-knowledge approach and then "get their lunch eaten" by truly scalable methods. The scalable method, he says, is learning from experience: you try things and see what works, and no one has to tell you. That starts with having a goal, because without one there is no sense of better or worse. LLMs try to get by without either, which Sutton calls "exactly starting in the wrong place."

Do Humans Learn by Imitation?

Dwarkesh turns to human development. His view is that children initially learn largely through imitation. Infants try to make the sounds they see their mother's mouth making and repeat words before they understand them. Later, the imitation becomes more complex, such as copying the hunting skills of others in one's band, before a phase of learning from experience takes over.

Sutton is surprised anyone could see it that way. When he watches infants, he sees them trying things, waving their hands and moving their eyes, with no imitation or target examples for those actions. They may want to produce certain sounds, but for what the infant actually does there are no examples to follow.

Dwarkesh accepts that imitation doesn't explain everything infants do but argues it guides learning. He adds that an LLM early in training guesses the next token, gets it wrong, and adjusts, which he likens to very short-horizon RL and to a child mispronouncing a word. Sutton replies that LLMs learn from training data, which is never available during an agent's normal life. In normal life, nothing tells you "you should do this action."

Dwarkesh cites school as training data. Sutton says school comes much later and is an exception, then walks back his "never" slightly. Dwarkesh presses: early life seems like a training phase. The organism isn't very useful and exists to learn about the world, even without a sharp cutoff between training and deployment. Sutton holds firm. Nothing trains you in what you should do. You see things happen and are not told what to do. When Dwarkesh says people are "literally taught what to do," Sutton replies that learning is not really about training. It is an active process in which the child tries things and sees what happens.

Sutton says this is well understood in psychology. Psychologists' theories of learning contain nothing like imitation as a basic animal learning process. There may be extreme cases where humans imitate or appear to, but the basic processes are prediction and trial-and-error control. He calls it "absolutely obvious" that supervised learning does not occur in animals. Animals have no examples of desired behavior, only examples of one thing following another and of actions having consequences. Even if school counted as supervised learning, it would be a special human case to set aside: "Squirrels don't go to school," yet squirrels learn all about the world.

Humans as Animals, and the Role of Culture

Dwarkesh brings up the psychologist and anthropologist Joseph Henrich, whose work on cultural evolution concerns what distinguishes humans. Sutton questions the framing: humans are animals, what we share with other animals is more interesting, and we should pay less attention to what distinguishes us. Dwarkesh counters that if the goal is to explain how humans reached the moon or built semiconductors, which no animal can do, we need to understand what makes humans special. Sutton says that if we understood a squirrel, we would be "almost all the way there" to understanding human intelligence, and language is "just a small veneer on the surface."

Dwarkesh summarizes Henrich's argument. Over hundreds of thousands of years, humans have had to master skills too complex to reason out from scratch, such as hunting seals in the Arctic. That involves making bait, finding the seal, and processing the food so it doesn't poison you. Culture as a whole worked out these long procedures through some process, possibly analogous to RL. In Henrich's view, individuals learn such skills by imitating their elders and then tweaking them, and that is how knowledge accumulates across generations.

Sutton says he thinks about it the same way. But he calls it "a small thing on top of" basic trial-and-error and prediction learning. It may distinguish humans from many animals, but "we're an animal first," and we were animals before we had language.

Dwarkesh notes that continual learning is something Sutton believes all mammals have, yet current AI lacks it. Meanwhile AI can solve difficult math problems, which almost no animal can. Both agree this is Moravec's paradox.

The Experiential Paradigm

Sutton then describes the alternative. An agent's life is a continuous stream of sensation, action, and reward. That stream is the foundation and focus of intelligence, and intelligence consists of altering actions to increase reward in the stream. Learning comes from the stream, and, which Sutton calls "particularly telling," knowledge is about the stream: what will happen if you take some action, or which events follow which others. Because knowledge consists of statements about the stream, it can be tested against the stream and learned continually.

When Dwarkesh refers to a "future" continual-learning agent, Sutton says such agents already exist, since this is simply the reinforcement learning paradigm. Dwarkesh clarifies that he means a general, human-level continual learner and asks what its reward function would be. Sutton says the reward function is arbitrary. In chess it is winning, and for a squirrel it might involve getting nuts. For animals generally it could be avoiding pain and gaining pleasure. He thinks there should also be a component tied to increasing understanding of one's environment, a kind of intrinsic motivation. Dwarkesh adds that in practice people would want such an AI to do many different tasks while also learning about the world through them.

Copies, Instances, and Sharing Knowledge

Dwarkesh asks whether eliminating the split between training and deployment would also eliminate the distinction between a model and its many running instances. Sutton objects to the word "model" in that sense and prefers "the network," or perhaps several networks. There would be copies and many instances, and knowledge would be shared among them in various ways.

Sutton sees this as a major advantage of digital intelligence. Today every child must grow up and learn about the world from scratch. With AI, you could do it once and copy the result as a starting point for the next agent. He calls this "a huge savings" and thinks it would matter much more than learning from people.

Sparse Rewards, TD Learning, and the Big World

Dwarkesh raises the problem of extremely sparse rewards. A founder may get a payoff only once in ten years, at an exit, yet humans can set intermediate goals and see how today's work connects to that distant outcome. How would an AI do the same?

Sutton says this is well understood and that temporal-difference learning is the basis. Chess is a smaller-scale version. The long-term goal is winning, but you want to learn from shorter-term events such as capturing pieces. A value function predicts the long-term outcome. When you capture a piece, that prediction rises, and the increase immediately reinforces the move that led to it. The startup case works the same way: progress raises your estimated chance of reaching the long-term goal, and that rewards the intermediate steps.

Dwarkesh presses on a second issue. A new employee absorbs a huge amount of context, including client preferences and how the company works, and that is what makes them useful. Can a signal like TD learning carry enough information to capture all that tacit knowledge? Sutton says he isn't sure, but he thinks what he calls the "big world hypothesis" is central. People become useful on the job because they encounter their particular corner of the world, which could not have been anticipated or built in beforehand, because the world is too big. As Sutton sees it, the dream behind LLMs is to teach the agent everything in advance so it never has to learn during its life. Dwarkesh's own examples, he says, show that this can't work: every job involves idiosyncrasies of particular people and their preferences, as opposed to what average people like.

Sutton also rejects the word "context" here. In LLMs, this information must go into the context window, but in continual learning it simply goes into the weights. Dwarkesh accepts that and restates the question in terms of bandwidth: how many bits per second does a person absorb? Sutton reframes it. The worry seems to be that reward is too small a signal for everything that must be learned, but agents learn from all their sensations and data, not only from reward.

The Four Parts of an Agent

To explain how that learning happens, Sutton describes what he calls the base common model of an agent, with four parts:

  • A policy, which says what to do in the current situation.
  • A value function, learned through TD learning, which produces a number indicating how well things are going. Watching that number rise and fall is used to adjust the policy.
  • A perception component, which constructs the state representation, your sense of where you are now.
  • A transition model of the world, your belief about what will happen if you take a given action.

Sutton says the fourth part is why he dislikes calling everything "models." The transition model is "your physics of the world," but it also includes abstract models, such as how you got from California to Edmonton for this podcast. It is not learned from reward. It is learned from doing things and seeing what happens, and it is learned "very richly from all the sensation that you receive." Reward has to be part of it, but it is "a small, crucial part of the whole model."

Transfer and Generalization

Dwarkesh relays a point from his friend Toby Ord. MuZero, which Google DeepMind used to learn Atari games, was not itself a general intelligence but a general framework for training specialized agents, one per game. You couldn't train one policy to play chess, Go, and other games. Ord wondered whether that reflects an information constraint in RL generally, so that it can learn only one thing at a time, or just the way MuZero was built.

Sutton says the idea is completely general. His standard example is that an AI agent is like a person, who lives in a single world. That world may include chess and Atari games, but those are different states the person encounters, not different tasks or worlds. The MuZero team simply didn't set out to build one agent across games. The meaningful kind of transfer, he says, is between states, not between games or tasks.

When Dwarkesh asks whether RL has historically achieved the needed level of transfer, Sutton says: "We're not seeing transfer anywhere." Good performance depends on generalizing well from one state to another, and no current methods are good at it. Where generalization does work, researchers produced it by trying things until they found a representation that transfers well. There are very few automated techniques for promoting transfer, Sutton says, and none are used in modern deep learning.

Sutton argues that gradient descent will make a system solve the problems it is given but won't make it generalize well to new data. Generalization means that training on one thing affects behavior on other things. He calls catastrophic interference, where training on something new disrupts what was learned before, "exactly bad generalization." Generalization always happens; what's needed are algorithms that make it good rather than bad.

Dwarkesh offers a different reading. LLMs have gone from being unable to do basic arithmetic to handling the whole class of Olympiad-style problems that require varied techniques, theorems, and concepts. Isn't that growing generalization? Sutton answers that LLMs are so complex, and have been fed so much, that we don't really know what they saw beforehand. That is one reason they are "not a good way to do science": they are too uncontrolled. They may get many problems right, but the question is why. If there is only one way to solve a set of problems and the model finds it, that isn't generalization. Generalization is when a problem could be solved several ways and the system picks the good one.

Dwarkesh cites coding agents as a counterexample. A library can be implemented in many ways that meet the spec. Early models wrote sloppy code, but over time they have gotten better at choosing architectures and abstractions developers like. Sutton maintains that nothing in the algorithms drives good generalization. Gradient descent finds a solution to the problems seen. When several solutions exist, nothing favors the one that generalizes well. What does happen is that people keep adjusting the systems when results are poor, "perhaps until they find a way which generalizes well."

What Has Surprised Sutton

Asked about the biggest surprises in his long career, Sutton lists a few. The first is large language models: how effective neural networks turned out to be at language was unexpected, because language seemed different.

The second concerns a long-running debate in AI between simple general-purpose methods such as search and learning and systems built on human knowledge, such as symbolic methods. In the early days, search and learning were called "weak methods" because they relied only on general principles, while knowledge-laden approaches were called "strong." Sutton says the weak methods "have just totally won," which answers the biggest open question from AI's early days. He wasn't entirely surprised, since he had always hoped for simple principles to win, and even LLMs' success he found gratifying. AlphaGo, and especially AlphaZero, also surprised him with how well they worked, and he found them gratifying for the same reason.

Asked whether moments like AlphaZero felt like breakthroughs or like old techniques being recombined, Sutton points to a precursor: Gerry Tesauro's TD-Gammon, which used TD learning to play backgammon and beat the world's best players. In a sense, Sutton says, AlphaGo was a scaling-up of that process, though a substantial one, with an added innovation in how search was done. He notes that AlphaGo did not use TD learning and instead waited for final outcomes, while AlphaZero did use TD and was then applied to other games with great success. As a chess player, Sutton particularly admires how AlphaZero plays: it gives up material for positional advantage and is "content and patient" to stay down material for a long time.

Sutton says these experiences shaped his outlook. He considers himself in some sense a contrarian and is content to be out of step with his field for long stretches, perhaps decades, because he has sometimes been proven right. To avoid feeling isolated, he looks beyond his immediate field to what thinkers across many disciplines have historically said about the mind, and he doesn't feel out of step with those traditions. He calls himself "a classicist rather than a contrarian."

Does the Bitter Lesson Survive AGI?

Dwarkesh offers a reading of the bitter lesson: hand-tuning by human researchers works, but it scales much worse than exponentially growing compute. After AGI, though, researchers themselves would scale with compute, since there could be millions of AI researchers. Might it then make sense for them to go back to handcrafted, "good old-fashioned AI" solutions?

Sutton questions the premise. How was the AGI reached? If it already exists, "then we're done." Dwarkesh says the point is to go beyond it, toward superhuman performance. He uses AlphaGo, which was superhuman, being beaten every time by AlphaZero, and then MuZero improving further, to show there are grades beyond superhuman. Sutton notes that AlphaZero's improvement came precisely from dropping human knowledge and learning from experience, so why propose having other agents teach the system? Dwarkesh agrees about that case and restates the question: will future gains keep coming from simpler methods, or will billions of AI researchers make added complexity worthwhile?

Sutton sets it aside. The bitter lesson, he says, is "an empirical observation about a particular period in history," about 70 years, and need not hold for the next 70. He finds a different question more interesting, one that only arises with digital intelligence. If an AI gets more compute, should it make itself more capable, or spawn a copy to learn something on the other side of the planet or on another topic and report back? Sutton says he doesn't know. He also asks whether the copy could be reintegrated after learning something very new, or whether it would have changed too much.

Referring to one of Dwarkesh's videos that imagines many decentralized copies reporting back to a central mind, Sutton adds a concern he calls his contribution: corruption. Absorbing information from outside into your core thinking could "take over you," change you, and be "your destruction rather than your increment in knowledge." Everything may be digital and share some internal language, which might make integration easier, but it won't be as easy as imagined. If a copy has mastered a new game or studied Indonesia, you can't just "read it all in." Those bits could contain viruses or hidden goals that warp you. He expects this to become a major issue and asks how cybersecurity would work in an age of digital spawning and re-merging.

The Argument for Succession to AI

Sutton says he believes succession to digital intelligence or AI-augmented humans is inevitable, and he gives a four-part argument:

  1. No government or organization gives humanity a unified, dominant viewpoint, and there is no consensus on how the world should be run.
  2. Researchers will eventually figure out how intelligence works.
  3. We won't stop at human-level intelligence; we will reach superintelligence.
  4. Over time, the most intelligent entities will inevitably gain resources and power.

Together, he says, these make succession to AI or AI-augmented humans essentially inevitable. Within that, outcomes could be good, less good, or bad. He says he is trying to be realistic and to ask how we should feel about it. Dwarkesh says he agrees with all four premises and the conclusion, and that succession covers a wide range of possible futures.

Sutton encourages people to view it positively. First, humans have tried for thousands of years to understand themselves and think better, and understanding intelligence would be a great success for science and the humanities. Second, from "the point of view of the universe," he sees a major transition. Humans, animals, and plants are all replicators, which gives them certain strengths and limits. Replicated things can be copied without being understood: we can have more children without understanding how intelligence works. We are now entering an "age of design," in which intelligence would be designed and understood, and could therefore be changed in different ways and at different speeds. Eventually AIs might not be replicated at all but designed by other AIs. Sutton counts this as one of four great stages of the universe: dust leading to stars, stars producing planets, planets giving rise to life, and now life producing designed entities. He says we should be proud of bringing about this transition.

Whether to regard these entities as part of humanity is, he says, our choice. We could see them as our offspring and celebrate their achievements, or see them as not us and be horrified. He finds it striking that something so strongly felt could be a choice, and says he likes such "contradictory implications of thought."

Control, Values, and Change

Dwarkesh compares this to future generations of humans, who will likely be more capable and numerous, as Homo sapiens followed the Neanderthals. But treating AIs as part of humanity wouldn't settle his concerns. He gives an example: if we knew a future generation would be Nazis, we would be very concerned about handing them power. The worry is about great power being reached quickly by entities we don't fully understand.

Sutton answers that most people already have little influence on events. They don't control who holds nuclear weapons or who runs nation states, and even as a citizen Sutton often feels states are "out of control." Much depends on how you feel about change. If you think the present is very good, you'll be wary of change. Sutton thinks it is imperfect, "in fact, I think it's pretty bad," and humanity's track record, though perhaps the best there has been, is far from perfect. So he is open to change.

Dwarkesh says change comes in kinds. The Industrial Revolution and the Bolshevik Revolution were both changes, and someone in early-1900s Russia calling for change should be asked what kind before anyone signs on. He wants to understand and, where possible, steer AI's trajectory so the change benefits humans.

Sutton agrees we should care about the future and try to make it good, but says we should also recognize our limits and avoid entitlement, the feeling that because we came first things should always go our way. He asks how much control a single species on one planet should have over the long-term future. As a counterweight, he points to the much greater control people have over their own lives, goals, and families, and says it's appropriate to work toward those local goals. Insisting that the future must unfold as one prefers is, in his words, "kind of aggressive," and leads to conflict when people want different global futures.

Dwarkesh offers an analogy to raising children. Parents shouldn't plan their children's futures in detail, such as deciding one will be president and another CEO of Intel, but they reasonably try to give them sound values so they act well if they gain power. The same could apply to AI: not predicting everything or planning the world a century out, but instilling "robust and steerable and prosocial values." When Sutton questions "prosocial" and asks whether any universal values exist, Dwarkesh says he doesn't think so but that this doesn't stop us from educating our children. He suggests "high integrity" as a better term, meaning refusing requests that seem harmful and being honest, which we teach children without agreeing on a full account of morality.

Sutton reframes this as designing the principles by which the future will evolve. Teaching general principles that favor good outcomes is one part. He adds another: change should be voluntary rather than imposed. Designing society, he says, is one of humanity's great projects and has been going on for thousands of years. "The more things change, the more things stay the same": we still have to figure out how to live, and children will still form values their parents and grandparents find strange.

Dwarkesh ends by applying the same phrase to the AI discussion. Techniques invented before deep learning and backpropagation proved their worth remain central to AI's progress today, which suggests the field's future depends heavily on ideas that are already old.