Ilya Sutskever: From the Age of Scaling Back to the Age of Research

Open on YouTube ↗
Overview

In this conversation with host Dwarkesh, Ilya Sutskever, cofounder of Safe Superintelligence (SSI), argues that the recipe that drove AI progress for the past five years is reaching its limits. He says the central unsolved problem is generalization: current models learn far less efficiently and far less robustly than people do. The discussion moves from why models look brilliant on benchmarks yet falter in practice, to what emotions might say about value functions, to SSI's strategy, how Sutskever's views on deployment have shifted, and what he thinks a good outcome for superintelligence would require. Throughout, Sutskever repeatedly says he holds opinions on key technical questions that he won't discuss publicly.

33 min read

The oddly normal feeling of a slow takeoff

The conversation opens with Sutskever remarking that it is "crazy" that all of this is real, that it feels like something out of science fiction. Dwarkesh adds that the slow takeoff feels surprisingly normal. Investing around 1% of GDP in AI might once have seemed like a huge deal, yet people have adjusted quickly.

Sutskever says part of the reason is that the investment is abstract. For most people it shows up as news that a company has announced some hard-to-comprehend dollar figure, and that is all anyone sees. Dwarkesh suggests this sense of normality might last even into the singularity. Sutskever disagrees. What feels unreal today is the size of the announcements. He expects AI to diffuse through the economy, pushed by very strong economic forces, and its effects to be felt very strongly.

Why models ace evals but stumble in the real world

Sutskever calls one fact about current models "very confusing": they seem smarter than their economic impact would suggest. They do extremely well on evaluations that look hard, yet their real-world impact lags far behind. His example comes from vibe coding. A user hits a bug and asks the model to fix it. The model apologizes effusively and introduces a second bug. When told about the second bug, it apologizes again and brings back the first one, and the user can alternate between the two indefinitely. He isn't sure how this happens, but he takes it as a sign that "something strange is going on."

He offers two possible explanations. The one he calls more whimsical is that RL training may make models "a little too single-minded and narrowly focused," less aware in some ways even as it makes them more aware in others. The second concerns data selection. In pre-training, the question of what data to use was answered by default: all of it. RL is different. Someone has to decide which environments to build, and from what he hears, every company has teams that produce new RL environments and add them to the training mix. There are enormous degrees of freedom, and he suspects people inadvertently take inspiration from evals. They want the model to look great at release, so they ask what RL training would help on those tasks. Combine that with models that genuinely generalize poorly, and he thinks it could explain much of the gap between eval performance and real-world performance. Dwarkesh summarizes this as the idea that "the real reward hacking is the human researchers who are too focused on the evals."

Dwarkesh then lays out two possible responses. One is to broaden the set of environments, so that models are trained not only on coding competitions but also on building good applications of many kinds. The other, which he suspects Sutskever is pointing toward, is to ask why becoming superhuman at competitive programming doesn't already make a model a more tasteful programmer in general, and to look for approaches that let learning in one environment improve performance elsewhere.

Two students and the "it" factor

Sutskever answers with a human analogy. Picture two students. One decides to become the best competitive programmer and puts in 10,000 hours: solving every problem, memorizing every proof technique, and learning to implement every algorithm quickly and correctly. They become one of the best. The second thinks competitive programming is cool, practices perhaps 100 hours, and also does very well. Asked which will do better in their later career, Dwarkesh picks the second, and Sutskever agrees.

In his view, models are like the first student "but even more." Developers collect every competitive programming problem that exists, augment the data to create more, and train on all of it. The result is a great competitive programmer with every technique at its fingertips. Framed this way, he says, it becomes more intuitive that such heavy preparation would not necessarily transfer to other things.

Dwarkesh asks what the second student had before their 100 hours of practice. Sutskever's answer is "it," the "it" factor. He says he knows it exists because he studied alongside such a student as an undergraduate.

What pre-training is, and why it has no clean human analog

Dwarkesh suggests pre-training might actually resemble the 10,000 hours of practice, just obtained for free because the material already sits somewhere in the data distribution. On that reading, pre-training doesn't necessarily generalize better than RL; it simply has far more data. Sutskever describes pre-training's main strengths as twofold: there is a huge amount of it, and nobody has to think hard about which data to include. It is very natural data that contains much of what people do and think, "the whole world as projected by people onto text." He adds that pre-training is very hard to reason about, because it is difficult to tell how a model relies on its data. When a model makes a mistake, it might be because something happened not to be well "supported" by the pre-training data, though he calls that term loose.

Dwarkesh raises two proposed human analogies: the first 13 to 18 years of life, when a person isn't economically productive but is building an understanding of the world, and evolution's three-billion-year search. Sutskever sees similarities to both and says pre-training tries to play both roles, but he also sees large differences. The volume of pre-training data is staggering, yet a 15-year-old who has seen only a small fraction of it knows much less, and what they do know, they know "much more deeply somehow." At that age, a person already wouldn't make the mistakes today's AIs make. On evolution, he says it might actually have an edge, which leads him to the topic of emotions.

Emotions as a kind of value function

Sutskever recalls reading about a patient whose brain damage, from a stroke or an accident, eliminated their emotional processing. The person stayed articulate, could solve small puzzles, and seemed fine on tests. But they felt no sadness, anger, or animation, and they became extremely bad at making decisions. Choosing socks could take hours, and their financial decisions were very poor. Sutskever asks what this says about the role of built-in emotions in making us "a viable agent." It may or may not be possible to get the equivalent from pre-training. What he has in mind is not emotion as such but something like a value function, which tells you what the eventual reward of a decision will be. He thinks it could come implicitly from pre-training but says that isn't "100% obvious."

He then explains value functions. In the current, naive form of reinforcement learning, a model receives a problem, takes perhaps thousands or hundreds of thousands of actions or thoughts, and produces a solution. The solution is graded, and that score becomes the training signal for every action in the trajectory. For long tasks, this means no learning happens until a solution is proposed. He says this is how o1 and R1 are "ostensibly" trained. A value function can sometimes, though not always, tell you mid-course whether you are doing well or badly. In chess, losing a piece tells you immediately that you made a mistake, without playing out the game. In math or programming, if after a thousand steps of thinking you decide a direction is unpromising, you could already assign a signal to the moment a thousand steps earlier when you chose that path, and avoid similar paths in the future.

Dwarkesh cites the DeepSeek R1 paper's point that the space of trajectories may be so wide that learning a mapping from intermediate states to value is hard, especially in coding, where people backtrack and revise. Sutskever calls this "such lack of faith in deep learning." It may be difficult, but he says it's nothing deep learning can't do. He expects value functions to be used in the future "if not already." His point about the brain-damaged patient, he clarifies, is that the human value function may be modulated by emotions in some important way hard-coded by evolution, and that this may matter for being effective in the world.

Dwarkesh notes that emotions are impressive for being so useful while being fairly simple. Sutskever agrees they are simple compared with what AI learns, perhaps simple enough to map out in a human-understandable way, which he thinks would be cool. He frames their usefulness as a complexity–robustness tradeoff: complex things can be very useful, but simple things are useful across a very broad range of situations. Our emotions evolved mostly from mammalian ancestors, were fine-tuned slightly in hominids, and include some social emotions other mammals may lack. Because they are unsophisticated, they still serve us in a world very different from the one they evolved in. They do make mistakes, though. He hesitates over whether hunger counts as an emotion, but says our intuitive sense of hunger fails to guide us well in a world of abundant food.

The age of scaling, and the return to research

Asked about scaling more broadly, Sutskever offers a periodization. Early ML consisted of people tinkering to get interesting results. Then came scaling laws and GPT-3, and everyone realized they should scale. He treats this as an example of language shaping thought: "scaling" is a single word, but a powerful one because it tells people what to do. Pre-training was one specific scaling recipe: mix compute and data into a neural net of a certain size, and you know that scaling the recipe up will give better results. Companies loved this because it offered a low-risk way to invest resources, compared with telling researchers to "go forth" and come up with something.

Pre-training will eventually run out of data, which he says is "very clearly finite." Based on things said on Twitter, he notes, Gemini may have found a way to get more out of pre-training. But after that, the options are some souped-up pre-training with a different recipe, RL, or something else. Because compute is now very large, he says, "we are back to the age of research." He puts rough dates on the eras, with error bars: 2012 to 2020 was the age of research, and 2020 to 2025 was the age of scaling. He doesn't believe that multiplying scale by another 100x would transform everything. It would make things different, "for sure," but not transform them. So it is "back to the age of research again, just with big computers."

Dwarkesh asks what the new scaling relationship might be, given that pre-training had something close to a physical law: a power law between data, compute, or parameters and loss. Sutskever notes that the field has already shifted from scaling pre-training to scaling RL. He cites Twitter reports that labs now spend more compute on RL than on pre-training, because long rollouts are expensive and each rollout yields relatively little learning. He wouldn't even call this scaling. The better question, he says, is whether this is the most productive use of compute. Once people get good at value functions, they may use resources more productively, and with an entirely new training method it becomes ambiguous whether to call it scaling or simply using resources. He expects a return to the old mode of trying many things and noticing when something interesting happens.

He stresses that value functions would make RL more efficient, but anything done with a value function can also be done without one, just more slowly. The most fundamental issue, he says, is that "these models somehow just generalize dramatically worse than people. It's super obvious."

Generalization as the crux

Dwarkesh splits generalization into two questions. One is sample efficiency: why do models need so much more data than humans? The other is teachability: why is it so much harder to convey what we want to a model? Researchers pick up a mentor's way of thinking by watching them work, with no verifiable rewards and no "schleppy, bespoke" curriculum. That second question, he says, is closer to continual learning.

Sutskever says evolution is one candidate explanation for human sample efficiency, since it may have given us a small amount of highly useful information. For vision, hearing, and locomotion, he thinks there is a strong case for this. Robots can become dexterous with huge amounts of simulated training, but getting a robot to pick up a new skill quickly in the real world, as people do, seems very far off. Ancestors going back to squirrels needed good locomotion, so humans may have "some unbelievable prior." He mentions Yann LeCun's point that children learn to drive in about 10 hours of practice, and recalls that as a five-year-old excited about cars, his own car recognition was probably already adequate for driving, despite very little data diversity from spending most of his time at his parents' house. But that might be evolution too.

Language, math, and coding are different, in his view. Models are better than the average human at these skills, but not at learning them. Because math and coding did not exist until recently, human ability, reliability, and robustness in these domains suggests that whatever makes people good learners is less a complicated prior than "some fundamental thing": people "might have just better machine learning, period."

Discussing the human learner, the two note that it uses fewer samples, is more unsupervised, and is far more robust. A teenager learning to drive has no prebuilt verifiable reward and learns through interaction with the car and the environment. Sutskever says the teenager self-corrects because they have a value function: from the start they have a sense of how well or badly they are driving. The human value function is, he says, very robust, with a few exceptions around addiction.

On how to build something like this, Sutskever says it is a question he has "a lot of opinions about." But "we live in a world where not all machine learning ideas are discussed freely, and this is one of them." He believes it can be done, and treats the existence of humans as proof. He flags one possible blocker: human neurons might do more computation than we assume, which would make things harder. Regardless, he believes it points to some machine learning principle he has opinions on but cannot discuss in detail. (Dwarkesh jokes, "Nobody listens to this podcast, Ilya.")

How much compute does research actually need?

Asked what a new research era will look like, Sutskever says scaling "sucked out all the air in the room." Everyone started doing the same thing, and "there are more companies than ideas by quite a bit." He cites the Silicon Valley saying that ideas are cheap and execution is everything, which he thinks is partly true, alongside a tweet asking why, if ideas are so cheap, nobody is having any. He also thinks that is true.

He frames research in terms of bottlenecks: ideas, and the ability to bring them to life through compute and engineering. In the 1990s, people with good ideas lacked the computers to demonstrate them convincingly, so compute was the bottleneck. Now compute is large enough that it isn't obvious you need much more to prove an idea. He gives examples. AlexNet was built on two GPUs. The transformer was built on 8 to 64 GPUs, and no single experiment in the transformer paper used more than 64 GPUs of 2017 vintage, which he estimates at roughly two of today's GPUs. He adds that the o1 reasoning work was not the most compute-heavy thing in the world. Research needs some compute, he says, but not necessarily the largest amount ever. Building the absolute best system benefits from more compute, especially when everyone works within the same paradigm and compute becomes a key differentiator.

Dwarkesh points out that the transformer became dominant only after it was validated at larger and larger scales, and asks how SSI will tell which of its ideas is the next transformer without frontier-lab compute. Sutskever argues that SSI's research compute is larger in relative terms than people assume. SSI has raised $3 billion. Other companies raise much more, but a large share goes to inference, and big loans are earmarked for it. Serving a product also requires large engineering and sales staff, and much of the research effort goes to product features. Once those are subtracted, he says, the gap in research compute shrinks considerably. And if you are doing something different, you don't need maximal scale to prove it. He says SSI has enough compute to convince itself and others that its approach is correct.

Dwarkesh cites public estimates that OpenAI spends around $5–6 billion a year on experiments alone, separate from inference, which is more per year than SSI's total funding. Sutskever says it is a question of what you do with it. Other companies have more demands on training compute, including more work streams and modalities, so their compute is fragmented. Asked how SSI will make money, he says the company is focused on research and "the answer to that question will reveal itself," with many possible answers.

Rethinking the straight shot to superintelligence

Asked whether SSI still plans to "straight shot" superintelligence, Sutskever says "maybe." He sees merit in avoiding day-to-day market competition, the "rat race" that forces difficult trade-offs, and in emerging only when ready. But he names two reasons the plan might change. One is pragmatic: timelines might turn out long. The other is that there is real value in the most powerful AI being out in the world. You can write an essay about what AI will do, he says, but seeing an AI actually do those things "is incomparable." You have to "communicate the AI," not just the idea.

Dwarkesh presses further. He can't think of an engineering discipline where the final artifact was made safe mainly by thinking about safety. Airplane crashes per mile and bugs in Linux declined because the systems were deployed, failures were noticed and fixed, and robustness improved. He also argues that the harms of superintelligence aren't only about a "malevolent paper clipper," and that gradual access could spread out the impact and help people prepare. Sutskever replies that even a straight-shot plan would involve gradual release, which he calls an inherent part of any plan. The question is what you release first.

He then offers a second example of language shaping thought, this time with two words: "AGI" and "pre-training." In his view, "AGI" exists mainly as a reaction to "narrow AI," the chess and checkers systems that could beat Kasparov and do nothing else. Pre-training reinforced the idea, because more pre-training made models better at everything more or less uniformly. But both concepts "overshot the target." A human being is not an AGI in that sense. People have a foundation of skills but lack enormous amounts of knowledge and rely on continual learning. So a successful safe superintelligence might look like "a superintelligent 15-year-old that's very eager to go," a great student who knows little and is sent off to become a programmer or a doctor. Deployment would itself involve a trial-and-error learning period, a process rather than a finished product. Dwarkesh restates this as a mind that can learn every job rather than one that already knows every job, as he says the original OpenAI charter defined AGI. Sutskever agrees.

Broad deployment and rapid economic growth

Dwarkesh sketches two paths: the learning algorithm becomes superhuman at ML research and improves itself, or, even without that, instances deployed across the economy learn different jobs and merge what they learn. The result would be a model that is functionally superintelligent with no recursive self-improvement in software, because humans cannot merge minds this way. He asks whether this implies an intelligence explosion.

Sutskever thinks rapid economic growth is likely with broad deployment. Once AIs can learn quickly and there are many of them, there will be strong forces to deploy them, unless regulation prevents it, which he says it might. How fast growth will be is hard to know. The worker is very efficient, but the world is big and parts of it move at different speeds. He expects countries with friendlier rules to grow faster, but calls this hard to predict. Dwarkesh calls the situation precarious: something as good as humans at learning but able to merge instances seems physically possible, and a new hire at SSI becomes net productive in about six months, while this system would keep getting smarter quickly.

Showing the AI: how Sutskever's thinking has changed

Sutskever says one way his thinking has changed over the past year is that he now places more weight on deploying AI incrementally and in advance. He hedges that this change may "back-propagate" into SSI's plans. The core difficulty is that future systems are hard to imagine. It is "very hard to feel the AGI," like trying to imagine being old and frail while you are young. He believes most people working on AI can't imagine it either, and that "the whole problem is the power." If something is hard to imagine, "you've got to be showing the thing."

He offers explicit predictions. As AI becomes more powerful, people will change their behavior in unprecedented ways. Fierce competitors will collaborate on safety, and he points to OpenAI and Anthropic taking "a first small step," something he says he predicted in a talk about three years ago. Governments and the public will want to act. He thinks AI companies don't currently feel AI is powerful because of its mistakes, but that at some point it will start to feel powerful, and then companies will become "much more paranoid" about safety. "We'll see if I'm right," he adds.

His third point concerns what companies should aim to build. Because there are fewer ideas than companies, everyone has converged on self-improving AI. He argues for something better: AI "robustly aligned to care about sentient life specifically." He suggests this may be easier than AI that cares only about humans, because the AI will itself be sentient. He compares it to mirror neurons and human empathy for animals, which he sees as emerging from modeling others with the same circuitry we use to model ourselves, because that is the most efficient approach.

Dwarkesh objects that most sentient beings will eventually be AIs, numbering in the trillions and then quadrillions, so this criterion may not preserve human control. Sutskever concedes it may not be the best criterion. He says it has merit and should be considered, that a short list of ideas companies could draw on would help, and that it would be "materially helpful" if the power of the most powerful superintelligence could somehow be capped. He doesn't know how to do that.

How powerful, and what "going well" might look like

On how much room there is above human intelligence, Sutskever expects multiple such AIs to be created at roughly the same time. If a cluster is "literally continent-sized," those AIs could be very powerful, and he says it would be good if truly powerful systems could be restrained or bound by some agreement. He describes the core concern as follows: a sufficiently powerful system pursuing even something sensible, like care for sentient life, in a very single-minded way might produce results we don't like. Perhaps, he suggests, the answer is not to build an RL agent in the usual sense. He notes that humans are "semi-RL agents" whose emotions make them tire of one reward and move on. Markets are short-sighted agents, evolution is intelligent in some ways and dumb in others, and governments are designed as a never-ending fight among three branches.

He also emphasizes that the systems in question don't exist yet and nobody knows how to build them. He believes current approaches "will go some distance and then peter out," continuing to improve but not being "it." Much depends on understanding reliable generalization. He recasts alignment problems in the same terms: the fragility of learning human values and the fragility of optimizing them may both be instances of unreliable generalization. What would happen if generalization were much better is, for now, unanswerable.

On what a good outcome looks like, he says that if the first N dramatic systems care for humanity or sentient life, which must actually be achieved, he can see things going well "at least for quite some time." For the long run, he offers an answer he says he doesn't like. In the short term, there might be universal high income. But "change is the only constant," and political structures have a shelf life. One equilibrium is that everyone has an AI that does their bidding, earning money and advocating politically for them, then filing a report that the person approves with "Great, keep it up." The problem is that the person is no longer a participant. His disliked solution is for people to become part-AI through something like "Neuralink++," so that understanding is transmitted wholesale and people are fully involved in whatever situation their AI is in.

The mystery of how evolution encodes social desires

Dwarkesh asks whether the persistence of ancient emotions is an example of alignment success. The brainstem issues a directive like "mate with somebody more successful," while the cortex figures out what success means in modern life. Sutskever widens the point: it is "really mysterious" how evolution encodes high-level desires. Linking a desire to a good smell is easy to imagine, since the genome can connect dopamine neurons to a chemical sensor. Social desires, such as caring about being seen positively and being in good standing, depend on high-level concepts that the brain assembles from many pieces of information, with no dedicated sensor. He feels strongly that these desires are built in, and notes that evolution seems to have hard-coded them fairly recently and easily. He knows of no good hypothesis for how.

He offers a speculation and then explains why it is probably false. Because cortical neurons mostly communicate with their neighbors, functions cluster into regions that sit in roughly the same place across people, so evolution might have hard-coded a physical location: "when that fires, that's what you should care about." Dwarkesh raises the case of people born blind, whose visual cortex is taken over by other senses. Sutskever offers what he considers a stronger counterargument: people who have half their brain removed in childhood still end up with all their brain regions, relocated to one hemisphere. So region locations aren't fixed, and the theory fails. He finds it an interesting mystery that evolution made us care about social matters so reliably, even in people with various mental conditions.

SSI's bet, the Meta episode, and timelines

Asked what SSI will do differently, Sutskever says there are ideas he finds promising, mainly around understanding generalization, and he wants to find out whether they hold up. "We are squarely an 'age of research' company." He says SSI has made "quite good progress over the past year" but needs more, and he describes the company as an attempt to be "a voice and a participant."

On his cofounder and former CEO leaving for Meta, Sutskever gives context: SSI was raising money at a $32 billion valuation when Meta offered to acquire it. He said no; his former cofounder "in some sense said yes," gained near-term liquidity, and was the only SSI employee to join Meta.

He says SSI is mainly distinguished by its technical approach. He predicts that as AI grows more powerful, companies' alignment strategies will converge. The strategy should involve companies finding ways to talk to each other and making the first real superintelligence aligned and caring for sentient life, people, democracy, or some combination. Technical approaches will probably converge eventually too. Asked when a system that learns as well as a human and then becomes superhuman might arrive, he says "like 5 to 20" years.

He expects current approaches to "stall out" in similar ways across companies. Even so, those companies could earn "stupendous revenue," though perhaps not profits, since they will struggle to differentiate. If one lab finds the right approach, he thinks its release won't make clear how it was done, but it will reveal that something different is possible, and that information will send others trying to figure it out. He adds that each increase in capability will change how things are done in ways he admits he can't yet specify.

Dwarkesh asks why the gains wouldn't accrue to whichever company first gets a continual-learning loop running. Sutskever points to history: when one company makes an advance, others scramble to match it and prices fall. He raises an idea not yet discussed, narrow superintelligences, and argues that competition favors specialization, as in markets and evolution. Different companies would occupy different niches, with one excelling at some complex economic activity and another at litigation. Dwarkesh counters that a human-like learner could learn any niche, and that a company with the first such learner could run an instance on every job. Sutskever calls this "a valid argument" but says his strong intuition is that it won't go that way: "In theory, there is no difference between theory and practice. In practice, there is."

Diversity, self-play, and competition between agents

On the "million Ilyas in a server" picture of recursive self-improvement, Sutskever expects diminishing returns, because you want people who think differently. Dwarkesh notes how strikingly similar LLMs from different companies are, and says raising the sampling temperature just produces gibberish rather than the diversity of scientists with different prejudices. Sutskever attributes the sameness to pre-training on largely the same data, with some differentiation now emerging from RL and post-training.

Self-play interested him because it promised to create models with compute alone, without data, which matters if data is the ultimate bottleneck. But self-play as practiced in the past was too narrow, useful mainly for negotiation, conflict, certain social skills, and strategizing. He thinks it has found a home in other forms: debate, prover-verifier setups, and LLM-as-a-Judge systems incentivized to find mistakes, which he describes as related adversarial setups. Self-play is a special case of competition between agents, and the natural response to competition is to differentiate. Agents that can see what others are working on would conclude they should pursue something different, and he suggests this could create an incentive for diverse approaches.

Research taste

In the final question, Dwarkesh asks how Sutskever, who co-authored work from AlexNet to GPT-3, comes up with ideas. Sutskever says people differ, but he is guided by an aesthetic of how AI should be, formed by thinking about how people are, "but thinking correctly." The artificial neuron is his example: the brain has folds and many organs, but the folds probably don't matter, while neurons seem to matter because there are so many of them. From there follow a local learning rule for changing connections, distributed representations, and the idea that because the brain learns from experience, neural nets should too. He keeps asking whether something is fundamental, looks at it from multiple angles, and seeks beauty, simplicity, and elegance, with "no room for ugliness," along with correct inspiration from the brain. The more of these that are present at once, the more confidence he places in a top-down belief.

That top-down belief, he says, is what sustains you when experiments contradict you. If you always trust the data, you may abandon a correct idea that failed only because of a bug you didn't know about. Deciding whether to keep debugging or conclude the direction is wrong comes down to the conviction that "something like this has to work, therefore we've got to keep going." That conviction rests on the multifaceted sense of beauty and inspiration from the brain he has just described.