The Data Black Hole: Why AI Progress May Owe More to Data Than to Learning Efficiency
Dwarkesh PatelIn this short video essay, Dwarkesh asks what is actually driving AI progress. Their answer is that models have not become much better at learning from limited data. Instead, labs have been pouring vastly more and better data into them. The speaker defines intelligence as sample efficiency: how much data you need in a domain before you can operate fluently and competently within it. By that measure, they argue, humans are orders of magnitude ahead. The essay then considers common objections to that comparison, and asks whether the gap even matters for what the labs are trying to achieve.
Progress through a wider data distribution, not better learning
The speaker argues that it is not clear the field has made much progress in training sample efficiency over the last few years. What has changed is that the data distribution has been dramatically widened and improved. According to the speaker, AIs have mainly been getting better through more and better data, along with the compute needed to generate that data.
Reinforcement learning is the main way this has happened. The speaker frames RL as a kind of synthetic data generation. A large amount of compute is spent against a verifier, or against a rubric when an LLM acts as judge, to find out what the good data is. The model is then trained to predict these correct rollouts, in much the same way it is trained to predict the next word in internet text.
The process has a precondition. The model must already have some prior probability of producing the correct solution. The speaker says this is why "mind-stretching amounts" of human expert trajectories are needed in every field and skill the model is meant to become competent in.
How bespoke the human expert data is
To show how task-specific this data is, the speaker suggests reading job listings on the websites of data vendors Mercor and Surge. Examples include Word specialists who convert legacy documents into polished Word files, legal experts who write realistic M&A diligence reports or securities filings, and management consultants who write template market research.
The data must also exist in large quantities. By the speaker's account, each skill requires at least hundreds of human experts who generate example completions, write rubrics, and explain their chain of thought. They note that the industry producing these expert labels, and the RL environments in which "these meticulously cataloged skills can congeal," earns billions a year in revenue. They expect that figure to reach tens of billions soon.
The speaker offers a thought experiment. Imagine needing a couple of decades of courses, hundreds of concurrent professors, and millions of practice tasks just to learn how to polish a Word file. They say even this understates the gap, because the models' tasks are both far more numerous and each far harder to get through. A human student might practice a textbook problem once or twice. With GRPO, a model generates hundreds to thousands of rollouts per task, and it needs that many to solve the credit assignment problem.
The speaker concludes that the right mental model is not a human who has learned many skills. It is "a Frankenstein's monster that has been built out of a billion grafts of carefully constructed examples, all sewn together."
Why open models catch up so fast
The speaker cites an Epoch report finding that open models lag the frontier by about four months. They read this as evidence that data is the real driver of progress. Data can easily be distilled from public APIs. Hyperparameters, training tricks, and architectural optimizations cannot. If those less visible factors drove most of the progress, the speaker argues, catching up would be far harder than it has turned out to be.
This is the image behind the title. People see AIs as a galaxy glittering with capabilities, but at the center, invisible and holding everything together, is "an unimaginably massive black hole of data."
Three comparisons of the sample-efficiency gap
The speaker gives three rough comparisons.
The first is about language exposure. Assume, generously, that a person sees and hears about 2,000 words an hour. Then from birth to adulthood they encounter roughly 200 million tokens. Frontier models are trained on somewhere between tens and hundreds of trillions of tokens. By the speaker's calculation, that is close to a millionfold difference.
The second is about robotics. A person can learn to teleoperate an arbitrary humanoid or robot arm within hours. If AIs could learn that fast, the speaker says, robotics would be a deca-trillion-dollar industry, with "an endless army of Unitree G1s" doing useful work. The reason this has not happened, in their view, is that AIs learn far less efficiently. Even the millions of hours of demonstrations collected so far are not enough for complex, open-ended tasks.
The third is about driving. A teenager can learn to drive in about 20 hours of practice. Even if you count their 16 years of growing up and building physical intuition, the speaker estimates this is still three to four orders of magnitude less data than Waymo and Tesla use to train their self-driving models.
Objection 1: Evolution already did the pretraining
The speaker attributes one objection to Andrej Karpathy, who raised it on the speaker's podcast. The argument is that billions of years of evolution effectively pretrained humans, so comparing a human's lifetime data to what a randomly initialized LLM needs is unfair.
The speaker disagrees. The human genome is only about three gigabytes, and only one to two percent of it codes for proteins. They argue there is simply not enough space in it to store the parameters of a network evolution supposedly pretrained. Their proposed analogy is different: evolution found the right hyperparameters and loss functions. Within each lifetime, the brain still builds up its connectome, the analogue of the network's weights, from scratch.
The speaker adds a second point. Suppose you grant that hundreds of trillions of pretraining tokens simply catch the model up to evolution. That still does not explain why each new capability requires so much more data. An educated human does not need a hundred professors to learn a new programming language. Pretrained models, however, still need enormous amounts of data for each additional skill.
Objection 2: You're ignoring multimodal input
A second objection is that the comparison leaves out sensory data. The speaker puts the full sensory stream from birth to adulthood at probably tens to hundreds of billions of tokens.
Their response is that blind and deaf people, who are cut off from parts of that stream, still have general intelligence. That suggests the billions of sensory tokens are not what makes humans smart. The speaker goes further: deaf people who communicate through sign language and reading probably take in far fewer than the estimated 200 million language tokens. If so, the millionfold gap may itself be an understatement.
Objection 3: We just haven't scaled enough
The third objection relies on scaling laws, which show that bigger models are more sample efficient. The human brain has about 100 trillion synapses. The speaker cites current frontier models at around five trillion parameters. So perhaps making models one to two orders of magnitude bigger would bring them to human-level sample efficiency.
The speaker calls this objection "off-mark" and explains why. In the scaling-law equations, the parameter term and the data term are added to the loss independently. Start with a compute-optimally trained model and try to minimize data by adding as many parameters as needed. Using the constants from the Chinchilla paper, the speaker says that even infinitely many parameters would only reduce the data needed for the same loss by a factor of ten. Since humans are, in the speaker's estimate, thousands to millions of times more sample efficient, scaling current models cannot close the gap. The speaker takes this as a sign that humans are on a different scaling curve altogether.
Does sample efficiency even matter?
The speaker then asks whether sample efficiency is needed for the labs' two main goals: automating white-collar work and automating AI research itself.
For white-collar work, the speaker describes the labs' bet as follows. The common tasks of software engineers, analysts, and accountants are, by definition, common. That makes them relatively easy to bring into the training distribution. The speaker says the labs' revenue curves over the last few months suggest there is enormous value in doing this, even without replicating whatever makes human learning special.
In the speaker's view, it may be less efficient to train AIs on these tasks than to train humans, "but so what?" A human lifespan does not allow for the quantity and breadth of training the models receive. They illustrate this with a hypothetical human who, because of a learning disability, had to read every public GitHub repository before becoming a competent software engineer. Training that person would make no sense. They would be on Social Security early in their education, and once trained they could still work on only one project at a time. AIs, by contrast, can absorb gigawatts of training at once and amortize what they learn across billions of simultaneous sessions. So, the speaker says, labs can be "ludicrously inefficient" in training and "still be wildly in the green."
Jobs that demand out-of-distribution thinking
The speaker says the harder question is how much out-of-distribution thinking a job requires, meaning thinking that cannot be trained for in advance. This depends on the job. Some jobs, such as bank teller or travel agent, were mechanical and predictable enough to be automated long before modern AI. Others involve daily problems far from any data distribution.
The speaker counts software engineering in the second group, even though it is the job AIs are supposed to take first. They say they would bet there will be more overall demand for human software engineers in 2028 than today, largely because AI will act as a complementary input.
The labs' plan, and the question left open
For jobs in this second group, the speaker describes the labs' plan as two steps: first automate AI research, then have the automated AI researchers solve the sample-efficiency problem. That raises the central open question. Can AIs that lack human-level sample efficiency still solve the research problems standing between them and human-like intelligence and learning?
The speaker calls this very complicated and defers a full answer to a longer future post. As a preview, they say current thinking about an intelligence explosion is "very clumsy." People either dismiss the idea that AI could speed up AI progress, or they assume "some kind of God pops out the other end." What is missing, in the speaker's view, is careful reasoning about a period in which AI progress is much faster than usual, but still built on LLMs and the particular kind of intelligence LLMs have.
So one definition of intelligence is sample efficiency. That is to say, how much data do you need in a given domain to operate fluently and competently? And it's actually not clear that we've made that much progress in training sample efficiency over the last few years. It seems more like we've just dramatically widened and improved the data distribution.
The main way that AIs have been getting better is from adding more and better data, and scaling the compute required to develop that data in the first place. Obviously, RL is the main way that this has happened. You can think of RL as basically a kind of synthetic data generation, where you dump a ton of compute against a verifier — or a rubric, if you have an LLM as a judge — in order to find out what the good data is in the first place. And then you train your model to predict these correct rollouts, much in the same way that you might train that model to predict the next word in internet text.
For this process to work, the model must have at least some prior probability of anticipating the correct solution in the first place, which is why you need mind-stretching amounts of human expert trajectories in every single field and skill that you want the model to eventually be competent in.
It's hard to overstate how task-specific and bespoke this human expert data is. If you want some intuition, I recommend checking out the job descriptions on Mercor or Surge's websites. There are listings for Word specialists who will convert legacy documents into polished Word files, and legal experts who will write realistic M&A diligence reports or securities filings, and management consultants who will write up template market research.
And it is not only that the data have to be so domain-specific, but there has to be so much of it. Each skill corresponds to at least hundreds of human experts who are generating example completions, writing rubrics, and explaining their chain of thought.
There's a reason that the data industry producing these expert labels, and the RL environments in which these meticulously cataloged skills can congeal, is earning billions a year in revenue, soon to be deca-billions.
Now imagine if it took a couple decades' worth of courses with hundreds of concurrent professors and millions of practice tasks for you to learn how to polish a Word file. Even the task-count difference here understates the gap, because the models have to grind through their far more numerous tasks, each far harder.
Whereas a human student might practice a textbook problem once or twice, with GRPO, these models are generating hundreds to thousands of rollouts per task, and they need to do this to solve the credit assignment problem.
The correct way to think about these models is not like a human who has learned all these different skills that you see the models displaying. It's more like a Frankenstein's monster that has been built out of a billion grafts of carefully constructed examples, all sewn together.
Epoch recently reported that open models lag state-of-the-art frontier models by four months. I think the reason it is relatively easy for open source and previous laggards to catch up to within months of the frontier is that data is the real driver of progress. And data can be easily distilled from public APIs, whereas hyperparameters, training tricks, and architectural optimizations cannot. If the latter were driving most of the progress, then catching up would be far harder than we are observing it to be.
It is easy to forget how much data these models are trained on, and how much more it is than what we humans see in our lifetimes. We see these AIs as a galaxy glittering with capabilities. But at their center, invisible to the naked eye, holding all the constellations together, is an unimaginably massive black hole of data.
I just want to make a couple points of comparison to illustrate just how big the sample-efficiency gap is. Here's one. If a person sees and hears on average, let's say generously, 2,000 words an hour, then between the time they're born and the time they're an adult, they'll see about 200 million tokens. Now, by contrast, these frontier models are trained on somewhere between tens to hundreds of trillions of tokens. That is close to a millionfold difference.
Here's another point of comparison. If you wanted to, you could learn to teleoperate any random humanoid or robot arm within hours. And if we could get AIs to learn just as fast, robotics would be a deca-trillion-dollar industry, and you'd have an endless army of Unitree G1s doing all kinds of useful work in the world.
But the reason we can't do this is that our AIs learn much less efficiently than we do, and even with the millions of hours of demonstrations that we've collected, this is not enough to allow them to perform complex, open-ended tasks.
And a final point of comparison: a teenager can learn to drive a car with about 20 hours of practice. And even if we include their 16 years of growing up and understanding how the world works and building physical intuition, that is still three to four orders of magnitude less data than Waymo and Tesla are using to train their self-driving car models.
Now I want to deal with a couple of common responses and objections that people have to these kinds of comparisons.
One thing people will say, and I think Karpathy said this when he came on my podcast, is that for humans, many billions of years of evolution had to go into basically pretraining us. And so we're being unfair when we're comparing how little data we see within our lifetimes to what these cold-started LLMs, which are just starting off with a totally random initialization, have to learn from.
I think this is not the right way to think about it. Our genome is only three gigabytes, and only one to two percent of it is protein coding. There is simply not enough space to store the parameters of this network that evolution supposedly pretrained.
I think the closer analogy is that evolution found the right hyperparameters and the right loss functions, and that within our lifetime, we are still building up the connectome in our brain from scratch. That is to say, the thing analogous to the weights and parameters of the neural net itself.
And even if you granted this comparison and said, "Yes, the hundreds of trillions of tokens these models see to get pretrained is similar to just catching up to evolution," that still doesn't explain why any new marginal capability that you want to give these models takes so much data.
Once you have been educated, again, you don't need a hundred different professors to teach you how to learn a new programming language. But these AIs, even once they're pretrained, still require enormous amounts of data to learn the next marginal skill, and the next marginal skill after that.
Another objection to this kind of comparison is that we're not including the multimodal data that we're seeing in our lifetimes. So if we include all this sensory information that we see from birth to adulthood, that's probably tens to hundreds of billions of tokens of data. And my response to this objection is simply that blind or deaf people, who are cut off from parts of this sensory stream, still have general intelligence. That suggests to me that all these billions of sensory tokens are not really the thing that is making humans smart.
In fact, deaf people who communicate through sign language and reading, and not through hearing, are probably ingesting far less than the 200 million language tokens that we ballparked earlier, which suggests that even the millionfold difference that we calculated earlier might be an understatement.
Okay, the third common objection people make is that we just haven't scaled enough. We have these scaling laws. They tell us that bigger models are more sample efficient. The human brain, we know, is about 100 trillion synapses, and we have frontier models that are currently around five trillion parameters. So maybe we could just achieve human-level sample efficiency if we made these models one to two orders of magnitude bigger.
The reason this objection is off-mark is actually quite interesting. If you look at the way the scaling-law equations work, they tell you that the parameter and data terms are added to the loss independently. Suppose you have a model, and you've trained it compute-optimally, and you say, "I want to be sample efficient. I want to use as little data as possible, and I'll throw in as many parameters as necessary to make that happen."
Take the constants from the Chinchilla scaling-law paper. Even if you increased the number of parameters by infinity, that would only decrease by a factor of ten the amount of data that you need in order to keep the same loss.
Humans are somewhere between thousands to millions of times more sample efficient than these models. So scaling the size of current models simply can't make up for that discrepancy, and this really does suggest that humans are on a different scaling curve altogether.
Okay, all these nerdy comparisons aside, you might ask: why do we even care about sample efficiency? Is this actually necessary for the labs to achieve the two overarching objectives they have, which are, one, to automate white-collar work, and two, to automate AI research itself?
The bet that the labs are making with white-collar work is that the common tasks that a software engineer or analyst or accountant needs to do are common, and as a result, you can bring them into the training distribution quite easily.
If you look at the revenue curves of these labs over the last few months, it does suggest that there's an enormous amount of value from bringing into distribution these kinds of common tasks, even if we can't replicate whatever is making human learning so special.
And it might be more inefficient to train AIs to do these kinds of tasks than it is to train humans, but so what? Human lifespan simply does not allow for the quantity and the breadth of training that these models experience.
If you, as a human, had some weird learning disability where you needed to read through every public repository on GitHub before you could be a competent software engineer, then it would simply not make sense to train you up. You'd be on Social Security by the early stages of your education, and even once you were trained, you would only be able to work on one project at a time.
But AIs can learn these skills by firehosing gigawatts of training at a time, and what they learn can be amortized across billions of sessions at once. So we can be ludicrously inefficient in training them up and still be wildly in the green.
And then there's a question of how much out-of-distribution thinking white-collar employees need to do that you simply can't train for in advance. This is more a question about the nature of different jobs than it is a question about AI research, and it also depends on which job you're talking about.
Some jobs are so mechanical and predictable that we were able to automate them long before the modern era of AI, for example, bank tellers or travel agents. But there are other jobs that require dealing on a daily basis with problems that are quite distant from the data distribution. I think software engineering is probably one such job. This is the job that AIs are supposed to take first, but I would be willing to bet that there's overall more demand for human software engineers in 2028 than there is right now, largely due to the complementary input of AI.
The labs' plan for this latter category of jobs is first to automate AI research and then have the automated AI researchers solve the sample-efficiency problem.
So then the question is: can AIs, which do not have human-level sample efficiency, nonetheless solve the remaining research problems that stand in the way of human-like intelligence and learning?
This is a very complicated question, and I'll have to address it in a much longer future blog post. But just to tease it a bit, I think that the way people currently think about an intelligence explosion is very clumsy, because either people dismiss the possibility of AIs speeding up AI progress altogether, or they assume that some kind of God pops out the other end. They don't reason carefully about what it looks like to have a period where AI progress is much faster than usual, but to have that happen on top of LLMs and the particular kind of intelligence that LLMs are.
But I'll save that for next time. In the meantime, if you want to read this blog post, or all the other blog posts I write, or be alerted when I write a future blog post, go sign up for my newsletter at my website, dwarkesh.com. All right, I'll see you later.
Article published
