The Data Black Hole: Why AI Progress May Owe More to Data Than to Learning Efficiency

Open on YouTube ↗
Overview

In this short video essay, Dwarkesh asks what is actually driving AI progress. Their answer is that models have not become much better at learning from limited data. Instead, labs have been pouring vastly more and better data into them. The speaker defines intelligence as sample efficiency: how much data you need in a domain before you can operate fluently and competently within it. By that measure, they argue, humans are orders of magnitude ahead. The essay then considers common objections to that comparison, and asks whether the gap even matters for what the labs are trying to achieve.

11 min read

Progress through a wider data distribution, not better learning

The speaker argues that it is not clear the field has made much progress in training sample efficiency over the last few years. What has changed is that the data distribution has been dramatically widened and improved. According to the speaker, AIs have mainly been getting better through more and better data, along with the compute needed to generate that data.

Reinforcement learning is the main way this has happened. The speaker frames RL as a kind of synthetic data generation. A large amount of compute is spent against a verifier, or against a rubric when an LLM acts as judge, to find out what the good data is. The model is then trained to predict these correct rollouts, in much the same way it is trained to predict the next word in internet text.

The process has a precondition. The model must already have some prior probability of producing the correct solution. The speaker says this is why "mind-stretching amounts" of human expert trajectories are needed in every field and skill the model is meant to become competent in.

How bespoke the human expert data is

To show how task-specific this data is, the speaker suggests reading job listings on the websites of data vendors Mercor and Surge. Examples include Word specialists who convert legacy documents into polished Word files, legal experts who write realistic M&A diligence reports or securities filings, and management consultants who write template market research.

The data must also exist in large quantities. By the speaker's account, each skill requires at least hundreds of human experts who generate example completions, write rubrics, and explain their chain of thought. They note that the industry producing these expert labels, and the RL environments in which "these meticulously cataloged skills can congeal," earns billions a year in revenue. They expect that figure to reach tens of billions soon.

The speaker offers a thought experiment. Imagine needing a couple of decades of courses, hundreds of concurrent professors, and millions of practice tasks just to learn how to polish a Word file. They say even this understates the gap, because the models' tasks are both far more numerous and each far harder to get through. A human student might practice a textbook problem once or twice. With GRPO, a model generates hundreds to thousands of rollouts per task, and it needs that many to solve the credit assignment problem.

The speaker concludes that the right mental model is not a human who has learned many skills. It is "a Frankenstein's monster that has been built out of a billion grafts of carefully constructed examples, all sewn together."

Why open models catch up so fast

The speaker cites an Epoch report finding that open models lag the frontier by about four months. They read this as evidence that data is the real driver of progress. Data can easily be distilled from public APIs. Hyperparameters, training tricks, and architectural optimizations cannot. If those less visible factors drove most of the progress, the speaker argues, catching up would be far harder than it has turned out to be.

This is the image behind the title. People see AIs as a galaxy glittering with capabilities, but at the center, invisible and holding everything together, is "an unimaginably massive black hole of data."

Three comparisons of the sample-efficiency gap

The speaker gives three rough comparisons.

The first is about language exposure. Assume, generously, that a person sees and hears about 2,000 words an hour. Then from birth to adulthood they encounter roughly 200 million tokens. Frontier models are trained on somewhere between tens and hundreds of trillions of tokens. By the speaker's calculation, that is close to a millionfold difference.

The second is about robotics. A person can learn to teleoperate an arbitrary humanoid or robot arm within hours. If AIs could learn that fast, the speaker says, robotics would be a deca-trillion-dollar industry, with "an endless army of Unitree G1s" doing useful work. The reason this has not happened, in their view, is that AIs learn far less efficiently. Even the millions of hours of demonstrations collected so far are not enough for complex, open-ended tasks.

The third is about driving. A teenager can learn to drive in about 20 hours of practice. Even if you count their 16 years of growing up and building physical intuition, the speaker estimates this is still three to four orders of magnitude less data than Waymo and Tesla use to train their self-driving models.

Objection 1: Evolution already did the pretraining

The speaker attributes one objection to Andrej Karpathy, who raised it on the speaker's podcast. The argument is that billions of years of evolution effectively pretrained humans, so comparing a human's lifetime data to what a randomly initialized LLM needs is unfair.

The speaker disagrees. The human genome is only about three gigabytes, and only one to two percent of it codes for proteins. They argue there is simply not enough space in it to store the parameters of a network evolution supposedly pretrained. Their proposed analogy is different: evolution found the right hyperparameters and loss functions. Within each lifetime, the brain still builds up its connectome, the analogue of the network's weights, from scratch.

The speaker adds a second point. Suppose you grant that hundreds of trillions of pretraining tokens simply catch the model up to evolution. That still does not explain why each new capability requires so much more data. An educated human does not need a hundred professors to learn a new programming language. Pretrained models, however, still need enormous amounts of data for each additional skill.

Objection 2: You're ignoring multimodal input

A second objection is that the comparison leaves out sensory data. The speaker puts the full sensory stream from birth to adulthood at probably tens to hundreds of billions of tokens.

Their response is that blind and deaf people, who are cut off from parts of that stream, still have general intelligence. That suggests the billions of sensory tokens are not what makes humans smart. The speaker goes further: deaf people who communicate through sign language and reading probably take in far fewer than the estimated 200 million language tokens. If so, the millionfold gap may itself be an understatement.

Objection 3: We just haven't scaled enough

The third objection relies on scaling laws, which show that bigger models are more sample efficient. The human brain has about 100 trillion synapses. The speaker cites current frontier models at around five trillion parameters. So perhaps making models one to two orders of magnitude bigger would bring them to human-level sample efficiency.

The speaker calls this objection "off-mark" and explains why. In the scaling-law equations, the parameter term and the data term are added to the loss independently. Start with a compute-optimally trained model and try to minimize data by adding as many parameters as needed. Using the constants from the Chinchilla paper, the speaker says that even infinitely many parameters would only reduce the data needed for the same loss by a factor of ten. Since humans are, in the speaker's estimate, thousands to millions of times more sample efficient, scaling current models cannot close the gap. The speaker takes this as a sign that humans are on a different scaling curve altogether.

Does sample efficiency even matter?

The speaker then asks whether sample efficiency is needed for the labs' two main goals: automating white-collar work and automating AI research itself.

For white-collar work, the speaker describes the labs' bet as follows. The common tasks of software engineers, analysts, and accountants are, by definition, common. That makes them relatively easy to bring into the training distribution. The speaker says the labs' revenue curves over the last few months suggest there is enormous value in doing this, even without replicating whatever makes human learning special.

In the speaker's view, it may be less efficient to train AIs on these tasks than to train humans, "but so what?" A human lifespan does not allow for the quantity and breadth of training the models receive. They illustrate this with a hypothetical human who, because of a learning disability, had to read every public GitHub repository before becoming a competent software engineer. Training that person would make no sense. They would be on Social Security early in their education, and once trained they could still work on only one project at a time. AIs, by contrast, can absorb gigawatts of training at once and amortize what they learn across billions of simultaneous sessions. So, the speaker says, labs can be "ludicrously inefficient" in training and "still be wildly in the green."

Jobs that demand out-of-distribution thinking

The speaker says the harder question is how much out-of-distribution thinking a job requires, meaning thinking that cannot be trained for in advance. This depends on the job. Some jobs, such as bank teller or travel agent, were mechanical and predictable enough to be automated long before modern AI. Others involve daily problems far from any data distribution.

The speaker counts software engineering in the second group, even though it is the job AIs are supposed to take first. They say they would bet there will be more overall demand for human software engineers in 2028 than today, largely because AI will act as a complementary input.

The labs' plan, and the question left open

For jobs in this second group, the speaker describes the labs' plan as two steps: first automate AI research, then have the automated AI researchers solve the sample-efficiency problem. That raises the central open question. Can AIs that lack human-level sample efficiency still solve the research problems standing between them and human-like intelligence and learning?

The speaker calls this very complicated and defers a full answer to a longer future post. As a preview, they say current thinking about an intelligence explosion is "very clumsy." People either dismiss the idea that AI could speed up AI progress, or they assume "some kind of God pops out the other end." What is missing, in the speaker's view, is careful reasoning about a period in which AI progress is much faster than usual, but still built on LLMs and the particular kind of intelligence LLMs have.