Beyond RLVR: Grindability, Continual Learning, and What the Next Training Paradigm Might Look Like
Dwarkesh PatelThis video essay, narrated from a post on dwarkesh.com, examines the research bet the major AI labs are making: that training models on millions of verifiable tasks across thousands of reinforcement learning environments will produce general intelligence. Dwarkesh doesn't dismiss the bet. The essay argues that it runs into a structural limit. Most of the world's valuable skills cannot be turned into replayable training environments, so AI will eventually need to learn from scarce real-world experience and fold that learning back into its weights. The essay then sketches which techniques might make that possible and what the path could look like by 2027 or 2028.
The labs' bet and the optimists' case
The essay opens by describing the prevailing view. If AIs are trained to accomplish millions of verifiable tasks across thousands of diverse RL environments, the thinking goes, the result will be a general problem-solving agent. It would be able to make progress on open-ended tasks for weeks at a time despite errors, mistakes, and ambiguity. Proponents argue that the deficits critics point to can simply be "steamrolled" by scaling training further. Those deficits include the data inefficiency of current models and their lack of continual learning. The proponents compare this to how the fundamental research problems of natural language processing collapsed once enough compute was thrown at LLMs.
Dwarkesh then gives the optimists' rebuttals to two criticisms from the previous essay. The first concerns sample efficiency. That essay claimed models are about one-millionth as sample-efficient as humans. The optimist reply is that this is only true during training. Training is a one-time cost amortized across billions of sessions. What matters is how smart, general, and sample-efficient the model is within a session, and by that measure things have clearly been improving with more RL. Agents solve increasingly ambitious problems over longer time spans, which anyone who codes with these models has seen.
The second concerns continual learning, meaning updating a model's weights based on what it learns in deployment. Optimists argue it may not be necessary. If in-context learning becomes good enough over long horizons, there is no need to distill on-the-job learning back into the weights. Dwarkesh gives the argument its strongest form. People often say new employees aren't net productive until six months or more into a job, so on-the-job learning is clearly necessary for competence. But what if those six months could fit into the context window? Architectural innovations have already greatly expanded how much context a transformer can store. A couple more years of progress might produce context windows that feel effectively infinite.
A puzzle: why is computer use lagging?
Before weighing the bet, the essay takes what Dwarkesh calls a "completely tangential" but revealing question. Why has progress on computer use been so much slower than progress in coding or math? Computer use seems highly verifiable. Did the Etsy item arrive? Is the event venue booked? Were the taxes submitted?
Dwarkesh grants there are many reasons. One is that models see far less high-quality multimodal data in pretraining. But the essay emphasizes a reason it considers underrated. Being verifiable is not enough. A domain must also be "grindable." That means you can run many parallel rollouts against a deterministic, replayable simulator, all starting from the same point. Dwarkesh describes this as revealing "the canyon walls" that the river of AI progress will only slowly erode.
Coding shows the contrast. You can define a container holding a software repository with a missing feature, then have a thousand agents attack the problem in parallel, each with an identical copy of the container. Computer use doesn't work that way, at least not trivially. You can't send a thousand agents through the same Amazon checkout flow, because, in Dwarkesh's words, "Andy Jassy will find your bots and shut your ass down." One workaround is to build clones of Slack, Gmail, and other common applications. For now, the essay says, that is labor-intensive and doesn't scale.
Dwarkesh expects this to change once AIs are good enough at coding to build high-fidelity clones themselves. That would kill two birds with one stone, since rebuilding entire applications from scratch is also a good RL objective for coding. So computer use itself may soon be solved. The lesson from its current slowness remains, though. Unless you can build a very replayable training target for a domain, models struggle to make progress. The essay attributes this to models' extreme sample inefficiency during training.
Domains that can't be simulated
For computer use, the essay suggests, farmable deterministic simulators might offset the sample-efficiency deficit. For many other skills, that option doesn't exist. Dwarkesh asks how you would train an AI to build a business from scratch, win court cases, have a profitable day of trading, or help a candidate win an election. These rollouts require interacting with the real world and can't be recreated in a datacenter. The outer-loop verification may take months or years of real-world action. You also can't perturb the model's actions slightly across thousands of parallel rollouts to isolate what actually worked.
Dwarkesh notes that reset-free, non-stationary environments are a known open problem in RL and says they aren't pointing out anything new. The point they want to emphasize is that data in most domains is idiosyncratic and sparse, so sample efficiency is required for proficiency. If AIs are to acquire all human skills, and some skills humans lack, they must learn from information revealed in unstructured, unverifiable, and ambiguous ways, from scarce real-world interaction. In many domains the relevant training information exists in no other form. The essay puts the challenge sharply. What RL environment would produce an AI as good at politics as Lyndon Johnson, or as good at building a space-launch business as Elon Musk?
Will RLVR generalize?
The labs, as the essay frames it, are betting that reinforcement learning from verifiable rewards (RLVR) will generalize. Enough training on containerized, reproducible environments would yield a very general agent. That agent could make and execute plans, learn rapidly from new information, and pick up new skills, all within a single session. Dwarkesh illustrates what that would mean. Drop this "endlessly RLVR'd" AI into Texas politics in 1948, and it would give better advice than LBJ on winning the Senate seat. Give it $100 million in 2002 and "let it cook," and it would build SpaceX.
Whether RLVR generalizes this well, Dwarkesh says, is an empirical question. Would going from billions of dollars on RL environments to a trillion produce fully human-like general intelligence within the context window?
The essay cites a quote from Dario Amodei in a podcast conversation with Dwarkesh as a hint that RLVR generalization is not infinitely strong. Explaining why model performance degrades at long context, Dario said: "There's two things. There's the context length you train at, and there's a context length that you serve at. If you train at a small context length and then try to serve at a long context length, maybe you get these degradations." Dwarkesh acknowledges they may be reading too much into this. Their interpretation is that short-horizon RL training doesn't necessarily generalize to long-horizon performance. If it can't, the essay asks, how would agents trained on white-collar tasks generalize to being dropped into the real world and building a business as well as Sam Walton?
The wasted value of deployment
Even if AIs could become like Henry Ford or Albert Einstein after enough in-context experience, the essay argues, that would be "ephemeral and wasted" unless the learning could be put back into the weights. Dwarkesh states that around 30 to 50 percent of a lab's compute goes to inference, and that this compute currently does nothing to improve the model. They call this a huge waste.
It's worse than it sounds, the essay continues, because deployment is where the most valuable information appears. What is actually happening in the organizations where the model is used? What is it being used for? What mistakes does it tend to make in the real world? Dwarkesh offers an analogy for the current approach: a genius grad student who has never been allowed a real internship, fed more and more classroom case studies in the form of RL environments. It strikes them as bizarre that AIs already deployed throughout the economy, doing many kinds of tasks and exposed to large amounts of domain- and organization-specific tacit knowledge, can't make use of any of it.
Why the learning has to reach the weights
Dwarkesh argues that continual learning can't just mean an ever-growing KV cache as a model learns from more users. That doesn't scale, and it isn't how humans work. The brain has no clean separation between parameters and activations, and no part of the skull keeps expanding as a person learns over a lifetime. Human learning clearly involves compression, and the essay says this aids generalization and grokking.
The essay points to people with autistic-savant-type abilities who can recall random tables of numbers or nonsense syllables years later. That is roughly the fidelity models have with information in context. According to the essay, that sheer volume cripples these people's ability to understand abstractions and metaphors. Human continual learning, Dwarkesh argues, is less about having every observation at the tip of your tongue and more about "chiseling the right intuitions and big-picture knowledge back into the weights."
The difficulty is that moving learning into the weights means giving up the sample efficiency of in-context learning. Gradient updates are very sample-inefficient. As a result, every successfully shipped online-learning model has had to learn the same thing across millions of users. The essay's example is Cursor's Tab model. It online-learns by predicting the same objective, namely which edits users actually accepted, across more than 400 million requests a day. Dwarkesh says we haven't yet seen models online-learn different things for different users. A single session may contain plenty of data for a human to learn from, but not enough to train a more capable AI.
That limits current online learning to a narrow set of use cases. The point of continual learning is that the world is complicated and every job, company, and problem differs. An intelligence needs to learn the specifics of a particular deployment, which can't be stuffed into a shared training run. The essay lists what on-the-job learning covers: how everything in an organization works and fits together, how to cooperate with the surrounding infrastructure and people to advance a larger project, what the common failure modes are, and similar knowledge.
Sample efficiency and continual learning as one problem
The essay connects the two threads here. On the job, relatively little data is available to the model. Learning from it requires sample efficiency. Models can achieve that in context, using the "fast weights" attention builds on the fly, but those scale very poorly in memory. That suggests a need for architectural innovations that provide some intermediate representation.
Dwarkesh doesn't think architecture is the fundamental bottleneck. Many working ideas already exist, such as sparse attention and KV cache compaction, and new architectural papers appear every week. Perhaps, then, the bottleneck is the loss function. How do you update the weights, and thus improve the model, based on what was learned in a particular session?
On-policy self-distillation
Here too, Dwarkesh says, many ideas seem naively like they should work. The essay focuses on one that has been getting attention: on-policy self-distillation (OPSD). Dwarkesh mentions having recorded an impromptu blackboard lecture about it with Sasha Rush. The core idea is to train the base model to make the same predictions on a real-world problem that the model would make after accumulating all the context of a long session. The goal is to distill what the model learned in a session back into its weights.
The essay gives two reasons OPSD beats RLVR for this purpose. First, it needs no outer-loop verifiable reward. It needs only a model that can learn the right things within its context window, and the base model is then trained to match this "veteran teacher" that built up experience during the session. Second, it provides a much denser supervision signal. Instead of projecting a single reward across a whole trajectory, you train on the per-token probability discrepancy between teacher and student.
Dwarkesh also argues OPSD beats supervised fine-tuning for continual learning. The most naive form of SFT would train the base model to predict every token observed during the session. The essay says this makes no sense as a learning target. You don't get better at your job by recalling a perfect transcript of every day. You get better by consolidating the handful of insights that matter for improving.
RL avoids this failure mode. It concentrates updates on what is relevant to getting the outcome right, which is why its updates are so sparse. The essay calls this very important for continual learning, because learning on the job shouldn't overwrite what the base model already knows. Dwarkesh recalls a post from a few months earlier arguing that RL learns much less information per sample than supervised learning. Here they suggest that may be a feature: you change the model only as much as necessary to achieve the outcome. OPSD preserves this property. Rather than "slingshotting" toward the teacher distribution as supervised learning would, it extracts only the knowledge needed to match the teacher's results on actual tasks. Dwarkesh summarizes it as taking scarce real-world experience and squeezing all its signal into a small, well-targeted update.
Dreaming
The essay then offers what it calls a much more speculative idea, "dreaming." If an AI could build a good simulation of reality to rehearse new skills, try alternative strategies, and reinforce what works, it could experience orders of magnitude more simulated samples in the same wall-clock time.
Dwarkesh draws on history. A couple of years after DeepMind released AlphaZero, researchers trained EfficientZero, which was designed to be data-efficient. Given two hours against a simulator of an unfamiliar Atari game, the essay says, EfficientZero would probably beat a novice human. Does that make it more sample-efficient than a human? That was the training goal, but Dwarkesh says it depends on how you measure. For every real game step, EfficientZero plays dozens of simulated games in its head. Future LLMs might similarly consume far less real-world data while practicing endlessly against environments they build themselves.
The obvious difference is that simulating the whole world is much harder than emulating Go, which is why Dwarkesh calls the idea speculative. If it works, the essay suggests it would become a fourth axis of scaling alongside pretraining, RL, and inference-time compute, called test-time training or dreaming. The model would spend compute writing RL environments and training against them, rehearsing the skills a specific user needs in production. Dwarkesh contrasts this with today's /compact command in Codex, Cursor, or Claude. That command spends a small amount of compute writing a summary, which the essay says gives only "the simulacrum of continual learning." A hypothetical /dream command would instead burn huge amounts of compute to build and train against a video-game version of what the model is seeing in the real world.
A scenario for 2027–2028
The essay ends with one possible path for continual learning by 2027 or 2028. In this scenario, RLVR produces an agent that can get its bearings on an unfamiliar problem, try different strategies, and iterate when it hits a roadblock. Dwarkesh calls this the crucial thing RLVR provides: an AI competent enough to start gathering real-world experience, if it could learn from it. That agent is then sent to do real work, including on projects outside its training distribution.
Suppose effective context lengths have grown enough for an AI to co-work with someone for a full week of wall-clock time. At the end of the week, the user gives a thumbs up or down, a kind of work review. With a thumbs up, the base model distills what the AI learned during the session. It might use OPSD, dreaming, some technique not yet known, or a combination. The AI gets better at domains adjacent to what RLVR explicitly trained it on. In the next round, it improves at things adjacent to what it previously learned online. In this way, the essay argues, AI capabilities could expand well beyond the verifiable domains used in pre-deployment training.
Dwarkesh frames this as a chain. Pretraining created a base intelligence smart enough to become a competent agent with enough RLVR on top. RLVR, in turn, created an agent competent enough to be broadly deployed. Once a training recipe for continual learning arrives, that agent could learn on the job from broad deployment. At that point, the main source of AI improvement would no longer be pre-release training. It would be the experience accumulated across the economy. Every interaction would find the AI smarter, having learned both from a user's previous sessions and from its interactions with every other user. Dwarkesh calls this prospect "very scary and exciting and different from the way that AI improves right now." The essay leaves open whether the needed techniques, OPSD, dreaming, or something not yet invented, will actually arrive.
So here's a big research bet that all the labs are making. They think that if we train AIs to accomplish millions of verifiable tasks across thousands of diverse RL environments, then we will have basically built AGI, because this kind of training will have created a kind of problem-solving agent: the kind of thing that can make progress on open-ended tasks for weeks on end in the face of errors and mistakes and ambiguity.
And the people who are optimistic about this vision will say that all these things we talk about as the fundamental deficits in the current training paradigm — for example, the data inefficiency of these models, or the fact that they lack continual learning — can just be steamrolled if we scale training more, in the same way that all the fundamental research problems in natural language processing collapsed when we threw enough compute into LLMs.
So in the previous essay, I talked about how these models are one one-millionth as sample-efficient as humans, and the people who are in favor of the current training paradigm will say, "Look, that might be true, but this is only true during training." Training is this one-time cost that is amortized across billions of sessions that a model will experience. What really matters is how smart and general and sample-efficient the model is during a session, and this has clearly been improving as we've been doing more RL training. AI agents are able to solve more and more ambitious problems over longer and longer time spans. Anybody who has used these models for coding knows that.
Similarly, people would say, look, continual learning — this capability I keep harping about, where the model's weights get updated based on what it's learning from deployment — may simply not be necessary. Because if in-context learning gets so good across longer and longer time horizons, then you don't need to distill everything the model is learning on the job back into the weights.
People often say that their employees are not net productive until six months or more on the job. So clearly, online learning is necessary for competence. But what if you could just fit those six months into the context window? There have been tons of architectural innovations that dramatically increase the amount of information, or the amount of context, that a transformer can store. And why not think that, with a couple more years of progress, we might have what feels like infinitely large context windows?
Okay, so before we discuss this research bet a bit further, I want to step back and ask a completely tangential question, which I find actually very interesting and confusing about the nature of current AI progress. Why has progress on computer use been so much slower than other domains?
Computer use is so clearly verifiable. You could ask a question like: did the desired Etsy item I ordered get delivered? Is the venue for an event I'm trying to organize booked? Have my taxes been submitted? So isn't it weird that computer use has been making so much slower progress than coding and math and these other verifiable domains?
I'm sure there are many reasons for this, and one of them, of course, is the fact that the models are exposed to far less high-quality multimodal data during pretraining. But one reason that I think is actually quite underrated by people, and which I think reveals the canyon walls against which this river of AI progress will only slowly chip away, is that it is not enough for a domain to be verifiable. It also has to be very grindable, in the sense that you have to be able to run lots of parallel rollouts against a deterministic and replayable simulator, and you have to run those rollouts from the same starting point.
If you're trying to make a model better at coding, you can define some container that has a software repo with some missing feature that you have tasked the AIs with creating. And then you have a thousand parallel agents go at the problem, each of which has an identical copy of the container. But this doesn't work with computer use, at least not trivially. You can't just have a thousand agents go try the same checkout flow on Amazon to get better at using websites, because Andy Jassy will find your bots and shut your ass down.
You can solve this by making clones of Slack and Gmail and all the other common applications and websites. But at least currently, this is a very labor-intensive and unscalable way to build environments. Of course, once AIs get good enough at coding themselves to build these clones with extremely high fidelity, then I'm sure computer use will make quicker progress than it is right now. And you're also killing two birds with one stone with this kind of procedure, because getting AIs to rebuild whole applications from scratch is also a great RL objective for coding.
So while computer use itself may soon be solved, its current lethargy is telling us the following: that unless you can build a very replayable training target for a domain, the models will struggle to make much progress. And the reason this is true, of course, is that the models are incredibly sample-inefficient during training. This is the point I was making in my last video essay.
So for computer use, we might be able to make up for the sample-efficiency deficit by building these farmable deterministic simulators. But for so many other different kinds of skills that we need AIs to have, we simply can't do this.
How do we train an AI to get really good at building a business from scratch? How about winning court cases, or having a profitable day of trading in the markets, or helping a candidate win an election? The rollout here requires interacting with the real world, and you can't recreate it from just within a datacenter. The outer-loop verification here may take months or even years of real-world actions to elicit, and you can't re-observe it by perturbing the model's actions slightly in thousands of parallel rollouts to isolate exactly what the model did that actually worked.
Now, dealing with such reset-free, non-stationary environments is a known open problem in RL. I'm not pointing out anything new. But I really do want to emphasize that because of the idiosyncratic and sparse nature of data in most domains in the world, you need sample efficiency in order to get proficient. If AIs are to develop all the skills that humans have, and even skills that humans don't have, then they need to be able to learn from information revealed in unstructured, unverifiable, and ambiguous ways from scarce amounts of real-world interaction. Because in many domains, the relevant training information simply doesn't exist in any other way.
What is the RL environment to make an AI that is as good at politics as Lyndon Johnson, or as good at building a space-launch business as Elon Musk?
The labs are betting that RLVR will generalize. That is, that if you train on enough containerized, reproducible environments, you will develop a very general agent that can make and execute plans and learn rapidly from new information, and even pick up new skills, all within a single session. If you drop this endlessly RLVR'd AI into Texas politics in 1948, it could give you better advice than LBJ about winning the Senate seat. And if you gave it a hundred million dollars in 2002 and let it cook, it would build SpaceX for you.
Now, whether RLVR can generalize this well is an empirical question. If the labs went from spending billions of dollars on RL environments to a trillion dollars, would you get the kind of thing that is a fully human-like general intelligence within the context window?
Dario gave a telling quote during our podcast together, which I think hints that RLVR generalization is not infinitely strong. When he was explaining why model performance tends to degrade at long context, he said: "There's two things. There's the context length you train at, and there's a context length that you serve at. If you train at a small context length and then try to serve at a long context length, maybe you get these degradations."
Now, maybe I'm reading too much into this, but it seems like he's saying that short-horizon RL training doesn't necessarily generalize to long-horizon RL performance. And if you can't generalize from short horizon to long horizon, then how are agents supposed to generalize from getting trained at a bunch of white-collar tasks to, say, having the ability to be dropped in the real world and build a business from scratch as well as Sam Walton?
And even if, after enough in-context experience, the AIs could become like Henry Ford or Albert Einstein, all that would be ephemeral and wasted if you couldn't get those learnings back into the weights.
Around 30 to 50 percent of a lab's compute goes to inference, and that compute is currently not playing any productive role in helping improve the model. This seems like a huge waste. And it's even worse than it sounds, because it is only in deployment that the most valuable bits of information which your model could learn from are actually revealed. What's actually happening in the organizations where I'm being used? What are they using me for? And what kinds of mistakes do I tend to make in the real world?
We've got some genius grad student who's never been allowed to take a real internship, and we keep giving it more and more classroom case studies in the form of RL training on environments. It's so bizarre that we have AIs that are broadly deployed through the economy already, and are participating in so many different kinds of tasks, and are privy to so much domain- and organization-specific tacit knowledge, and they're not able to make use of it.
But this kind of continual learning requires going back to the weights. AIs can't just keep building up a bigger and bigger KV cache as they learn from more and more users. That's just not scalable, and that's also not how humans do it. There's no clean separation in our brain between parameters and activations, and it's not like some part of your skull keeps expanding as you learn more things throughout your lifetime.
When we learn stuff, there's clearly some kind of compression, and this aids our generalization and grokking. There are, in fact, some humans who have this autistic-savant-type ability to recall random tables of numbers or nonsense syllables years later — basically the kind of fidelity of information that models have in context. And such sheer volume cripples these humans' ability to understand abstractions and metaphors.
Human continual learning is less about having all your observations at the tip of your tongue and more about chiseling the right intuitions and big-picture knowledge back into the weights.
But the moment you move into the weights, you have to give up on in-context learning's sample efficiency. Because gradient updates are super sample-inefficient, all of the successfully shipped online-learning models have had to learn the exact same thing across millions of users.
For example, the Cursor Tab model online-learns by predicting the same exact objective for over 400 million requests a day. The objective here is which edits actually got accepted by the user. At least so far, we haven't seen models online-learn different kinds of things for different users, because while a single session may generate more than enough data for a human to learn from, it's not enough to train a more capable AI.
Current online learning can work for a very limited number of use cases. But the whole point of continual learning is that the world is very complicated, and each job and company and problem is different, and you need your intelligence to be able to learn the specific information related to a particular deployment, which simply can't be stuffed into some shared training run.
These are all the things we're talking about when we talk about on-the-job learning: things like how everything in your organization works and fits together, how to cooperate with all the infrastructure and the other people around you to make progress on some larger project, what the common failure modes are, and many other things like this.
As the podcast has grown, I've had to deal with more and more operational overhead. Take paying bills. In the past, contractors would just email me their invoices. Every few weeks, I'd dig through my inbox, create a folder with all the bills, and manually pay each one.
At this point, though, I just give everybody an email address that goes straight to Mercury, which is my banking platform. Whenever anybody sends an invoice to that address, Mercury automatically downloads it, scans it, and extracts all the relevant information — things like the contractor name, address, payment amount, invoice number, and due date — and then uses all of this to create a draft payment. Mercury then stores a list of these drafts for me to review. I just go through the list and double-check that they've been built correctly. I don't have to track anything or enter any information myself.
Mercury does all the fundamental things for your business extremely well, and it puts them all in one place. If you want to learn more, go to mercury.com. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N.A., Members FDIC.
In this way, sample efficiency and continual learning are actually deeply connected problems. Relatively little data is available to the model on the job. Now, to learn from this data requires sample efficiency, and models can do that in context, but using the fast weights that are built on the fly by attention, which allow for this sample efficiency, scales very poorly in terms of memory.
So we need architectural innovations that allow for some kind of intermediate representation. I talked before about how we already have many different working ideas for this kind of thing, from sparse attention to KV cache compaction. And every week, somebody releases a new paper suggesting some kind of other architectural optimization. It doesn't seem to me that architecture is fundamentally what is bottlenecking continual learning.
So perhaps the bottleneck is the loss function. How do we update the weights, AKA how do we improve the model itself, based on information that was learned from one particular session?
Even here, naively, it seems like there are many ideas that ought to work. A lot of people are talking about this technique called on-policy self-distillation recently. If you want to learn more about it, I recorded a little impromptu blackboard lecture on my iPhone with Sasha Rush a couple weeks ago, and it's in the link in the description.
But to summarize the explanation, the idea is that we encourage the base model to make the same predictions when trying to solve some real-world problem as the model with all the context accumulated after a long session would have made. The whole point of this procedure is to distill what the model learned in a session back into the weights themselves.
This is better than RLVR for two reasons. One, OPSD doesn't require us to have some outer-loop verifiable reward. We just need a model that can learn the right things within the context window. And as long as we have that, we can train the base model to match our veteran teacher model, which has built up all this experience during the session.
And two, OPSD provides a much denser supervision signal than naive RL. Instead of projecting a single reward through the whole trajectory, you can train on the per-token probability discrepancy between the teacher and student.
For continual learning, OPSD is also superior to supervised fine-tuning. The most naive version of SFT for this application that you can imagine is just to train the base model to predict all the tokens that are observed during the session.
But this makes no sense if you think about it as a learning target. The way you get better at your job is not by recalling the transcript of every single thing that happened every day with perfect fidelity. Rather, it's by consolidating the handful of insights and pieces of knowledge that are actually relevant to you getting better at your job.
RL training doesn't suffer from this failure mode. RL is great at concentrating the update to only what is relevant to getting the outcome right. That's why the updates from RL are incredibly sparse. And this is a very important property for continual learning, because as you're learning on the job, you don't want to overwrite and forget all the other things that the base model knows.
I wrote a post a few months earlier arguing that RL learns much less information per sample than supervised learning. But this may be a good thing rather than a bad thing. You only change the model as much as is absolutely necessary to achieve the outcome, and no more. OPSD preserves this property of RL, where instead of slingshotting towards the teacher distribution as supervised learning would have you do, you only extract the knowledge that is necessary to achieve the same results as the teacher on actual real-world tasks.
OPSD is one way to attack the sample-efficiency problem. You take this scarce real-world experience, and you squeeze all the signal into a tiny, well-targeted update. But there's also another much more speculative idea. Let's call it dreaming. If the AI can build a good simulation of reality against which to rehearse new skills, or try alternative strategies and reinforce what actually works, then AIs could experience orders of magnitude more simulated samples in the same wall-clock time.
Let's go back into history a bit. A couple years after DeepMind released AlphaZero, a group of researchers trained a model called EfficientZero, and the whole point of this model is to be very efficient with data. So if this model and a human both got two hours to play against a simulator of an Atari game that they hadn't seen before, this model would actually probably beat the novice human.
Does this mean that the model was more sample-efficient than the humans? Well, that was the goal of the training, but it depends on how you measure sample efficiency. Because for each step in the real game, EfficientZero is playing dozens of simulated games in its head. In a similar way, future LLMs might be able to consume far less real-world data while practicing endlessly against environments that they build for themselves.
The big difference, of course, is that it will be much harder to build a simulation of the whole world than it is to emulate the game of Go. That's why I said this is a much more speculative idea. If it works, it would become a fourth axis of scaling alongside pretraining, RL, and inference-time compute. You could call it test-time training or dreaming.
The model spends compute writing up RL environments and then training against them, and it's rehearsing all the skills that will actually be used in production for a specific user. So instead of hitting /compact in Codex or Cursor or Claude, which kindles a small amount of compute to write up a summary, and which gives you the simulacrum of continual learning, you hit /dream. And this incinerates huge amounts of compute to build and train against a video-game version of what the model is witnessing in the real world.
So what might continual learning look like by 2027 or 2028? And how do we get there? Here's one scenario. All of this RLVR training is producing an agent that can get its bearings when it's thrown at an unfamiliar problem, and it can try different strategies, and it can iterate when it hits a roadblock. This is the crucial thing that RLVR has given you: an AI that is at least competent enough to start getting some real-world experience, if it could learn from it. And once you have that, you send it out into the world to do real work, even on projects that are off the training distribution.
Now let's say at this point, the effective context lengths have expanded such that AIs can jam and co-work with you for a full week of wall-clock time. At the end of a week, you give it a thumbs up or a thumbs down, you give it a work review. And if you give it a thumbs up, the base model distills everything that the AI learned during the session, and it may use OPSD, it may use dreaming, it may use some other technique that we aren't even aware of, or it'll use a combination of all of the above.
And AI can get better at domains that are adjacent to what it was explicitly trained for beforehand with RLVR. And in the next round it gets better at the thing adjacent to what it was previously online learned. In this way, the gamut of AI skills and knowledge and capabilities can expand far beyond the verifiable domains that the model was originally trained against before it was deployed.
Just as pretraining created a base intelligence that was smart enough to become a competent agent with enough RLVR on top, so RLVR has created an agent that is competent enough to actually be broadly deployed in the world, and from this broad deployment to learn on the job once the training recipe for continual learning actually arrives.
By this point, the main way that AIs get better is not from the training they have received before they are released to the public. Rather, it's from all this experience that they'll be accumulating from being broadly deployed in the economy and engaging in so many different kinds of tasks. Every time that you interact with an AI, it'll be smarter, not only because it's been learning from your previous sessions, but also because it's been learning from all its interactions with all the other users in the world. And that's very scary and exciting and different from the way that AI improves right now.
This was a narration of a blog post that I also released on my website at dwarkesh.com. Go there if you want to read all the footnotes, or if you want to sign up so you can find out when I release the next blog post. Otherwise, I'll see you on the next episode.
Article published
