Beyond RLVR: Grindability, Continual Learning, and What the Next Training Paradigm Might Look Like

Open on YouTube ↗
Overview

This video essay, narrated from a post on dwarkesh.com, examines the research bet the major AI labs are making: that training models on millions of verifiable tasks across thousands of reinforcement learning environments will produce general intelligence. Dwarkesh doesn't dismiss the bet. The essay argues that it runs into a structural limit. Most of the world's valuable skills cannot be turned into replayable training environments, so AI will eventually need to learn from scarce real-world experience and fold that learning back into its weights. The essay then sketches which techniques might make that possible and what the path could look like by 2027 or 2028.

18 min read

The labs' bet and the optimists' case

The essay opens by describing the prevailing view. If AIs are trained to accomplish millions of verifiable tasks across thousands of diverse RL environments, the thinking goes, the result will be a general problem-solving agent. It would be able to make progress on open-ended tasks for weeks at a time despite errors, mistakes, and ambiguity. Proponents argue that the deficits critics point to can simply be "steamrolled" by scaling training further. Those deficits include the data inefficiency of current models and their lack of continual learning. The proponents compare this to how the fundamental research problems of natural language processing collapsed once enough compute was thrown at LLMs.

Dwarkesh then gives the optimists' rebuttals to two criticisms from the previous essay. The first concerns sample efficiency. That essay claimed models are about one-millionth as sample-efficient as humans. The optimist reply is that this is only true during training. Training is a one-time cost amortized across billions of sessions. What matters is how smart, general, and sample-efficient the model is within a session, and by that measure things have clearly been improving with more RL. Agents solve increasingly ambitious problems over longer time spans, which anyone who codes with these models has seen.

The second concerns continual learning, meaning updating a model's weights based on what it learns in deployment. Optimists argue it may not be necessary. If in-context learning becomes good enough over long horizons, there is no need to distill on-the-job learning back into the weights. Dwarkesh gives the argument its strongest form. People often say new employees aren't net productive until six months or more into a job, so on-the-job learning is clearly necessary for competence. But what if those six months could fit into the context window? Architectural innovations have already greatly expanded how much context a transformer can store. A couple more years of progress might produce context windows that feel effectively infinite.

A puzzle: why is computer use lagging?

Before weighing the bet, the essay takes what Dwarkesh calls a "completely tangential" but revealing question. Why has progress on computer use been so much slower than progress in coding or math? Computer use seems highly verifiable. Did the Etsy item arrive? Is the event venue booked? Were the taxes submitted?

Dwarkesh grants there are many reasons. One is that models see far less high-quality multimodal data in pretraining. But the essay emphasizes a reason it considers underrated. Being verifiable is not enough. A domain must also be "grindable." That means you can run many parallel rollouts against a deterministic, replayable simulator, all starting from the same point. Dwarkesh describes this as revealing "the canyon walls" that the river of AI progress will only slowly erode.

Coding shows the contrast. You can define a container holding a software repository with a missing feature, then have a thousand agents attack the problem in parallel, each with an identical copy of the container. Computer use doesn't work that way, at least not trivially. You can't send a thousand agents through the same Amazon checkout flow, because, in Dwarkesh's words, "Andy Jassy will find your bots and shut your ass down." One workaround is to build clones of Slack, Gmail, and other common applications. For now, the essay says, that is labor-intensive and doesn't scale.

Dwarkesh expects this to change once AIs are good enough at coding to build high-fidelity clones themselves. That would kill two birds with one stone, since rebuilding entire applications from scratch is also a good RL objective for coding. So computer use itself may soon be solved. The lesson from its current slowness remains, though. Unless you can build a very replayable training target for a domain, models struggle to make progress. The essay attributes this to models' extreme sample inefficiency during training.

Domains that can't be simulated

For computer use, the essay suggests, farmable deterministic simulators might offset the sample-efficiency deficit. For many other skills, that option doesn't exist. Dwarkesh asks how you would train an AI to build a business from scratch, win court cases, have a profitable day of trading, or help a candidate win an election. These rollouts require interacting with the real world and can't be recreated in a datacenter. The outer-loop verification may take months or years of real-world action. You also can't perturb the model's actions slightly across thousands of parallel rollouts to isolate what actually worked.

Dwarkesh notes that reset-free, non-stationary environments are a known open problem in RL and says they aren't pointing out anything new. The point they want to emphasize is that data in most domains is idiosyncratic and sparse, so sample efficiency is required for proficiency. If AIs are to acquire all human skills, and some skills humans lack, they must learn from information revealed in unstructured, unverifiable, and ambiguous ways, from scarce real-world interaction. In many domains the relevant training information exists in no other form. The essay puts the challenge sharply. What RL environment would produce an AI as good at politics as Lyndon Johnson, or as good at building a space-launch business as Elon Musk?

Will RLVR generalize?

The labs, as the essay frames it, are betting that reinforcement learning from verifiable rewards (RLVR) will generalize. Enough training on containerized, reproducible environments would yield a very general agent. That agent could make and execute plans, learn rapidly from new information, and pick up new skills, all within a single session. Dwarkesh illustrates what that would mean. Drop this "endlessly RLVR'd" AI into Texas politics in 1948, and it would give better advice than LBJ on winning the Senate seat. Give it $100 million in 2002 and "let it cook," and it would build SpaceX.

Whether RLVR generalizes this well, Dwarkesh says, is an empirical question. Would going from billions of dollars on RL environments to a trillion produce fully human-like general intelligence within the context window?

The essay cites a quote from Dario Amodei in a podcast conversation with Dwarkesh as a hint that RLVR generalization is not infinitely strong. Explaining why model performance degrades at long context, Dario said: "There's two things. There's the context length you train at, and there's a context length that you serve at. If you train at a small context length and then try to serve at a long context length, maybe you get these degradations." Dwarkesh acknowledges they may be reading too much into this. Their interpretation is that short-horizon RL training doesn't necessarily generalize to long-horizon performance. If it can't, the essay asks, how would agents trained on white-collar tasks generalize to being dropped into the real world and building a business as well as Sam Walton?

The wasted value of deployment

Even if AIs could become like Henry Ford or Albert Einstein after enough in-context experience, the essay argues, that would be "ephemeral and wasted" unless the learning could be put back into the weights. Dwarkesh states that around 30 to 50 percent of a lab's compute goes to inference, and that this compute currently does nothing to improve the model. They call this a huge waste.

It's worse than it sounds, the essay continues, because deployment is where the most valuable information appears. What is actually happening in the organizations where the model is used? What is it being used for? What mistakes does it tend to make in the real world? Dwarkesh offers an analogy for the current approach: a genius grad student who has never been allowed a real internship, fed more and more classroom case studies in the form of RL environments. It strikes them as bizarre that AIs already deployed throughout the economy, doing many kinds of tasks and exposed to large amounts of domain- and organization-specific tacit knowledge, can't make use of any of it.

Why the learning has to reach the weights

Dwarkesh argues that continual learning can't just mean an ever-growing KV cache as a model learns from more users. That doesn't scale, and it isn't how humans work. The brain has no clean separation between parameters and activations, and no part of the skull keeps expanding as a person learns over a lifetime. Human learning clearly involves compression, and the essay says this aids generalization and grokking.

The essay points to people with autistic-savant-type abilities who can recall random tables of numbers or nonsense syllables years later. That is roughly the fidelity models have with information in context. According to the essay, that sheer volume cripples these people's ability to understand abstractions and metaphors. Human continual learning, Dwarkesh argues, is less about having every observation at the tip of your tongue and more about "chiseling the right intuitions and big-picture knowledge back into the weights."

The difficulty is that moving learning into the weights means giving up the sample efficiency of in-context learning. Gradient updates are very sample-inefficient. As a result, every successfully shipped online-learning model has had to learn the same thing across millions of users. The essay's example is Cursor's Tab model. It online-learns by predicting the same objective, namely which edits users actually accepted, across more than 400 million requests a day. Dwarkesh says we haven't yet seen models online-learn different things for different users. A single session may contain plenty of data for a human to learn from, but not enough to train a more capable AI.

That limits current online learning to a narrow set of use cases. The point of continual learning is that the world is complicated and every job, company, and problem differs. An intelligence needs to learn the specifics of a particular deployment, which can't be stuffed into a shared training run. The essay lists what on-the-job learning covers: how everything in an organization works and fits together, how to cooperate with the surrounding infrastructure and people to advance a larger project, what the common failure modes are, and similar knowledge.

Sample efficiency and continual learning as one problem

The essay connects the two threads here. On the job, relatively little data is available to the model. Learning from it requires sample efficiency. Models can achieve that in context, using the "fast weights" attention builds on the fly, but those scale very poorly in memory. That suggests a need for architectural innovations that provide some intermediate representation.

Dwarkesh doesn't think architecture is the fundamental bottleneck. Many working ideas already exist, such as sparse attention and KV cache compaction, and new architectural papers appear every week. Perhaps, then, the bottleneck is the loss function. How do you update the weights, and thus improve the model, based on what was learned in a particular session?

On-policy self-distillation

Here too, Dwarkesh says, many ideas seem naively like they should work. The essay focuses on one that has been getting attention: on-policy self-distillation (OPSD). Dwarkesh mentions having recorded an impromptu blackboard lecture about it with Sasha Rush. The core idea is to train the base model to make the same predictions on a real-world problem that the model would make after accumulating all the context of a long session. The goal is to distill what the model learned in a session back into its weights.

The essay gives two reasons OPSD beats RLVR for this purpose. First, it needs no outer-loop verifiable reward. It needs only a model that can learn the right things within its context window, and the base model is then trained to match this "veteran teacher" that built up experience during the session. Second, it provides a much denser supervision signal. Instead of projecting a single reward across a whole trajectory, you train on the per-token probability discrepancy between teacher and student.

Dwarkesh also argues OPSD beats supervised fine-tuning for continual learning. The most naive form of SFT would train the base model to predict every token observed during the session. The essay says this makes no sense as a learning target. You don't get better at your job by recalling a perfect transcript of every day. You get better by consolidating the handful of insights that matter for improving.

RL avoids this failure mode. It concentrates updates on what is relevant to getting the outcome right, which is why its updates are so sparse. The essay calls this very important for continual learning, because learning on the job shouldn't overwrite what the base model already knows. Dwarkesh recalls a post from a few months earlier arguing that RL learns much less information per sample than supervised learning. Here they suggest that may be a feature: you change the model only as much as necessary to achieve the outcome. OPSD preserves this property. Rather than "slingshotting" toward the teacher distribution as supervised learning would, it extracts only the knowledge needed to match the teacher's results on actual tasks. Dwarkesh summarizes it as taking scarce real-world experience and squeezing all its signal into a small, well-targeted update.

Dreaming

The essay then offers what it calls a much more speculative idea, "dreaming." If an AI could build a good simulation of reality to rehearse new skills, try alternative strategies, and reinforce what works, it could experience orders of magnitude more simulated samples in the same wall-clock time.

Dwarkesh draws on history. A couple of years after DeepMind released AlphaZero, researchers trained EfficientZero, which was designed to be data-efficient. Given two hours against a simulator of an unfamiliar Atari game, the essay says, EfficientZero would probably beat a novice human. Does that make it more sample-efficient than a human? That was the training goal, but Dwarkesh says it depends on how you measure. For every real game step, EfficientZero plays dozens of simulated games in its head. Future LLMs might similarly consume far less real-world data while practicing endlessly against environments they build themselves.

The obvious difference is that simulating the whole world is much harder than emulating Go, which is why Dwarkesh calls the idea speculative. If it works, the essay suggests it would become a fourth axis of scaling alongside pretraining, RL, and inference-time compute, called test-time training or dreaming. The model would spend compute writing RL environments and training against them, rehearsing the skills a specific user needs in production. Dwarkesh contrasts this with today's /compact command in Codex, Cursor, or Claude. That command spends a small amount of compute writing a summary, which the essay says gives only "the simulacrum of continual learning." A hypothetical /dream command would instead burn huge amounts of compute to build and train against a video-game version of what the model is seeing in the real world.

A scenario for 2027–2028

The essay ends with one possible path for continual learning by 2027 or 2028. In this scenario, RLVR produces an agent that can get its bearings on an unfamiliar problem, try different strategies, and iterate when it hits a roadblock. Dwarkesh calls this the crucial thing RLVR provides: an AI competent enough to start gathering real-world experience, if it could learn from it. That agent is then sent to do real work, including on projects outside its training distribution.

Suppose effective context lengths have grown enough for an AI to co-work with someone for a full week of wall-clock time. At the end of the week, the user gives a thumbs up or down, a kind of work review. With a thumbs up, the base model distills what the AI learned during the session. It might use OPSD, dreaming, some technique not yet known, or a combination. The AI gets better at domains adjacent to what RLVR explicitly trained it on. In the next round, it improves at things adjacent to what it previously learned online. In this way, the essay argues, AI capabilities could expand well beyond the verifiable domains used in pre-deployment training.

Dwarkesh frames this as a chain. Pretraining created a base intelligence smart enough to become a competent agent with enough RLVR on top. RLVR, in turn, created an agent competent enough to be broadly deployed. Once a training recipe for continual learning arrives, that agent could learn on the job from broad deployment. At that point, the main source of AI improvement would no longer be pre-release training. It would be the experience accumulated across the economy. Every interaction would find the AI smarter, having learned both from a user's previous sessions and from its interactions with every other user. Dwarkesh calls this prospect "very scary and exciting and different from the way that AI improves right now." The essay leaves open whether the needed techniques, OPSD, dreaming, or something not yet invented, will actually arrive.