Eight Predictions for a World Where AI Models Actually Learn on the Job

Open on YouTube ↗
Overview

This video is a narration of an essay by Dwarkesh (published at dwarkesh.com). It asks what changes once AI systems can genuinely learn from experience instead of starting fresh every session. The speaker's position is that real continual learning is needed for AIs to do whole jobs as well as humans. Its arrival, they argue, would reshape AI regulation, alignment research, the competitive dynamics between labs, and the economics of serving models. The essay offers eight predictions and admits that the most important changes are probably the ones hardest to foresee.

10 min read

Why Notes Between Sessions Aren't Enough

The speaker says they have argued elsewhere that actual continual learning is necessary. In their view, AIs that are forced to write markdown files from one session to the next will not perform whole jobs as competently as humans.

To show why, they describe a thought experiment about learning the saxophone. A student who has never played walks into a music hall, tries, fails as any beginner would, and writes notes on what went wrong. The next student waits outside, comes in, reads the notes, also fails because they have never played either, and adds more notes. The process repeats with an endless line of students, each passing written notes to the next.

The speaker doubts that any sequence of text could let a later student play the saxophone perfectly on the first try. At some point the relevant experience has to be accumulated "into your brain." They expect the same to hold for many skills we want AIs to pick up from the workplaces where they are deployed. With that premise in place, the essay turns to what would follow if continual learning really works.

1. The "Train, Then Deploy" Assumption in Regulation Breaks Down

The first prediction is about regulation. Many proposals for regulating AI, the speaker notes, assume a sequence: a model is trained and then deployed. On that assumption, running checks before deployment could confirm that the model won't help with cyberattacks or "do something crazy."

The speaker does not think this assumption will necessarily hold. Suppose a model improves every day based on the millions of work sessions it does that day. Then there is no single, stable artifact to certify before release. This is one of several reasons the speaker says they worry about locking in a safety regulatory regime now. We don't know what kind of technology we will be dealing with in a year, let alone in five or ten. Locking in rules today could mean committing to an "archaic and potentially counterproductive" approach to AI threats.

The speaker does offer an alternative. If governments want some form of safety evaluation of model providers, periodic monthly or quarterly risk inspections would make more sense than focusing on a special moment after training ends and before deployment begins. That moment, they argue, will not be a meaningfully distinct category in the future.

2. Technical Alignment Would Have to Change

Second, the speaker expects labs' approach to technical alignment to need a thorough overhaul. Much current research, they say, asks how to make a frozen set of weights behave well during deployment. They say they are not aware of much research on a different question: how to ensure that, with constant weight updates, a system never falls prey to jailbreaks or drifts into a deceptive or evil persona.

Pooling adds a further concern. If AIs consolidate what they learn across users, how do you stop users from injecting backdoors or malicious inclinations into the base model?

The speaker compares this to the human alignment problem. Humans improve in a self-directed way. Noting that they don't have kids themselves but imagine this is how it goes, the speaker describes children going out into the world and learning new things. Sometimes they go off the rails: they get "one-shotted by crazy ideologies," take the wrong drug, or become very strange. Parents hope they have instilled enough common sense and basic values that their children keep improving on their own without ending up with bizarre beliefs or misanthropic ideas. Continual learning, on this framing, would make AI alignment look more like that problem.

3. More Diversity Among AI Minds

The third prediction is that the variety of AI minds will increase. Today, the speaker counts fewer than five prominent AI minds. By this they mean the base models served at once to millions, hundreds of millions, or billions of users. They add that these models are quite similar to each other because they were trained on roughly the same data.

If AIs learn from experience, and that experience differs between companies and even between instances of the same model, the speaker thinks much more diversity could emerge. They call this a net good outcome. One risk they see for the future is a "monolithic singleton that's quite boring." A world with continual learning would, they hope, be more interesting than the current "mode collapse" across models.

4. Being Ahead in the Race Pays Off Faster

Fourth, once deployment becomes part of training, the speaker argues that the returns to leading the AI race accelerate. If you have the best model, more people use it for more complex and useful work. They then give it more feedback that it can integrate beyond the session window, which makes the model smarter still. The lead feeds on itself.

5. Pressure to Ship the Smartest Models Sooner

Fifth, if models learn mainly from deployment, the speaker expects labs to feel strong pressure to release their best models earlier. As an example, they cite a report that Anthropic had been using Mythos internally since February but shipped it publicly only in June.

Under real continual learning, the speaker argues, such a gap would not be viable. A lab could not keep four months between internal and external deployment and stay competitive. A rival that ships a worse model on release day would end up with a smarter one, thanks to real-world experience.

6. A Moat Through Switching Costs

Sixth, the speaker argues that continual learning would give leading labs a clear moat, something they currently lack. The speaker says they have wondered, like many others, how AI labs will actually make money. When they had Dario on the podcast, they asked him this question, and he compared the labs to cloud providers. Cloud providers offer many undifferentiated services yet earn high margins, which, as the speaker notes, is visible in Amazon's and Google's quarterly earnings. The reason, according to this account, is that switching from one cloud to another is slow and expensive.

Today, the speaker observes, nothing stops them from starting a software repository with Codex, continuing it in Cursor, and finishing it with Claude Code. Once a model improves as it works with you session after session, switching becomes costly. Changing AIs would be like firing an employee who has built up months of context on your organization and replacing them with a very fresh, inexperienced intern you must retrain from scratch. Given that lock-in, the speaker expects model providers to be able to demand hefty margins.

7. Carrots and Sticks for Training on User Sessions

Seventh, the speaker expects enterprises to see this dynamic coming and try to avoid lock-in. But the choice may come down to accepting lock-in or giving up a very valuable feature: a model that keeps getting better for you.

If real usage becomes the main way models improve, the speaker predicts that labs may subsidize users and enterprises that let the model train on their sessions. They say this is already happening, pointing to the deals offered to new users of coding products. They compare it to why Google gives away search. Conversely, labs might refuse access to their best models to enterprises that won't allow training on their sessions. With both carrots and sticks, the speaker argues, labs have a lot of leverage to get users to let AIs learn from experience.

The speaker acknowledges glossing over a technical distinction. Updating one user's set of weights is different from merging many weight forks back into the main model, and the latter may be harder. They nonetheless expect it will "in due time" be solved.

8. Economies of Scale Move Into Inference

The final prediction concerns scale economics. The speaker notes that AI training already has large economies of scale, because expensive training is amortized across more users. They cite as evidence that lab revenues are growing far faster than their compute.

Continual learning, the speaker argues, could also create economies of scale in inference for end users, mainly through batching. They refer to their episode with Reiner Pope, where they discussed this in detail. If per-company information requires full weight updates rather than living in low-rank adapters, batching brings large advantages. The speaker cites back-of-the-envelope math suggesting that the optimal inference batch size for a sparse model such as DeepSeek v3 is more than 2,400 concurrent sequences generated at once. Below that, compute is underutilized. For the reasoning behind this, the speaker again recommends the Reiner Pope episode on inference economics.

The upshot is that a given set of weights is served efficiently only when thousands of sequences are decoded against it simultaneously. A large company with many employees and agents doing varied work can serve its own weight fork efficiently. An individual user running a batch size of one could suffer more than two orders of magnitude worse compute efficiency. The speaker concludes that the economics of serving personalized weights strongly favor large organizations.

What Remains Unknown

The speaker closes by acknowledging that much more will have changed by the time continual learning actually works. The most important changes, they suggest, are probably the hardest to anticipate. Still, they maintain that the eight predictions above already seem clear.