Why Continual Learning Keeps AGI Further Off Than It Looks: Dwarkesh's Timelines as of July 2025

Open on YouTube ↗
Overview

On the Dwarkesh Podcast, guests have given very different estimates of how far away AGI is. Some say 20 years, others say two. In this solo episode, adapted from a blog post, Dwarkesh sets out their own view as of July 2025. They argue that today's large language models are impressive but lack a basic capability, the ability to learn on the job. They think this makes near-term economic transformation less likely than some researchers expect. They also think that once the problem is solved, the change could be sudden and enormous.

16 min read

Why today's models aren't already transforming the economy

Dwarkesh starts by rejecting a common claim: that even if all AI progress stopped today, current systems would still change the economy more than the internet did. They call current LLMs "magical." But they do not think the Fortune 500 is slow to reorganize around them because management is too conservative. Their view is that getting ordinary, humanlike work out of these models is genuinely hard, because the models lack some fundamental capabilities.

This view comes partly from direct experience. Dwarkesh calls themselves "AI forward" and estimates they have spent more than a hundred hours building small LLM tools for the podcast's post-production. They have asked models to:

  • rewrite auto-generated transcripts so they read the way a human editor would make them read,
  • pick out clips from a transcript,
  • co-write essays with them passage by passage.

They describe these as simple, self-contained, short-horizon, language-in/language-out tasks that should sit right at the center of what LLMs do well. They rate the models about 5 out of 10 on them. They say that is impressive. They also say the experience of trying to make these tools useful has lengthened their timelines.

The core problem: models don't get better over time

Dwarkesh names the lack of continual learning as "a huge huge bottleneck." On many tasks an LLM's starting level may be higher than an average person's. But there is no way to give the model high-level feedback and have it improve. You are stuck with whatever it can do out of the box. You can keep adjusting the system prompt, but in practice that produces nothing close to the learning and improvement human employees go through.

In Dwarkesh's view, humans are useful mainly because of this learning, not because of raw intelligence. People build up context, examine their own failures, and pick up small improvements and efficiencies as they repeat a task.

They illustrate this with an analogy about learning the saxophone. Normally a child blows into the instrument, hears the result, and adjusts. Now imagine a different method. A student makes one attempt. At the first mistake, the student is sent away and the teacher writes detailed notes on what went wrong. The next student reads the notes and tries to play Charlie Parker cold. When that student fails, the teacher refines the notes and calls in another. Dwarkesh says this would never work, however well the instructions are written. Yet it is essentially the only way we currently have to "teach" an LLM anything.

They acknowledge that RL fine-tuning exists. They argue it is not a deliberate, adaptive process the way human learning is. They point to their own editors, who have become extremely good. The editors did not get there because someone built custom RL environments for each subtask of their job. They noticed many small things themselves and thought hard about what resonates with the audience, what content Dwarkesh likes, and how to improve their daily workflow.

Could a smarter model train itself?

Dwarkesh considers one way out. A smarter model might build its own RL loop in a way that looks organic from the outside. The user gives high-level feedback, and the model invents verifiable practice problems to train on, or even builds a whole environment to rehearse the skills it thinks it lacks. Dwarkesh says this "just sounds really hard." They are unsure how well such techniques would generalize across different kinds of tasks and feedback.

They expect models to eventually learn on the job the way humans do. They find it hard to see that happening in the next few years, because they see no obvious way to add continual learning to the kind of models LLMs are.

Learning within a session, and losing it

Dwarkesh notes that LLMs do become fairly smart and useful within a single session. When co-writing an essay, they give the model an outline and ask for drafts one passage at a time. Up to about the fourth paragraph, the suggestions are all bad. Dwarkesh rewrites each paragraph from scratch and tells the model, in their words, "Look, your shit sucked. This is what I wrote instead." After that, the model starts giving good suggestions for the next paragraph. But this subtle grasp of their preferences and style is gone once the session ends.

One possible fix is a long, rolling context window, like the one Claude Code already has, which compacts session memory into a summary every 30 minutes. Dwarkesh expects that turning rich tacit experience into a text summary will be brittle outside software engineering. Software is heavily text-based, and the codebase itself already serves as an external memory. They return to the saxophone: imagine teaching a child to play from text alone. They add that even Claude Code often undoes a hard-won optimization the two of them built together before a /compact, because the reason for the change didn't make it into the summary.

Disagreeing with Sholto Douglas and Trenton Bricken

This reasoning is why Dwarkesh disagrees with a claim that Anthropic researchers Sholto Douglas and Trenton Bricken made on the podcast. Dwarkesh quotes Trenton: "Even if AI progress totally stalls (and you think that the models are really spiky, and they don't have general intelligence), it's so economically valuable, and sufficiently easy to collect data on all of these different white collar job tasks, such that to Sholto's point we should expect to see them automated within the next five years."

Dwarkesh's own estimate is that if AI progress stopped today, less than 25% of white-collar employment would disappear. Many tasks would be automated. For example, Claude 4 Opus can technically rewrite auto-generated transcripts for them. But because it cannot improve over time or learn their preferences, they still hire a human for the job. They expect the same pattern across other white-collar work, even with more data, unless continual learning improves. AIs will be able to do many subtasks somewhat satisfactorily. But because they cannot build up context, Dwarkesh argues, they will not be able to work as real employees inside a firm.

Bearish now, very bullish later

Dwarkesh says the same reasoning that makes them bearish on transformative AI in the next few years makes them especially bullish over the coming decades. When continual learning is solved, they expect a large jump in how valuable these models are.

They say this could happen even without a "software-only singularity," where models rapidly build ever-smarter successors. The result might instead look like a broadly deployed intelligence explosion. AIs would work throughout the economy, doing different jobs and learning as they go the way humans do. Unlike humans, the copies could combine what they learn, so in effect one AI would be learning every job in the economy. Dwarkesh suggests that an AI capable of this kind of online learning might quickly become superintelligent even with no further algorithmic progress.

They do not expect this to arrive all at once, for example through an OpenAI livestream announcing that continual learning is solved. Labs have strong incentives to release innovations quickly. So Dwarkesh expects to see a broken early version of continual learning, or "test time training, or whatever you want to call it," before anything that truly learns like a human. They expect plenty of warning before this bottleneck is fully removed.

Skepticism about reliable computer-use agents by next year

Sholto and Trenton also told Dwarkesh they expect reliable computer-use agents by the end of the following year. Computer-use agents already exist, but Dwarkesh calls them "pretty bad." The researchers had something far more capable in mind. You would tell an AI, "Go do my taxes," and it would go through your email, Amazon orders, and Slack messages. It would email everyone you need invoices from, compile your receipts, decide which items are business expenses, ask for your approval on edge cases, and then submit Form 1040 to the IRS.

Dwarkesh is skeptical. They stress that they are not an AI researcher and do not want to contradict researchers on technical details. With that caveat, they give three reasons they would bet against the forecast.

First, longer horizons mean longer rollouts. The AI may need to do two hours of agentic computer work before anyone can tell whether it succeeded. Computer use also involves processing images and video, which already costs more compute, before counting the longer rollouts. Dwarkesh thinks this should slow progress.

Second, there is no large pretraining corpus of multimodal computer-use data. They quote a post from Mechanize on automating software engineering: "For the past decade of scaling, we've been spoiled by the enormous amount of internet data that was freely available for us to use. This was enough to crack natural language processing, but not for getting models to become reliable, competent agents. Imagine trying to train GPT-4 on all the text data available in 1980—the data would have been nowhere near enough, even if you had the necessary compute." Dwarkesh names possibilities that would weaken this argument. Text-only training might already give models a strong sense of how UIs work and how their parts relate. RL fine-tuning might be sample-efficient enough that little data is needed. Models might be good enough at front-end coding to generate millions of toy UIs to practice on. But Dwarkesh says they have seen no public evidence that models have suddenly become less data-hungry, especially in domains where they have had much less practice.

Third, even ideas that look simple in hindsight take a long time to get working. The RL procedure DeepSeek described in its R1 paper looks simple at a high level. Yet it took two years from GPT-4's development and launch to the release of o1. Dwarkesh admits it would be "insanely and hilariously arrogant" to call R1 or o1 easy. Reaching them took a great deal of engineering, debugging, and ruling out alternative ideas. They say that is exactly their point. If it took that long to implement "train a model to solve verifiable math and coding problems," we are probably underestimating how hard computer use will be, since it involves a different modality and much less data.

The progress that is real

Dwarkesh then deliberately turns away from the skepticism. They say they don't want to be like the "spoiled children on Hackernews" who, handed a goose that lays golden eggs, would complain about how loud it quacks.

They point to the reasoning traces of o3 and Gemini 2.5 and say the models really are reasoning. They break down problems, think about what the user wants, respond to their own internal monologue, and correct themselves when they notice a line of thought is unproductive. Dwarkesh finds it strange that people treat this as ordinary, as if machines naturally go off, think, and return with a smart answer.

They suggest some people are too pessimistic because they haven't used the strongest models in the areas where those models are best. As an example, they describe giving Claude Code a vague spec and waiting ten minutes while it builds a working application on the first try, which they call "a wild experience." You could explain this in terms of circuits, training distributions, or RL. Dwarkesh says the most direct, concise, and accurate explanation is that it is "powered by a baby general intelligence." At that point, they say, part of you has to think: "It's actually working. We're making machines that are intelligent."

The predictions: 2028 and 2032

Dwarkesh says their probability distributions are very wide and that they take probability distributions seriously. So they think preparing for a misaligned superintelligence in 2028 still makes a lot of sense, and they consider that a fully plausible outcome. They then give the years at which they would take a 50/50 bet.

Computer use: 2028. The milestone is an AI that can do their small business's taxes end to end, as well as a competent general manager could in a week. That includes tracking down receipts on various websites, finding missing pieces, emailing people for invoices, filling out the forms, and filing with the IRS. Dwarkesh thinks computer use is currently in its "GPT-2 era." There is no pretraining corpus, and models must optimize for a much sparser reward over a much longer horizon, using action primitives they are not familiar with. On the other side, the base model is already fairly smart and may have a useful prior for computer-use tasks. There is also far more compute and many more AI researchers than before, so these factors might balance out. They see small-business tax preparation as the computer-use equivalent of what GPT-4 was for language, and it took four years to get from GPT-2 to GPT-4. They clarify that they expect impressive computer-use demos in 2026 and 2027. GPT-3 was very impressive but not very useful in practice. Their claim is only that models before 2028 won't be able to handle, end to end, a week-long, fairly involved project that requires computer use.

On-the-job learning: 2032. This milestone is an AI that can learn on the job as easily, organically, seamlessly, and quickly as a human, in any white-collar role. Their example is an AI video editor that, after six months, understands their preferences, the channel, and what works for the audience as deeply and usefully as a human editor would. They still see no obvious way to add continuous online learning to LLMs. But they note that seven years is a long time: GPT-1 had just come out seven years earlier. It does not seem implausible to them that some method will be found in that period.

Why timelines are "this decade or bust"

Dwarkesh anticipates an objection. They have argued that the lack of continual learning is a major handicap, yet they are predicting that in seven years we could see something that at minimum looks like a broadly deployed intelligence explosion. They accept this. They say they are forecasting a very strange world within a fairly short time.

Their reason is that AGI timelines are "very lognormal." They sum it up as "It's either this decade or bust," then qualify it: the more accurate claim is that the probability per year falls over time, which is "less catchy." Over the past decade, AI progress has come from scaling training compute on frontier systems by more than 4x per year. Dwarkesh argues this cannot continue past this decade, whether you look at chips, power, or the share of GDP spent on training. After 2030, progress will have to come mostly from algorithmic advances, and even there the easy gains will be used up, at least within the deep learning paradigm. As a result, the yearly probability of AGI drops sharply.

Dwarkesh draws two conclusions from this. If the longer side of their 50/50 bets turns out to be right, the world could stay relatively normal into the 2030s or even the 2040s. In all the other scenarios, even if we stay clear-eyed about AI's current limits, they say we should expect some truly crazy outcomes.

Dwarkesh ends by explaining where the essay came from. It began as a blog post after their conversation with Sholto and Trenton. During the interview they found they disagreed on timelines, and it took several weeks of reflection afterward to work out exactly where they disagreed and why their own timelines were longer.