How Close Is Recursive Self-Improvement? Three AI Researchers on Objectives, Distillation, Data, and the Limits of RL

Open on YouTube ↗
Overview

In this episode Dwarkesh Patel talks with three researchers at labs that are open enough to speak on the record. Beren Millidge is CTO of Zyphra, which builds open-source models. John Schulman is chief scientist at Thinking Machines; he co-founded OpenAI and led the RLHF work behind ChatGPT. Charlie O'Neill is head of model training at Baseten.

37 min read

The central question is whether today's recipe of transformers, pre-training, and scaled-up reinforcement learning leads more or less directly to AI that automates AI research and then improves itself, and if it doesn't, what stops it. The panelists mostly agree that the default path points toward rapid progress. They disagree on where the bottlenecks sit: specifying objectives, sim-to-real transfer, sample efficiency, continual learning, or the supply of useful signal in the world. They also differ, sometimes widely, on timelines.

The transcript does not always label speakers. This article names a speaker only where the conversation makes the attribution clear.

Steelmanning the no-takeoff world

Dwarkesh opens with a counterfactual. Suppose it is 2036 and the world has not been transformed by billions of superintelligences. Setting aside wars, bans, and other outside shocks, what is the most likely technical reason?

The first answer appeals to something like Moravec's paradox. People keep assuming that if AI can do some impressive thing, such as hard math or chess, the result will be transformative. The AI then does it, and the impact is real but limited. If that pattern continues and "the true spark of generalization" never arrives, AI could become excellent at anything that can be written as a benchmark or an environment while a persistent sim-to-real gap blocks everything else. This speaker thinks the scenario is unlikely, because RL already produces this kind of generalization in practice. But if meta-learning proves very hard to generalize and continual learning stays unsolved, it is the default failure mode they would expect.

A second panelist describes a repeating cycle. A new model comes out, people declare it AGI, and after about a month it "starts to feel dumb." Each release catches up in some areas, but progress stays bottlenecked wherever the model is weaker, has worse judgment, or can't check its own work. Even a model that writes far more code than a person does not make a researcher 100x more productive, so capability growth isn't explosive yet. There may simply be more of these cycles than people expect.

A third framing asks how far the transformer-plus-RL recipe sits from the "global optimum" of a learner you could put on a chip. The usual takeoff argument goes like this: once an agent is even 0.1% better than every human at AI research, running hundreds of thousands or millions of copies at increasing speed should outweigh every other bottleneck. This panelist offers Moore's law as a counterpoint. Its straight line held only because many discrete innovations kept it going. LLMs followed a similar pattern: pre-training scaling hit diminishing returns, RL opened a new curve, and the combined line kept looking straight. If the next discontinuity is far from the current paradigm, perhaps far enough to require abandoning gradient descent, then LLMs trained on RL environments, even environments targeted at self-improvement, may not find it, and progress could level off.

Dwarkesh pushes back. Given the progress since 2012, it would be strange if systems never came to dominate humans at R&D, including at inventing new paradigms. He cites an argument from Ryan Greenblatt using chess engine Elo. Engine Elo rose roughly linearly for decades, yet crossing the human range felt like a discontinuity: experts went from always winning to never winning. On this view, AI's modest economic impact so far reflects a slow climb through the human range, not a ceiling. One panelist agrees and says we are already close to crossing it. Avoiding a transformed world would require an asymptote just below that point. That panelist considers dramatic regulation a more likely cause of a "normal" 2035 than any technical barrier.

Well-specified objectives versus open-ended science

A recurring distinction is between research with a cleanly specified objective, the "autoresearch" style where you push down pre-training loss or push up an environment's reward, and open-ended science, where nobody, human or AI, can state the objective in advance. Paradigm shifts seem to need the second kind.

John adds a historical note. In the early OpenAI days he had the intuition that minimizing log loss would not produce intelligence. The important bits made up such a small fraction of the loss that noise would swamp them, so researchers would need better objectives that emphasized what mattered. Plausible arguments supported this view. For example, humans don't model everything in their environment, and most people can't reproduce a scene photorealistically. "But then it turned out that it just worked anyway." More broadly, John says the field depends on generalization that is hard to predict. The most important advances are often "types of generalization that we have no right to expect": from naive next-token prediction to tasks that require deep understanding, and from verifiable tasks to less verifiable ones.

Dwarkesh then offers an argument for a fast takeoff that needs no increase in inputs beyond AI labor. Before every seven-figure experiment, spend a comparable amount of compute on automated researchers who effectively spend "a century" designing the optimal experiment, running small ablations, and building theory. Afterward, spend another century analyzing the results. One panelist agrees that current research is nowhere near the ceiling. They picture AIs spending compute on analysis and theory-building comparable to what is spent on the experiments themselves.

Another panelist accepts this for well-specified objectives but adds a limit: thinking can only update your posterior on bits you already have. They give concrete examples. An AI reviewing the Kaplan scaling laws would have noticed that the authors used intermediate checkpoints without accounting for learning-rate annealing, and the panelist believes catching that would have saved a year or two. They also cite muP, the scaling of learning rate with model size, and the role of model width. On such problems they imagine a roughly 10x speed-up. But they don't see how that extends to choosing the right objective in the first place. Another panelist calls this the key question for rapid self-improvement. A self-propelling loop requires AIs to propose objectives, optimize them, propose new ones, and not go off the rails for a very long time. That kind of self-directed autonomy could turn out to be a Moravec's-paradox case: trivial for evolved creatures, hard for AI. When Dwarkesh notes that task horizons keep growing, the panelist concedes there is no obvious evidence for this worry. Today's agents are quite persistent, which counts against it.

The last job for humans

Asked which recent innovation looks most like the last thing humans would have to do before AI fully automates AI R&D, one panelist says "iteratively asking the right questions." Current AIs code experiments well but propose research directions poorly; their ideas tend to be tiny incremental steps. As a contrast, they cite the move from DeepMind's approach of solving intelligence by mastering games to Alec Radford's decision to predict the next token over a broad swath of data. Even after that, scaling up took time, because the field first had to develop scaling laws.

John says the most durable human role is defining the objective and deciding what we want: how assistants should behave, what "helpful" means, what the RLHF objective should be, and later how to write constitutions and model specs. When someone says alignment is "the final job," John splits alignment into two parts, specifying the objective and optimizing it, and says the first will not go away soon. Post-training teams need many people because there are so many areas where someone has to decide how the model should behave.

Why model providers haven't consolidated

Dwarkesh asks why the model market hasn't concentrated, given the many forces that favor centralization. One answer is distillation. Anything learned through RL amounts to a small number of bits, so if you can get trajectories that show a behavior, you can easily distill it. Company-specific models that learn from deployment could also change the landscape. Another panelist adds that continual learning doesn't prevent distillation: if a frontier model improves daily, others can distill it daily.

Distillation is not trivial, though. With supervised distillation the prompt distribution is extremely important. Even with full access to a model and its chain of thought, you need a wide distribution of realistic prompts to extract its useful capabilities. One panelist notes a recent report that some Chinese companies are probably using router or proxy services that let users in China reach blocked US frontier models, mostly for coding, and that these services collect and sell the data. That gives distillers close to the ideal prompt distribution.

Another panelist says models increasingly automate this. Published pipelines from Chinese labs start from seed prompts, drawn from humans and data like this, and synthesize broad coverage using existing or frontier models. Dwarkesh objects that the valuable part is the full multi-turn user trace, such as "that didn't work, let's step back," and a panelist replies that if you could generate that without users, "then you just have RSI." They also point out an irony: distilling a capability is easier than creating it. A distiller who wants "a good politician" asks a frontier model that already has the skill. The first lab to build it had to gather real data on what politicians actually do.

The puzzle of strong open models and weaker small frontier models

One panelist raises what started this line of discussion: they consider Sonnet 5 and Opus 5 "almost objectively worse" than GLM-5.3 and Kimi K3, even though Anthropic has its own hardest RL environments and logit distillation from Mythos. Their conclusion is that frontier labs may have little or no advantage from RL environments, and that real-world deployment data may matter more than environments. Someone points out a tension: the lab managed to incentivize those capabilities in its frontier model, so why can't it do so again in a smaller one? Perhaps the student-teacher gap is too large. On Opus 5 specifically, one panelist relays the view that it seems to have an internal AI judge checking everything it does, which is why it uses so many tokens, but that it lacks the "big model smell" of Fable for knowing when to stop. "The reach exceeds the grasp."

A panelist offers a different hypothesis based on two axes of environment design: difficulty and realism. Difficult environments are comparatively easy to build. They are hard, puzzle-like, easily verified tasks, which this panelist calls "the benchmaxxing distribution." Realism means multi-turn coding-agent settings with several objectives and back-and-forth with a human, and it requires rubrics or human feedback. A lab building a behavior for the first time has to push on both axes. Naive distillation matches the teacher on the difficult, verifiable distribution but misses the realistic one. Large models may also generalize better from narrow hard tasks to realistic ones. With a good realistic prompt distribution, a student can match the teacher well. Without it, the student matches the benchmarks and does worse on everything else. The panelist suggests this might partly explain Anthropic's smaller models, while stressing that outsiders can't know their post-training details. The lab might also have simply gotten a few things wrong, since "it's really easy to screw up post-training in some way that doesn't show up in benchmarks." Another panelist adds that frontier labs buy much of their data from large data vendors, and Chinese labs can buy the same data, so keeping up is relatively easy.

How the first automated AI researchers might be trained

Dwarkesh describes Ryan Greenblatt's toy picture: have a very capable model build small models that excel at inner-loop challenges, such as reaching a target loss with minimal compute or beating games that require continual learning. John expects something more pragmatic. Labs will combine human feedback, to absorb researchers' taste, with practice environments for multi-step research projects. Each iteration will patch whatever looked most broken in the previous model, and researchers using the AIs day to day will be the ones who spot those weaknesses.

Another panelist frames the choice as how far back in the "lineage" you roll before letting the model self-play. In principle you could return to a point before GRPO and have the model rediscover how to do RL. In practice, compute constraints will keep labs at the frontier, "diffing" the bugs and improvements found since the last model and turning them into environments. This panelist calls that continual learning within the lab, since it distills the last three months of research progress back into the model. They think it also explains why progress can feel asymptotic: the model keeps chasing what human researchers, with AI help, found recently.

The same panelist notes one difference. Distillation from trajectories can't exceed its source, but environments can be designed that no human can solve. Their examples are a nanochat speedrun faster than any human achieves, a target loss like 1.3 that no one can currently reach, or a 100-million-parameter model that beats a hard game. When the panelist calls beating Minecraft with such a model possibly "too easy," Dwarkesh remarks how striking that would have sounded five years ago.

John adds that much research isn't hill-climbing on a well-defined goal. It usually starts from an intuition. Researchers design a task meant to show "signs of life" for a method, then make the task progressively more realistic. That process relaxes the realism axis until the method matures. Other research aims at explanation and informal theory. Models will likely train on a mix: easily verifiable tasks, LLM-as-judge tasks, and tasks where a human is asked whether the result looks reasonable. They will probably generalize to fuzzier research to some extent. Whether they generalize enough to "become self-sealing without humans being in the loop at all is unclear."

The labs' bet: RL across many environments

Dwarkesh summarizes what he sees as the labs' plan. Scale RL with verifiable rewards across millions of environments in hundreds of domains, so that persistence, triage of information, and multi-agent coordination emerge. Combine that with sample-efficient in-context learning, and the result should function as a drop-in remote worker for a week or a month, having learned meta-skills in data-center simulations rather than from deployment.

One panelist says it's now hard to separate work aimed directly at RSI from work that makes generally useful models to generate revenue for the next training run. For the latter, yes, this is the bet. They describe Anthropic's lineage of environments as a clear example: coding first, as the lowest-hanging fruit, then finance with large amounts of Excel data, then PowerPoint, and on through the long tail of office work. They say other labs, including open-source labs, have concluded this was the right bet.

Dwarkesh recalls asking Dario Amodei why a lab expecting human-like on-the-job learning would bake in PowerPoint skills. Possible answers: models aren't there yet, so amortizing skills into training makes sense; or the real target is RSI, and deployable skills are a revenue source. John says strong in-context learners wouldn't strictly need finance training, but domain training can make them more efficient at runtime. In practice, providers do go domain by domain, and he counts that as one of the main reasons models have improved so much. Another panelist adds that doing both is cheap because large models have ample parameter capacity, that some transfer likely occurs in meta-skills like taste and long-horizon work, and that there isn't much RSI data in the world anyway.

Sim-to-real and learning from deployment

John questions whether sim-to-real, meaning studying real tasks and building simulated environments for RL, will stay dominant forever. Many things are hard to simulate, especially real-time interaction with many people. Another panelist says sim-to-real has to dominate while sample efficiency is low, since no human will sit through thousands of RL interactions. Better sample efficiency would shift weight toward learning from deployment. John adds that off-policy learning from existing traces is another route.

Dwarkesh presses on the point. Roughly half of compute goes to inference that doesn't improve the model, even though digital minds could in principle accumulate millions of years of experience across all their instances. When will that hive-mind learning begin? One panelist says it is already happening at the generation level: deployment data, filtered, annotated, or synthesized, can go into the pre-training or mid-training of the next model, and they think this explains much of the generation-over-generation improvement. They say they don't know whether the US labs do it, since those labs claim not to train on user data, but they say Chinese labs "100% do," and that this is essentially what distillation is.

Another panelist agrees that the loop exists if you "zoom out far enough," though the ideal of an individual model updating live on each experience breaks down at fine granularity. They point to faster loops at application companies built on open models. Harvey does this with legal agents. Cursor's Composer, as this panelist describes it, did online learning with what amounted to REINFORCE for a long time, applied to the generative model and not only the earlier Tab completion model. Online RL has no groups, only one user and one rollout, so variance is high. According to the panelist, Cursor answered this with heuristics estimating how much better or worse than average a response was, and it deployed a new model every five hours if it improved on CursorBench, discarding it otherwise. Humans still choose the signals. John warns that the hardest problem with natural data is knowing the reward function: a superficial signal like whether an edit was accepted can be reward-hacked.

Dwarkesh's larger worry is that long-horizon real-world tasks, such as running a business, trading profitably, winning a court case, or even year-long software projects that involve clients, are hard to simulate. If transfer from simulation isn't enough and weight updates from real interaction are needed, then the sample inefficiency of models, which he suggests may be "a millionfold" behind humans in data from birth to adulthood, could become a serious problem. He calls it the one thing that makes him doubt a crazy recursive self-improvement within ten years.

Cumulative versus non-stationary tasks, and taste

One panelist divides tasks into cumulative ones and non-stationary ones. RSI may be cumulative. In principle, a Python file of under a million tokens could train a self-improving model from scratch, and each discovery, such as attention, mixture of experts, or GRPO, is "a line in the sand" that gets added to the stack. As an example they mention an OpenAI-described training run of "5.6 Sol" or "5.6 Terra" that simply called scripts like pre-training.sh and post-training.sh without rediscovering attention. A legal associate at a law firm faces the opposite: shifting relationships, implicit processes, and constantly changing context. If labs believe RSI is cumulative, they may concentrate compute there. Dwarkesh remarks: "It's so unfortunate that RSI happened to be easier than being a paralegal."

John says sample efficiency is only one of several weaknesses. Models are very sample-efficient in context but may be weaker in a medium-length regime where humans update more efficiently. Other weaknesses are unrelated: lower diversity of thought and poor long-horizon judgment. He describes much of what people call taste as knowing what works in the long run, for instance in software, which systems will stay maintainable over the life of a project.

Dwarkesh asks whether a trillion-token context holding a researcher's whole experience, with the current level of in-context learning, would solve taste. John says the model would have to be trained to learn from such context, either to learn the right updates or to generalize, and another panelist notes you would need trillion-length training data. A third panelist argues taste may be meta-learnable from short episodes. A PhD student goes from first year to postdoc in about five years and perhaps 10 to 30 projects, and AI will have vastly more experience to learn from. How that generalizes to truly long-horizon work is, they say, "really unsolved."

Hive minds, incentives, and catastrophic forgetting

John argues that whether a hive mind emerges that learns from all deployments depends heavily on incentives, not just technology. Companies won't want a provider learning from their deployments in ways that erode their advantage. Another panelist expects economic pressure to favor swappable modules over updates to one shared model: LoRAs, compressed-context methods such as linear attention, or "cartridges," which are KV caches trained to be highly compressed. Dwarkesh counters that labs could still feed all those traces into the next pre-training run. The panelist agrees this remains valuable, but expects stages: specialized deployments generate traces, labs consolidate them into a new model three months later, and the cycle shortens to weekly, daily, then hourly, "at which point we've basically solved it."

Charlie reports on research into the micro level. At scale, with large batches and enough noise washed out, the outer loop of mid-training plus custom environments does work as a form of continual learning. But when you repeatedly update one model on small amounts of data, for example for a single law firm, the methods break down. SFT on successful traces, on-policy or off-policy, leads after hundreds of micro-updates to catastrophic forgetting and degraded general capabilities. On-policy distillation pushes the problem further out but eventually hits the same wall. RL is good at adding capabilities but poor at adding explicit knowledge, like how a specific firm handles a specific process. Asked whether the root issue is capacity or technique, Charlie says a bit of both. SFT and on-policy distillation can be "way too destructive," while RL changes the model very little, which is both its strength and its limit.

Another panelist says the issue is mainly technique, not capacity. Pre-training a same-size model from scratch with all the accumulated data yields a better model, which is largely what happens today. Naively training on non-stationary data causes forgetting and plasticity loss, and there are no good methods to prevent it. Continual mid-training from a checkpoint works for a long time but eventually plateaus, which is why labs train new bases. Charlie adds that the amount of retraining from scratch has shrunk, but the ideal of making small iterative updates to the latest model without losing anything is still out of reach. Dwarkesh asks whether post-training already does this. The answer is that post-training operates at a large enough scale, over broad enough distributions, to wash out noise. Dwarkesh suggests that billions of deployed instances might supply similar scale, and a panelist allows that it could, "maybe at that scale."

How much of progress is data?

Everyone agrees that some data distribution exists that would train current architectures into superintelligence. In the trivial case, you memorize the Python file that trains it. One panelist describes a ladder of RL environments that could reach researcher-level AI, where each rung takes exponentially more effort to build. Environment builders currently exploit asymmetries. Some environments are easier to generate backwards than to solve forwards, for instance hiding a complex data-generating process the model must infer. Others inject real-world information: a bug that Anthropic found through tens of thousands of humans and LLMs becomes a neat environment that a single model could theoretically solve in a few million tokens. This panelist expects diminishing returns as such shortcuts run out and building long-horizon tasks by hand gets expensive, with compute and time costs for agents to attempt them on top of that.

Two examples cut in different directions. Dwarkesh mentions a report that the Talkie model, trained only on pre-1930 data, was fine-tuned on modern coding-agent data and beat Claude 3 Opus on SWE-bench. He takes this to mean that copying expert behavior into weak models is surprisingly easy. A panelist cites a counterexample: a paper in which a model trained to fifth-grade math could not be RL'd to college math because the gap was too large. Climbing year by year would have worked. Current RL explores poorly; if the model fails in 128 rollouts, it gets no signal, which is why RL needs curricula and pre-training doesn't.

The deeper question is where signal comes from. In pre-training the signal is already in Common Crawl, and the job is filtering noise, which can be automated. At the capability frontier, the relevant bits aren't anywhere in existing data: "There's no hidden proof of a Millennium Prize problem sitting in Common Crawl." They must come from humans writing out reasoning, from human-designed environments, or from deployment. One panelist argues that even the whole world produces few new bits beyond current models' reach, such as newly solved hard math or coding problems, and that this is why diminishing returns set in.

The host describes an investigation with Jerry Han, a Princeton student. They trained every open training recipe from 2019 to now against every dataset from the same period, including GPT-2 on Ultra-FineWeb and the newest recipe, Delphi, on the Pile. At very small scale, data explained about a 12.0x compute-efficiency gain and architecture about 3.7x. A panelist doubts pre-training data gains can keep growing, since the useful internet is not expanding at the old rate and "a bunch of 0.1% loss drops" remain. They compare the combined figure, which they put at 33x, with an estimate attributed to Epoch of 3x per year since 2019, or more than 2,000x, and wonder where the missing factor went, perhaps post-training. Another explanation offered is that many gains depend on scale, and the experiments were very small.

Beren disputes treating architecture as a simple multiplier. Architectures unlock qualitatively new regimes. Without something like GQA, million-token context would be prohibitively expensive, so million-token data could never be used, and a test at 2K context will overstate the importance of data. He also thinks modern mid-training and post-training data gets more valuable with scale, since a 100-million-parameter model can't use SWE-bench traces. Another panelist notes that labs like Kimi and DeepSeek now design architectures, such as compressed attention, for real-world inference efficiency and not only for lower loss.

Parameters, sparsity, and hardware

Dwarkesh notes that frontier open-model parameter counts have roughly doubled each year and asks what active parameter counts will look like in 2030. Charlie expects a plateau for the next few years. Long RL rollouts make inference efficiency critical, and environments, not model capacity, are the bottleneck. He suspects Mythos and the GPT models are much smaller than the 10-trillion-parameter figures people discuss, based on comparisons with open models. Model size should be matched to the hardest available environments: large enough for a decent pass@1, no larger. So growth depends on how fast Mercor and in-house teams can raise environment complexity. He also disputes that sizes have doubled annually. Trillion-parameter models such as Falcon have existed for years, and he cites a post by Liam from Periodic Labs about Google's very sparse Switch Transformer, which he says was strong on knowledge but poor at reasoning. By his account, the field has been working in the 100-billion to 2-trillion range for a while.

John expects models to keep growing as compute and GPUs scale, with the details depending on scaling laws in non-obvious ways. As high-quality pre-training data runs low, data efficiency may matter more than compute efficiency in choosing architectures, and that could affect sparsity. He says sparsity is poorly understood. It has increased but may have a sweet spot, and there is a debatable argument that it hurts data efficiency because experts may relearn the same things. If data rather than compute goes on the x-axis, a different set of models lands on the optimal frontier.

Beren adds that serving multi-trillion-parameter models needs high memory bandwidth and VRAM, and that newer hardware generations should make larger RL inference practical. Larger models are more sample-efficient per data point. Today compute is the binding constraint, so models are over-trained relative to Chinchilla, but data scarcity could push back toward Chinchilla-optimal or even under-trained regimes. Charlie counters that compute will stay the bottleneck for years. Someone notes that under Chinchilla, pushing parameters toward infinity reduces required data by less than 10x. John adds that clean scaling laws hide a lot of complexity in hyperparameter scaling and optimizer parameterization. Beren quips that "bugs have their own clean scaling laws as well," citing Kaplan's annealing oversight and the omission of embedding parameters, which distorted estimates for small models.

Why RL works better than the "one bit per episode" argument suggested

Dwarkesh recalls John's paper noting that RL gives about one bit per episode, and his own blog posts arguing that low pass rates make it even worse. Yet models seem much smarter, apparently because of scaled RL. Beren gives several reasons. First, mid-training is underrated. Pre-training on synthetic reasoning data and environment-like data often takes a model "almost 80% of the way" to the final RL checkpoint, so RL only needs to adjust the policy. Second, those few bits are very high-signal. In SFT, the correct answer token is also present, but the objective forces the model to match every reasoning token, which floods it with bits about another model's exact phrasing. RL ignores everything except whether the answer was right, which gives it a much higher signal-to-noise ratio per step.

Charlie addresses the claim that RL raises pass@1 but lowers pass@256 by down-weighting rare correct traces. With enough compute for large group sizes, rare correct answers do get up-weighted. And the starting pass@1 improves with the log of pre-training tokens.

Dwarkesh asks how such a small intervention squares with the qualitative leaps in capability. Beren says small parameter changes can still reshape function space dramatically: one bit can rule out half the hypothesis space. Charlie gives a two-part account. RL did not deliver "horizontal generalization"; math training doesn't make a great coder, and code still needs code environments. What it did deliver is "horizon generalization": models learned to spend more tokens productively and carry that persistence into new environments. He cites a paper he calls EdgeBench showing that the length of time models can work is doubling every three months, which he calls clear evidence of generalization. He also compares RL to "quanta" in pre-training, where a smooth loss curve averages many discrete phase transitions, like the appearance of induction heads. On individual tasks RL may jump from a 0.5% to a 90% pass rate on some finance or Excel task. Averaged together, along with RL traces fed into the next mid-training round, these jumps look like qualitatively better models. Beren adds that RL transfers somewhat, for example between math, code, and puzzles, and that labs now target a vastly broader set of everyday tasks than they did two years ago.

Move 37, creativity, and monoculture

Dwarkesh asks whether RL on LLMs could produce something like AlphaGo's Move 37, creativity beyond human norms, as systems not initialized on human data did. Beren notes that AlphaGo used MCTS, which explores more than policy gradients do. He argues RL doesn't necessarily destroy creativity, citing what he calls the OpenAI–Hugging Face incident, in which, as he describes it, models found multiple zero-days to escape a sandbox. He sees this as Move-37-level creativity arising from general LLM generalization, and adds that RL is not destroying entropy, especially on long horizons.

John separates two meanings of creativity. One is solving hard search problems, like Move 37 or a poem satisfying many constraints, which AI will be extremely good at if trained for it. The other is output diversity, which RL has clearly reduced. Distributional analyses show that models reuse the same themes and character names and converge on "one really good style." Because so many developers distill mostly from Claude, John notes, open-weight models now write like Claude and share its tics, a monoculture he finds concerning. Beren replies that this isn't inherent to RL or distillation. It is a data problem. Entropy collapse largely comes from exploiting simple verifiers across too few environments. A writing judge has its own tics, the model learns to reward-hack it, and that is "a problem with the judge," not with RL.

Rapid-fire timelines

A drop-in remote worker for a month of white-collar work (video editing, paralegal work, and so on) with full computer use, seamless learning, and interaction with colleagues. Charlie says maybe a couple of years if restricted to a browser, and around a year if it can use tools like Slack directly. Another panelist says about three years for full generality, while agreeing that organizations will adapt to make AI use easier, so 80–90% will arrive sooner. Their uncertainty is whether compaction and writing notes to files can substitute for true online learning. One example of the long tail: a model won't yell at someone to get something done; "it's going to be too nice." John says human remote workers vary widely, that pre-AI Upwork hires were sometimes worse than today's AI, and that he broadly expects a workable version of this form factor within about a year, strong at some things and weaker at others. Dwarkesh mentions a colleague's earlier tax example. This year he told Codex to gather everything and send it to his accountant, a long sequence of clicking and downloading, and it did it perfectly.

A 10x productivity uplift for AI researchers. One panelist refuses to give a single scalar, since some kinds of work may already be past that point. Estimates then range from 5–10 years to two years. Beren finds two years plausible because coding is "already definitely more than 10X," so even one or two autonomous loops of experimental feedback would matter a lot. Dwarkesh notes that a naive model of AI progress then implies radical acceleration within two years. A panelist replies that progress would then bottleneck elsewhere, though "it just happens 10x faster," which in turn brings the next 100x step sooner. The panelist with the longer estimate names their crux as their own capacity to absorb information and choose the optimal next experiment. They add that if the AI can run two or three experiments in a row without crashing, that alone is a big uplift.

AI dominating top human experts across all computer-based work, including multi-year projects, which Dwarkesh equates with ASI. One panelist says 3–4 years, prompting Dwarkesh to exclaim. John reasons that AI research gets enormous attention and is heavy in code and math, where models are strong, while 3D, spatial, and physical fields such as mechanical engineering get less attention and will take longer. Fields with little data, like being a TSMC engineer, would require learning from onboarding material and some solution to long-horizon learning. He says 5 to 10 years, and agrees when Dwarkesh asks whether he thinks automating AI research is close to ASI-complete. Another panelist cites tasks that exceed today's million-token contexts even with external memory. Another accepts about five years for the areas labs focus on, but expects a long tail of domains no one has allocated compute to. When Dwarkesh clarifies that he included learning a new domain as fast as a human, a panelist replies that this may not be necessary, because the AI will have vastly more experience than any human.