Andrej Karpathy on Why Agents Will Take a Decade: Ghosts, Terrible RL, and the March of Nines

Open on YouTube ↗
Overview

In this long conversation with Dwarkesh, Andrej Karpathy explains why he thinks current AI timelines are too optimistic, even though he calls himself very bullish on the technology. He calls this "the decade of agents" rather than "the year of agents." His view is that today's models are impressive but cognitively incomplete. Reinforcement learning as currently practiced is crude. Deployment follows a slow "march of nines." He expects AI to blend into the long-running trend of automation rather than break it. The conversation then turns to his education project, Eureka, and his views on how to teach technical material well.

38 min read

Why "the decade of agents"

Karpathy says his phrase was a reaction to someone else's claim that this would be "the year of agents." He saw that as over-prediction in the industry. He uses early agents like Claude and Codex every day and finds them "extremely impressive," but he thinks there is still a great deal of work to do.

His test for an agent is whether you would hire it like an employee or intern. Nobody would do that today, and his explanation is simply that "they just don't work." They lack enough intelligence and multimodality. They cannot really use a computer. They have no continual learning, so you can't tell them something and expect them to remember it. He calls them "cognitively lacking."

When Dwarkesh asks why the number is a decade and not one year or fifty, Karpathy admits it is intuition. He has spent roughly 15 years in AI, in both research and industry, watching people make predictions and seeing how they turned out. The problems feel "tractable" and "surmountable" but still difficult. "If I just average it out, it just feels like a decade to me."

Fifteen years of seismic shifts, and agents attempted too early

Karpathy describes several shifts that reoriented the whole field. He got into deep learning by chance, working near Geoff Hinton at the University of Toronto, when neural networks were a niche interest. AlexNet was the first major shift. After it, everyone trained neural networks, but each network did one task, such as an image classifier or a translator.

Next came an early push toward agents. The Atari deep reinforcement learning work around 2013 tried to build systems that act and collect rewards, not just perceive. Karpathy now calls this "a misstep," and says early OpenAI, where he worked, shared it. For several years the field focused on reinforcement learning in games. He was always suspicious that games would lead to AGI, because what he wanted was something like an accountant interacting with the real world.

His own OpenAI project was part of the Universe effort: an agent that used a keyboard and mouse to operate web pages. In hindsight, he says it was "way too early." An agent mashing keys and clicking randomly gets rewards so sparse that it never learns. "You're going to burn a forest computing, and you're never going to get something off the ground." The missing ingredient was representation power. Today's computer-use agents are built on top of large language models, and that pretraining had to come first.

He sums up the history as three buckets: per-task neural networks, a premature first round of agents, and then LLMs that supplied representations before anything else was added. Agents are much more capable now, but he suspects parts of the stack are still missing.

"We're building ghosts, not animals"

Dwarkesh presents Richard Sutton's view from an earlier episode. Animals learn everything from scratch without labels, so perhaps AGI should also learn from raw sensory data rather than from millions of years' worth of training data. Karpathy says he is careful with animal analogies, because animals come from a very different optimization process. A zebra runs and follows its mother minutes after birth. He says that is not reinforcement learning but something "baked in." Evolution somehow encodes neural network weights in DNA, and "I have no idea how that works."

His framing, from a blog post he wrote about the Sutton interview, is that we aren't building animals. We are building "ghosts or spirits," trained by imitating human data on the internet. These are fully digital entities that mimic humans and start from a different point in the space of possible intelligences. He thinks making them somewhat more animal-like over time is a good goal. A single algorithm that learns everything from the internet would be "incredible," but he isn't sure one exists, and he points out that animals don't work that way either, because they have evolution as an outer loop. Much of what looks like animal learning, he suggests, is really brain maturation. He adds that humans probably use reinforcement learning mostly for motor tasks like throwing a ball, not for intelligence tasks like problem solving.

Dwarkesh notes that DNA holds only about three gigabytes, far too little to specify every synapse. So evolution seems to supply a learning algorithm rather than knowledge. Karpathy agrees there is "miraculous compression" and that learning algorithms are encoded which then learn online. His own stance is practical: "I have a hard hat on." We can't run evolution, but imitating internet documents works and produces something with a lot of built-in knowledge and intelligence. That is why he calls pretraining "crappy evolution": the version of that starting point our technology can actually reach.

When Dwarkesh presses on the knowledge-versus-algorithm distinction, Karpathy says pretraining does two unrelated things. It absorbs knowledge, and it also becomes intelligent by picking up algorithmic patterns and building internal circuits for things like in-context learning. He thinks the knowledge may hold models back, because it makes them lean on memory and struggle to go "off the data manifold." A research direction he wants is stripping away knowledge while keeping what he calls the "cognitive core": the algorithms and problem-solving strategies of intelligence.

In-context learning as working memory

Dwarkesh suggests that models seem most intelligent in context, such as when they catch their own mistake and back up. In-context learning emerges from gradient descent during pretraining but may not itself be gradient descent. Karpathy partly disagrees. In-context learning is pattern completion inside a token window, but he cites work on in-context linear regression. In that work, researchers found analogies to gradient descent in the learned weights, and one paper even hand-coded a transformer's weights to perform gradient descent through attention. He concludes that in-context learning is "probably doing a bit of some funky gradient descent internally," though nobody knows for sure.

Dwarkesh offers a numerical contrast. Llama 3's 70B model, trained on about 15 trillion tokens, stores roughly 0.07 bits per training token in its weights. By comparison, the KV cache grows by about 320 kilobytes per token in context, which is about a 35-million-fold difference. Karpathy agrees and frames it this way: anything learned in training is a "hazy recollection," because the compression is extreme. The context window works like working memory that the network can access directly. His practical example: ask an LLM about a book such as one of Nick Lane's and you get roughly correct answers, but paste in the full chapter and the answers get much better.

Which brain parts are still missing

Asked what part of human intelligence models most fail to replicate, Karpathy answers "just a lot of it," and gives a loose brain analogy. The transformer is general and plastic, so it resembles cortical tissue. He notes experiments where rewiring visual cortex to auditory cortex still produced a functioning animal. Reasoning traces in thinking models might resemble the prefrontal cortex. RL fine-tuning might loosely correspond to the basal ganglia. But he asks where the hippocampus is, and says the amygdala and other ancient nuclei for emotion and instinct are absent. Some parts, like perhaps the cerebellum, may not matter for cognition. As an engineer, he isn't sure we should try to build a brain analog. The practical point is that you still wouldn't hire this thing as an intern.

On continual learning, Dwarkesh raises the idea that it might emerge on its own if models are trained with an outer RL loop spanning many sessions. Karpathy says he doesn't really resonate with that. Models restart from zero tokens every time. His human analogy is that during the day he builds up a context window, and during sleep something distills it into the brain's weights. LLMs have no equivalent phase, one where they would analyze experiences, generate synthetic data, and distill it back, perhaps into a per-person LoRA that changes only a sparse subset of weights. He also expects very long contexts with elaborate sparse attention, and points to DeepSeek v3.2's sparse attention as an early sign. His broader view is that researchers are rediscovering cognitive tricks evolution found through a completely different process and may converge on a similar cognitive architecture.

What models will look like in ten years

Karpathy uses "translation invariance in time" to forecast. Ten years ago, in 2015, convolutional networks dominated, residual networks were new, and the transformer didn't exist. So he expects that in ten years we will still train giant neural networks with forward and backward passes and gradient descent, but with different details and at much larger scale.

He describes reproducing Yann LeCun's 1989 convolutional network, which he believes was the first modern-style network trained with gradient descent, on digit recognition. Applying 33 years of algorithmic improvements halved the error. Further gains required a training set ten times larger, more computational optimization, and longer training with dropout and other regularization. His lesson is that data, hardware, kernels and software, and algorithms all have to improve together, and "no one of them is winning too much." Dwarkesh is surprised that 30 years of progress only halved the error. Karpathy replies that "half is a lot," and that what struck him was that architecture, optimizer, and loss function all had to improve at once.

nanochat and the limits of coding agents

Karpathy had just released nanochat, about 8,000 lines of code covering the whole pipeline for building a ChatGPT clone. He says he didn't learn much new from it, since he already knew how to build it. The work was making it clean enough for others to learn from. His advice for learning from it: put it on a second monitor and rebuild it from scratch, referring to it but never copy-pasting. He notes that the final repository hides how it was actually built, chunk by chunk, and says he would like to add that process later, probably as a video. He cites the Feynman idea that if you can't build it, you don't understand it. Building forces you to confront things you didn't know you didn't understand. "Don't write blog posts, don't do slides… Build the code."

He had tweeted that coding models were of little help with nanochat, and he explains why. He describes three ways people write code today: fully by hand, which he thinks is no longer right; by hand with autocomplete, which is where he sits and where he remains "the architect"; and vibe coding with agents. Agents do well with boilerplate and with code patterns that appear often on the internet. nanochat is neither. It is "intellectually intense code" where everything must be precisely arranged.

His concrete example: he didn't use PyTorch's DistributedDataParallel container. Instead he wrote his own gradient synchronization inside the optimizer step. The models kept trying to get him to use DDP and "couldn't get past that." They also added defensive try-catch statements, tried to turn his code into a production codebase, used deprecated APIs, and bloated the code. "It's just not net useful." He also finds typing out requests in English too slow. Going to the right spot in the code and typing a few characters for autocomplete is a much higher-bandwidth way to say what he wants.

He did use models in two places. He partly vibe-coded the report generation, which is boilerplate and not mission-critical. He also used them when rewriting the tokenizer in Rust, a language he is new to. There he had a Python version he fully understood, plus tests, so he felt safe. He says models make unfamiliar languages much more accessible.

Dwarkesh ties this to forecasts of an intelligence explosion driven by AI automating AI research. Karpathy agrees this is why his timelines are longer: models are "not very good at code that has never been written before," which is exactly what building new models requires. Even well-known tweaks like RoPE embeddings are hard for them, because "they know, but they don't fully know" how to fit them into a specific repository's style and assumptions. He calls GPT-5 Pro, which he sometimes consults with his whole repo pasted in, surprisingly good compared with a year ago. Still, he thinks the industry is "trying to pretend like this is amazing, and it's not. It's slop," possibly for fundraising reasons.

Karpathy has trouble drawing a line between AI and computing in general. He sees a continuum of programmer speedups: compilers, syntax highlighting, type checking, search engines, now better autocomplete and "loopy" agents that sometimes go off the rails. He calls it an "autonomy slider." Humans do less of the low-level work and move up a layer of abstraction.

"Reinforcement learning is terrible"

Dwarkesh asks how humans gain rich understanding from experience without the kind of end-of-episode reward RL uses. His example is a founder who learns a great deal from ten years running a business, but not by re-weighting every action according to the final outcome. Karpathy repeats that humans don't really use reinforcement learning. RL is "a lot worse than I think the average person thinks." It only looks good because the earlier approach, pure imitation, was worse.

His example is a math problem. RL tries hundreds of attempts in parallel and checks each against the answer. Say three succeed and 97 fail. Every token in the successful attempts, including wrong turns that happened to precede the right answer, gets upweighted as "do more of this." He calls this "sucking supervision through a straw": minutes of work reduced to one number, which is then spread across the whole trajectory. "It's just stupid and crazy." A human would never make hundreds of attempts. After finding a solution, a human reviews which parts went well and which didn't. Nothing in current LLMs does that.

He puts this in historical context. InstructGPT "blew my mind" by showing that fine-tuning a base autocomplete model on conversation-like text quickly made it an assistant while keeping its knowledge. RL was the next step. It lets models hill-climb on reward functions without expert demonstrations and even find solutions humans wouldn't. "Yet, it's still stupid. We need more." He mentions a Google paper he saw the day before with a reflect-and-review idea, possibly a "memory bank" paper, and expects major algorithmic updates in this direction. His guess is that "three or four or five more" such changes are needed.

Why process supervision is hard: the "dhdhdh" problem

Process-based supervision would reward each step rather than only the final result. Karpathy says the difficulty is assigning partial credit automatically. Labs use LLM judges prompted to grade partial solutions, but those judges are huge, gameable models. Optimize against them and "you will find adversarial examples… almost guaranteed." It might work for 10 or 20 steps but not for 100 or 1,000.

He gives an example he believes was public. While training against an LLM-judge reward, the reward suddenly jumped to perfect. The model's outputs began reasonably and then turned into "dhdhdhdh." That string was an out-of-distribution adversarial example the judge rated at 100%. Dwarkesh compares this to prompt injection. Karpathy says it's simpler than that: the outputs are nonsensical adversarial examples. You can add such cases to the judge's training data, but each new judge still has adversarial examples. Iterating may make them harder to find, but he isn't sure it converges, given a model with "a trillion parameters or whatnot." He assumes labs are trying, but thinks "we need other ideas."

Asked what those ideas look like, he points to reviewing solutions, generating synthetic examples that improve the model, and meta-learning this process. He says he is "only at a stage of reading abstracts." The papers are interesting ideas, but he hasn't seen anyone show convincingly that they work at frontier-lab scale, though he acknowledges the labs are closed.

Reflection, synthetic data, and model collapse

Dwarkesh asks what the machine-learning analog of daydreaming, sleep, or reflection would be. Karpathy uses reading a book as his example. LLMs just predict the next token of the text. For a human, a book is "a set of prompts for me to do synthetic data generation," or material to discuss at a book club. Understanding comes from manipulating the information. He would like to see a pretraining stage where the model thinks through material and reconciles it with what it already knows.

The subtle obstacle is collapse. A single synthetic reflection may look great, but training on many of them makes the model worse. Model samples are "silently collapsed": each one looks fine, but together they cover a tiny slice of the possible space. His example is asking ChatGPT for a joke, which gives you roughly three jokes. Ask for reflections on a chapter ten times and you get ten nearly identical answers. Humans are noisier but maintain much more entropy.

He suggests humans also collapse over a lifetime. Children say surprising things because they haven't overfit yet, while adults repeat the same thoughts. Dwarkesh mentions a paper proposing that dreaming evolved to prevent overfitting by placing us in strange situations. Karpathy finds that interesting and adds that you "always have to seek entropy in your life." Talking to other people is one good source.

Dwarkesh offers a tentative idea about a spectrum. Young children learn best but remember almost nothing. LLMs can recite Wikipedia but struggle with fast abstraction. Adults fall in between. Karpathy agrees strongly. Humans' poor memorization is "a feature," because it forces them to learn generalizable patterns. LLMs can memorize a hashed random sequence after one or two passes, and their memories distract them. That is why he wants a cognitive core with less memory, which would look things up and keep only the algorithms of thought. He sees this as mostly separate from collapse.

On fixing collapse, he says naive approaches such as entropy regularization or wider logit distributions haven't worked well empirically. Part of the reason is that most tasks don't require diversity, and RL actively penalizes creativity. Labs optimize for usefulness, so diversity isn't built in, which then hurts synthetic data generation. "We're shooting ourselves in the foot." When Dwarkesh says Karpathy had implied the problem was fundamental, he pulls back. He hasn't done the experiments and thinks entropy could probably be raised, but controlling the distribution is tricky, because too much entropy produces invented language and rare words.

How small can a cognitive core be?

Karpathy has predicted that a cognitive core of about a billion parameters could be very capable. In 20 years, he says, you might have a productive, human-like conversation with one. It would know when it doesn't know something and look it up. Dwarkesh thinks this is too large. gpt-oss-20b already beats the original GPT-4, which had over a trillion parameters, so why not tens of millions?

Karpathy's answer is about data. A random document in a frontier pretraining set is "total garbage," with stock tickers and slop from every corner of the internet, "not like your Wall Street Journal article." Big models are needed to compress all of that, and most of that compression is memory rather than cognition. Smarter models could help clean pretraining data down to its cognitive parts, and small models would probably be distilled from better large ones. He grants it might go somewhat below a billion, but thinks a model needs some basic knowledge so it can "think in your head" without looking up everything.

On frontier model size, he has no strong prediction. He says labs are being practical with their FLOPs and cost budgets. Pretraining turned out not to be the best place to spend, so models got smaller while reinforcement learning, mid-training, and later stages grew. He expects no giant paradigm shift in the near term, but many incremental gains: much better datasets, better hardware such as Nvidia's work on Tensor Cores, better kernels, better optimization and architecture. "Nothing dominates. Everything plus 20%."

Measuring progress: AGI definitions, radiologists, and call centers

Asked how to chart progress toward AGI, whether by education level or task horizon length, Karpathy is "almost tempted to reject the question," since AI is an extension of computing and nobody plots a single y-axis for computing. He holds to the original OpenAI-era definition: a system that can do any economically valuable task at human level or better. He notes that people usually drop physical work from it, which is a big concession. He guesses knowledge work is roughly 10–20% of the economy, which is still trillions of dollars in the US.

He points out that Geoff Hinton's prediction that radiologists would disappear "turned out to be very wrong." Radiology is growing, because the job is messy and involves patients and context beyond image recognition. He thinks call center work is a better early candidate. Tasks are similar and repeated, short, closed, purely digital, and involve little outside context. Even there, he expects an autonomy slider rather than full replacement: AIs handling 80% of the volume, humans taking the remaining 20% and supervising teams of perhaps five AIs, with new companies building interfaces to manage imperfect AI.

Dwarkesh offers a hypothesis. If AI automates 99% of a job, the human doing the last 1% becomes the bottleneck and may earn more, until that last piece is automated too. Karpathy doesn't think radiology shows this and repeats that it's a poor example. He would watch call center trends. He would also wait a year or two after companies swap in AI, since some might rehire, and he mentions evidence that this has already happened at some companies adopting AI.

Dwarkesh notes that a supposedly general technology is mostly used for coding: API revenue is dominated by it. Karpathy's explanation is that coding is built around text, which LLMs process well, and that infrastructure already exists, such as IDEs and diffs an agent can plug into. Slides are spatial and visual and have no diff tool, so someone has to build one. Dwarkesh isn't convinced this explains everything. His own attempts to get models to rewrite transcripts or suggest clips were disappointing, and Andy Matuschak tried few-shot prompting, supervised fine-tuning, and retrieval without getting good spaced-repetition cards. Karpathy concedes he has no great answer. Code is more structured, text may have more entropy, and code is hard enough that simple help feels empowering.

Superintelligence, loss of control, and the 2% growth debate

Karpathy sees superintelligence as an extension of gradual automation, first digital and later physical. Invention is one of the things that gets automated. He expects the result to feel "extremely foreign." The outcome he considers most likely is "a gradual loss of control and understanding": layers of automation spread everywhere, understood by fewer and fewer people.

Dwarkesh argues that control and understanding differ; a US president or elderly corporate board has power without understanding. Karpathy expects to lose both. His speculative sci-fi picture is not one AI taking over, but many competing entities becoming gradually more autonomous, some going rogue and others fighting them. Even if each acts on someone's behalf, society as a whole may lose control of outcomes.

On an intelligence explosion, Karpathy says we have been in one for decades, even centuries. The GDP curve is an exponential that reflects ongoing automation, from the Industrial Revolution to compilers. He had tried to find AI in GDP data and concluded it wouldn't show up. He couldn't find computers or the iPhone there either. The first iPhone lacked the App Store, and technologies spread slowly enough to average into the same curve. AI is "a new kind of computer," and he expects it to keep growth near the same roughly 2% trajectory. Recursive self-improvement is "business as usual." Engineers using Google Search, IDEs, autocomplete, and Claude Code are all part of it.

Dwarkesh argues the opposite. True AGI is labor itself, and in a labor-constrained world, billions of human-like minds that start companies and integrate themselves would be like adding 10 billion people. He points to Hong Kong and Shenzhen, where decades of 10%+ growth came from many capable people, and to the Industrial Revolution's jump from about 0.2% to 2% growth as a precedent. Karpathy says he is "pretty willing to be convinced," but pushes back. Computers were already labor, and so is self-driving. He thinks the "God in a box" assumption is wrong: AI will succeed at some things, fail at others, and diffuse gradually. He is somewhat suspicious of pre-industrial growth data. Dwarkesh clarifies that the Industrial Revolution wasn't one magical invention but an overhang being unlocked, and that population growth historically drove growth. Karpathy says he understands the view but doesn't "intuitively feel" it. The disagreement is left open.

Evolution of intelligence and the absence of LLM culture

Discussing Nick Lane's work, Karpathy says he is surprised intelligence evolved at all. He would have expected worlds full of animal-like life doing animal-like things. He suggests judging difficulty by how long a bottleneck lasted. Bacteria and archaea went about two billion years without becoming complex life. Animals have existed for a few hundred million years, perhaps 10% of Earth's history. The separate emergence of cleverness in birds such as ravens, with very different brain structures, hints that intelligence may have arisen more than once.

Dwarkesh relays Gwern's and Carl Shulman's view that humans found a niche that rewarded marginal intelligence. A bird with a bigger brain couldn't fly. Hands rewarded tool use. External digestion freed energy for the brain. Karpathy adds that a dolphin couldn't have fire. Dwarkesh also cites Gwern's point that intelligence requires environments unpredictable enough that evolution can't bake solutions into DNA, yet important enough to learn. Karpathy agrees that humans have to "figure it out at test time."

Dwarkesh mentions Quintin Pope's argument that humans needed about 50,000 years to build a cultural scaffold, which AI gets for free through shared corpora and distillation. Karpathy says "yes and no": LLMs have no real culture. He imagines a scratchpad that an LLM edits for itself, or an LLM writing a book that other LLMs read and are inspired by. He names two unclaimed multi-agent ideas: culture, and self-play like AlphaGo's, where one LLM creates harder and harder problems for another. He hasn't seen either done convincingly. Organizations of AIs are missing too. The bottleneck, he says, is that the models are still cognitively like children. Claude Code and Codex feel like elementary-school students, "savant kids" with perfect memory who can pass PhD quizzes and produce convincing slop but "don't really know what they're doing."

Self-driving and the march of nines

Karpathy led self-driving at Tesla from 2017 to 2022. He first pushes back on the idea that the problem is solved. Demos date back to CMU's self-driving truck in 1986. Around 2014, a friend gave him a perfect Waymo ride around Palo Alto, and he thought it was close, but it took much longer. Some fields have a large "demo-to-product gap," especially when failure is costly. He argues that production software shares this property, since one mistake can leak millions of people's Social Security numbers. In that sense software's potential harm is "almost unbounded."

His central idea is the "march of nines." Working 90% of the time is the first nine, and each additional nine takes about the same amount of work. In his five years at Tesla, the team got through perhaps two or three nines, with more still to go. This is why he is "very unimpressed by demos," especially staged ones. Dwarkesh adds that humans crash about once every 400,000 miles, or seven years. A coding agent would hit an equivalent number of tokens much faster in wall-clock time, and software engineering covers far more ground than driving.

Dwarkesh raises a counterargument: much of self-driving's difficulty was basic perception and common sense, which LLMs and VLMs now provide. Karpathy is "not 100% sure" he agrees. General models bring more generalizable intelligence, but they are still fallible and full of gaps. He also argues that self-driving is far from done. Waymo's deployments are small and not yet economical given capital costs. Driverless cars can be misleading, since teleoperation centers keep more humans in the loop than people expect. "We haven't actually removed the person, we've moved them." He speculates, calling it "just making stuff up," that Waymo avoids some areas because of poor connectivity. He says he loves Waymo and rides it often, and thinks Tesla has the more scalable approach. For him the timeline runs from the 1980s and isn't finished, since "self-driving at scale" means people no longer needing driver's licenses.

Dwarkesh notes two differences: knowledge work lacks the tight latency limits of driving, and copying a model is much cheaper than building a car. Karpathy agrees that "bits are a million times easier" than the physical world, though at scale compute will still impose practical latency constraints. He adds the social layers: law, insurance, public reaction. He asks what the AI equivalent is of people putting cones on Waymos, or of the hidden teleoperator.

Is compute being overbuilt?

Asked whether slower adoption means compute is being overbuilt, as with railroads or the telecom boom of the late '90s, Karpathy says he doesn't think so. He is "actually optimistic," and only sounds pessimistic in reaction to Twitter claims he attributes to fundraising and attention-seeking. Claude Code and Codex barely existed a year ago, and demand like ChatGPT's suggests the buildout will be used. His concern is calibration. He has seen reputable people get timelines wrong many times over 15 years, and some of these questions have geopolitical consequences.

Eureka and the Korean tutor standard

Why education instead of another AI lab? Karpathy feels some "determinism" in what labs will do and doubts he would change it uniquely. His "big fear" is humanity being sidelined, as in WALL-E or Idiocracy. He cares about "what happens to humans," not only the Dyson spheres AI might build, and sees education as the way to help. His project, Eureka, aims to be something like Starfleet Academy: an elite, up-to-date institution for frontier technical knowledge.

He thinks AI will change education fundamentally, but that the obvious approach of asking an LLM questions still feels "a bit like slop." His standard comes from learning Korean. He studied alone online, then in a class of about ten in Korea, then with a one-on-one tutor. After a short conversation, she understood exactly what he knew, probed his mental model, and always gave him material at the right difficulty. "I felt like I was the only constraint to learning." No LLM does this today, "not even close." So he thinks it isn't yet the time to build that kind of AI tutor, just as his consulting advice in computer vision was often "don't use AI."

For now, he is building something more conventional, with physical and digital parts. The first product is an AI course, LLM101N, with nanochat as its capstone. He is building the intermediate material and hiring a small team of TAs. He calls education a hard technical problem of "building ramps to knowledge," measured in "eurekas per second." AI already helps a lot. Compared with building CS231n at Stanford, which he believes was Stanford's first deep learning class, he now moves much faster with LLMs doing tedious work. But the models can't yet create the content, and asking ChatGPT to "teach me AI" produces slop. Later he expects AI TAs for basic questions, with human faculty still designing course architecture, and perhaps eventually AI designing courses better than he could. He plans to hire faculty for fields outside his expertise. He imagines a full-time physical program as the top tier, plus a cheaper digital tier reachable by 8 billion people.

Pre-AGI education is useful, post-AGI education is fun

Before AGI, Karpathy says, motivation is simple: people want to earn money, and many are reskilling in AI. After AGI, he compares education to the gym. Machines lift heavy things, yet people still work out because it's healthy, fun, and attractive. Learning is hard today because people bounce off material that is too easy or too hard. Solve that technical problem, as his Korean tutor did, and learning becomes easy and enjoyable. "Anyone will speak five languages because why not?" Dwarkesh compares this to how common strength-training feats have become compared with a century ago. Karpathy is betting on "the timelessness of human nature," and points to aristocrats and ancient Greece as small "post-AGI" pockets where people flourished physically and intellectually.

Dwarkesh asks how this fits with keeping humanity in control. Karpathy admits that in the long run "it's a bit of a losing game," though there may be a transition period where people who understand a lot can stay in the loop. Eventually it might become a sport, a cognitive version of powerlifting. He believes today's geniuses "are barely scratching the surface" of what a mind can do. He mentions that he loved school and stayed through his PhD, and that he enjoys learning both for its own sake and as empowerment. Dwarkesh asks why online courses haven't already achieved this. Karpathy says bouncing off material "feels bad," like negative reward, and fixing that takes AI and human collaboration first, and perhaps AI alone later.

How to teach, and how to learn

Karpathy credits physics for much of his teaching approach, and says everyone should learn it early because school is about "booting up a brain." Physics teaches model-building, first-order approximations with higher-order corrections, and pulling signal out of noise. The "spherical cow" joke is, in his view, "brilliant." He recommends Geoffrey West's book Scale, which relates animal properties like heartbeat to size, and notes the example that heat dissipation grows with surface area while heat generation grows with volume.

His method is to find the first-order term and serve it "on a platter." micrograd is his example: about 100 lines of Python that do backpropagation on arbitrary networks through the recursive chain rule. "Everything else is just efficiency": tensors, strides, kernels, memory movement. He finds untangling knowledge into a ramp, where each step depends only on the one before, deeply satisfying. Dwarkesh praises his transformer tutorial for starting from a bigram lookup table and motivating each added piece. Karpathy describes this as presenting the pain before the solution. He also always prompts the learner to try first. Giving the answer before a student has attempted it is "a dick move," because trying first shows the action space and the objective and makes the solution meaningful.

On why experts often explain badly, he names the curse of knowledge and says he suffers from it too. One fix: when a biology paper confused him, he asked ChatGPT his "dumb questions" with the paper in context, then shared the conversation with the author so they could see where newcomers struggle. He would welcome people sharing such conversations about his own material. He agrees with Dwarkesh's observation that people explain things better over lunch than in writing. From his PhD days, he recalls authors summing up a paper in three perfect sentences over beers at a conference. "Why isn't that the abstract?"

As a learner, Karpathy says he has no special tricks and calls learning "a painful process." He alternates between depth-first, on-demand learning tied to a project that pays off, and breadth-first "101" learning of the "trust me, you'll need this later" kind. He prefers the first. His other habit is explaining things to others. If he can't explain something, he doesn't understand it yet, and that forces him to go back, fill the gaps, and make sure he knows what he's talking about.