Grant Sanderson on AI That Disproves Conjectures, and What Mathematics Still Can't Measure

Open on YouTube ↗
Overview

Dwarkesh Patel sits down with Grant Sanderson, creator of 3Blue1Brown, who is now working on a series documenting AI's progress in mathematics. The host's premise is that AI has advanced faster in math than in any other field, so how that progress unfolds may preview what happens elsewhere. The conversation keeps returning to a few questions. What kind of mathematical achievement would actually be a turning point? Can the things that make mathematics great, such as new definitions, good conjectures, and unifying concepts, be measured or trained at all? And what is left for human mathematicians, explainers, and students?

32 min read

The IMO Turned Out to Be "Just Another Benchmark"

Three years earlier, Dwarkesh had asked Sanderson whether an AI that won gold at the International Math Olympiad would effectively be AGI. Sanderson had predicted it would be one more benchmark, with no "aha" moment. Asked why that proved right, Sanderson points to what he calls the dirty secret of the IMO: although problem designers try to write problems that can't be trained for, students can in fact train for many of them.

He then describes AI's jagged frontier as fractal. Math is one of the spikes, but within math some things are far easier than others. The IMO has four categories: geometry, number theory, algebra, and combinatorics. By Sanderson's account, AI systems essentially "cold-solved" geometry by 2024, handling it in about nineteen seconds with a brute-force solver, and he notes that students have a brute-force route into geometry too. Combinatorics is the wild card, with more playful, puzzle-like problems. Six problems are spread across four categories, so which category gets two problems is a toss-up. In 2024 combinatorics got two. Sanderson says that had there been more geometry that year, the systems would have won gold, but they struggled on the combinatorics problems. Someone defending math as a last human holdout, he says, might argue those are the problems requiring more creativity.

Three Ways the Riemann Hypothesis Might Fall

Dwarkesh's follow-up: if an AI solves a Millennium Prize problem, could many economic tasks still resist automation? Sanderson says the answer depends on what the solution looks like, and he sketches distinct possibilities.

The first is the "lightning bolt" between fields. He tells the story of Hugh Montgomery and Freeman Dyson at the Institute for Advanced Study. Montgomery, a number theorist, was studying the statistical correlation between pairs of zeros of the Riemann zeta function. The Riemann hypothesis concerns whether all those zeros lie on a straight line. Montgomery wrote down a formula that looked something like one over sine squared. Dyson, a physicist, recognized it from the eigenvalue statistics of random Hermitian matrices, which arise in studying nuclear energy levels. That coincidence prompted exploration of whether random matrix theory might bear on the zeta function. Sanderson considers it somewhat open whether there is fruit there. If the eventual proof extended an idea like that, it would fit how one would expect LLMs to be good at math. They are expert in quantum physics and in analytic number theory at once, so they shouldn't need a chance lunch conversation to notice the similarity. That kind of skill, he argues, is quite distinct from what makes someone good at white-collar work. If you struggle to use AI as an editor, the reason isn't that it knows everything and just needs to connect two things.

The second possibility is "mountain building." Fermat's Last Theorem can be stated simply (no integer solutions to xⁿ + yⁿ = zⁿ for larger exponents), and one might expect an elementary proof, but as far as anyone can tell there isn't one. The actual proof rests on centuries of work around elliptic curves and another mountain of ideas around modular forms, and both had to exist before anyone could ask the question that connected them. If a Riemann hypothesis proof required building a new mountain, Sanderson thinks that capacity for the right new ideas is so different from how current systems are intelligent, and so powerful, that it would be surprising if it didn't permeate the economy. Even if it couldn't do every white-collar task, it would be transformative in a way IMO gold was not.

The third possibility, raised later, is that an AI simply works harder: a thousand-page chain of reasoning with no new theory, like a hypothetical elementary proof of Fermat's Last Theorem spelled out incoherently over thousands of pages.

Conjectures and Definitions as the Next Frontier

Dwarkesh admits he is moving the goalposts. When he interviewed Dario Amodei two or three years ago, he asked why models with so much knowledge weren't connecting ideas to make discoveries. The disproof of the unit distance conjecture now looks like exactly that kind of connection. So what comes next? He proposes two candidates: generating interesting problems in the first place, as Riemann did when he suspected the zeta function's zeros relate to the density of primes, and inventing new objects or conceptualizations that create or unify fields.

Sanderson recommends a video on the unit distance conjecture by the channel Polylog. He cites a quote from someone in it: good mathematicians prove theorems, great mathematicians come up with conjectures, and the greatest come up with definitions. But he doubts this can take the form of a benchmark. A benchmark is a goalpost the ball either passes or doesn't, which matters both for RL with verifiable rewards and for knowing no one moved the goalposts. OpenAI can headline a disproof because it is a clear, discrete result. Imagine instead a headline about a model producing "a really good conjecture," with a promise that everyone agrees it's good. That doesn't land the same way.

Instead, Sanderson expects the signal to be a tone shift among mathematicians. His series, which he says is months from release, consists largely of interviews with mathematicians that began more than a year ago. He already hears a shift in how they talk about AI between mid-2025 and 2026, a short time in the real world but "eons" in AI. Conjecture-generating ability will show up, he predicts, when mathematicians say a conversation with some model genuinely helped them decide what their research field should be.

Dwarkesh observes that what can't be benchmarked also can't easily be trained, since there's no fundamental difference between a benchmark and a training environment. He acknowledges that claims of "deep reasons AI can't do X" often collapse quickly, and he expects ways to train these abilities in the near term, but thinks they would have to differ from current RLVR.

Galois and the Hundred-Year Verification Loop

Dwarkesh's example is Galois inventing group theory to explain why the quintic has no general formula, while Abel had proved the same impossibility a few years earlier without it. If you wanted to verify that group theory was a valuable concept, the loop might be a century long, running through cryptography and symmetry in physics.

The example "struck a nerve" for Sanderson. He spent a year of his life on a Galois project in 2022 that he shelved, and he walks through the history. The quadratic formula was known (he says the Greeks could solve quadratics without writing algebra, and Arab mathematicians wrote down the formula). Dueling Italian mathematicians secretly found formulas for the cubic and then the quartic. The quartic formula is a monster, and for hundreds of years nobody resolved the quintic. Abel, a young Norwegian who initially thought he had found a formula, is usually credited with proving it impossible.

Sanderson gives real credit to Lagrange, who recognized that solving polynomials is tied to how algebraic expressions behave under permutation. The expression a + b + c + d is unchanged by any permutation, while a + b·c + d is changed by some. Lagrange saw that an expression in four variables that takes only three distinct values across all permutations has an unexpected relationship to reducing degree four to degree three. Extending the method to quintics would require an expression in five variables taking four or fewer values across all 5! permutations. Sanderson calls that a puzzle a twelve-year-old could engage with, and one that quickly feels impossible. Lagrange didn't solve anything, but he was the first to sense that symmetry was the right lens, and he planted the idea of proving impossibility rather than hunting for a formula. At the time this was nearly a recreational side topic, since most people cared more about physics.

Roughly fifty years later, Abel read Lagrange. Galois loved Lagrange when he was falling in love with math. Abel was advised to focus on elliptic functions and died of tuberculosis at twenty-six. Galois, a teenager, had his papers rejected. Sanderson says frankly that they weren't very coherent or complete: "the verified reward there is, 'No good.'" In prison Galois wrote a piece arguing that mathematics undergoes shifts of abstraction, as when algebra freed people from interpreting expressions as numbers, and that the next layer would be to think about the symmetries underlying formulas rather than the formulas themselves. Even describing what problem Galois solved is tricky. Abel had already proved general unsolvability, and while Galois theory in principle tells you whether a specific polynomial (x⁵ − 1 is solvable, x⁵ − 2 has the fifth root of two) can be solved by radicals, Galois didn't exhibit a specific unsolvable example. Sanderson notes the myth that Galois wrote everything down the night before his fatal duel, when in fact he had tried to publish five times. He asked his brother and a friend to get his notes to Gauss and others. About twenty years passed before Liouville saw promise in them, and about twenty more before Jordan produced something like a modern group theory attributed to Galois. Practical payoff came much later still. Sanderson highlights Gell-Mann anticipating quarks from a purely group-theoretic question in the twentieth century.

His point is that for much of this span the idea did not even pass human review. The real question is how to capture the instinct in Lagrange's, Galois's, or Liouville's mind that "there's something here."

Compression as a Possible Reward for Elegance

Sanderson connects this to another series he's making on the idea that "compression is intelligence." A smaller expression that is more predictive feels more intelligent. He wonders whether a reward could target the smallness of the concepts a proof requires, perhaps bringing in Kolmogorov complexity to quantify elegance. He doesn't think it's easy, but he believes something like it is needed to reward a Galois-like instinct rather than just problem-solving. It also bears on the end goal. Even if automated engineers built starships nobody understands, many people will still want understanding, the equivalent of Newton's law of gravitation distilled from a complicated way of thinking. Dwarkesh agrees that humans somehow manage this heuristic and that AIs will eventually do it too.

Will We Understand an AI's Proof?

Dwarkesh challenges the fear that AI will prove the Riemann hypothesis without improving human understanding. Isn't finding natural abstractions and subgoals simply the useful way to attack hard problems? And empirically, the unit distance counterexample's chain of thought, as he understands it, used known concepts in natural language that mathematicians could follow.

Sanderson says it depends on which of the three modes a solution takes. He cites another result from this year, Erdős problem #1196 on primitive sets, as a lightning-bolt case. Say to an expert something like "use a Markov chain process to show this probabilistically from the bottom up rather than top down, and use the von Mangoldt function," and they know how to run with it. Such results are very human-parsable because you only need to show both endpoints of the connection. Mountain building demands far more time to understand. Raw-hustle proofs would bring the full digestion worry.

The closest precedent for an alien mountain, he suggests, is the attempted proof of the abc conjecture by an otherwise reputable mathematician in Japan, built on what Sanderson calls "inter-universal geometry." Mathematicians took a long time even to parse it, and Sanderson thinks it probably isn't correct. The biggest fear would be an AI doing something similar: people spend years climbing the mountain only to find it isn't right after all. Even if it were right, the climb is enormous.

Dwarkesh brings up David Bessis's blog post "The Fall of the Theorem Economy." The argument is that theorem-proving gets the credit but is parasitic on the work of coming up with definitions, and this was never a credit problem because whoever made the definition usually proved the theorems too. If AI produced Abel-style direct proofs of many conjectures, humans or future AIs would be left to consolidate them. Sanderson thinks such proofs would still help enormously. Most of discovering new math is being wrong, "a random drunken walk," so knowing that digestion leads to a correct result is itself progress. He notes that recent mathematics already has results proven long before they're understood. His favorite example is Timothy Chow's expository paper on forcing, the method behind showing that the continuum hypothesis (whether an infinity lies between the naturals and the reals) is independent of the usual axioms. Chow proposes the idea of an "unsolved expository problem": proven, but not understood. Sanderson says that framing is his whole life. There is a difference between proof and explanation.

Discoverers Are Often Great Explainers, and What That Means for Sanderson's Job

Asked whether a conceptualization like Minkowski spacetime diagrams is distinct from the idea itself, Sanderson observes a strong correlation between people who produce genuinely novel insights and people who communicate clearly. That runs against the university experience of experts being poor teachers. Einstein, Claude Shannon, and Feynman wrote lucid work that doesn't need to be hacked through "with a machete." Perhaps the same faculty produces both.

This has changed his beliefs. He used to think AIs would become theorem provers and mathematicians would shift toward his kind of job, explaining. Now he suspects AIs will also explain well, likely better than most humans, so digesting and explaining is probably not what's left for mathematicians.

What remains, in his view, is relational: motivation and curation. He mentions a view he's heard that mathematicians may become like art museum curators. The AI made the art and can explain it, but people still want someone to help them navigate a near-infinite space of ideas worth engaging with. Even if AIs curated better, he thinks people would prefer a human they have a relationship with, because interest is socially motivated. Viewers trust his choice of topics, and he says much of his video time goes not into visuals but into deciding what is worth saying. He compares this to human musicians keeping a role through their stories even if a model's MP3 is objectively better.

Filling In the Map of Connections

Dwarkesh asks whether many great breakthroughs are themselves connections, such as general relativity joining Riemannian geometry and special relativity, in which case better connecting might go far. Sanderson says there's much more to come on connections alone. Only a couple of lightning bolts have been thrown. He also stresses that most mathematicians don't describe their work as targeting the next problem. He introduces the Langlands program as "a research ethos" more than a field. Fermat's Last Theorem is one instance, but Langlands wrote a famous letter suggesting there were many more such connections and got somewhat specific about their nature, like a map with valleys, mountains, and plains. Many mathematicians see their work as tracing threads on that map, preemptively finding connections because big problems have repeatedly fallen that way. He suggests asking any mathematician whether their work resembles Langlands or targets one problem, and he says you get a bifurcated split.

AI as a supercharged connector could amplify that work. But again it's hard to score, with no headline or "we did it" PR moment, and it will need much more human-in-the-loop judgment. Sanderson's guess is that most useful progress from these models over the next five years will be filling in that landscape of connections available to someone expert in multiple fields.

Why Haven't Connections Come Sooner? Autoregression Versus Data

Sanderson offers a thought experiment. You're locked in a box, handed slips of paper, asked to predict what comes next, and then your memory is wiped, over and over. Shown the resulting essay, you might say it's not what you would have written. Repeated prediction makes you "a slave to your context." Answering a question in one field draws on that field's context, while the substantive connection is by nature an unlikely one. What in RL specifically rewards unlikely connections when most aren't the predictable next token? He wonders whether questioning how tokens are generated could help (he doubts it's as simple as temperature), or whether more intelligence alone will make the model predict that it should draw the lightning bolt.

Dwarkesh thinks data is more productive to reason about than architecture or loss function. He notes that diffusion text models don't produce something of a wholly different character, though they're less explored. Agents got better because environments rewarded steps like "let's step back and search the whole codebase" or "reassess my mistake." He guesses that labs use FrontierMath-like problems designed to require connecting fields, plus partially synthetic ways to make them harder, such as removing assumptions. The key is an environment that incentivizes the ability. He'd be surprised if the next three years didn't bring many more lightning bolts.

Parallelism, Fresh Context, and Engineered Entropy

Dwarkesh emphasizes advantages digital minds have beyond raw smarts: they can be parallelized and scaled across every accessible problem, rather than being one idiosyncratic genius who "dies in a duel." They can potentially merge knowledge and spawn identical copies. With AI companies pouring billions into math for PR reasons, "quantity has a quality all of its own." Sanderson imagines automating the Montgomery–Dyson lunch: agents representing different fields, and engineering the serendipitous conversations that institutes exist to create.

But he suggests the opposite of pooling may matter more: the ability to deliberately discard context. AIs get stuck in bad chains of thought, and so do humans, and breakthroughs sometimes come from asking, "What if I tried to prove the opposite?" His planned first episode centers on an IMO problem the AIs failed. He says many very smart students and Terry Tao also failed it, and people called it a "troll problem." He avoids spoiling it. In the IMO context, an elegant-seeming approach is enticing but hard to prove optimal because it isn't. An almost "brain-dead" solution is best. Solving it requires escaping the context of contest training, and someone given it as a street brain teaser might do well. Spinning off agents with deliberately different contexts, one proving and one disproving, could systematize that. He wonders how many headline results in three years will have that character.

Dwarkesh ties this to the worry about entropy collapse, where similarly trained models all think alike, which he says is why they write poorly. He notes that the unit distance conjecture reportedly resisted disproof because people assumed it was true. AIs might increase entropy by systematically trying both a statement and its negation, or by assigning agents different biases, the way Einstein's conviction that things should look the same across reference frames shaped his work. Sanderson adds that Einstein also believed "God should not play dice," so if every LLM were Einstein you might stall quantum mechanics. There is no single correct heuristic for science, only multiple independent research programs. Implementing this, he says, feels like old-school software that amplifies entropy. The design challenge is describing the ontology of approaches. Prove-or-disprove is easy. Enumerating every possible tactic with enough breadth is hard.

Grindability, Not Just Verifiability, and the Role of Lean

Dwarkesh, stressing that he's outside the labs and offering a naive theory, argues people overemphasize verifiability. Computer use is highly verifiable (did my package arrive, is my event booked?), yet progress there has been slow. What it lacks is "grindability." Bot detectors and compute costs make it hard to run a thousand parallel rollouts of the same Amazon checkout, and cloning every website is labor-intensive. Massive parallel rollouts are needed because sample efficiency is unsolved, "sucking supervision through a straw," as Karpathy puts it. Code can be containerized and run deterministically in hundreds of copies, so credit assignment reduces to the diff between success and failure. Starting a business or trading for a day can't be replayed. Math and code are exceptions. He also thinks Lean matters less than people think for current progress, noting that the released chain of thought (or rewrite of it) for the unit distance disproof contained no Lean.

Sanderson finds the grindability point interesting and corroborates the Lean point: DeepMind's IMO effort used Lean one year and natural language the next. But he sees an underexplored benefit. Today a human still has to check a result like the unit distance counterexample, which bounds exploration. AlphaGo- and AlphaZero-style systems can explore their own universe with automated rewards and no check-ins. With Lean, one could run an endless process extending Mathlib, the GitHub repository that aims (and is very far from managing) to contain all of math in checkable code. It would likely be a fork, since the community has taste about what belongs. It might invent its own conjectures and definitions, mostly useless, and you could "look away for ten years" and ask what it found. He'd be surprised if nothing interesting came out. Math is unique in allowing this.

Dwarkesh mentions Karpathy's autoresearch setup, where agents modify a single LLM-training file and keep changes that speed up a speedrun. He also mentions Eric Jang's similar effort building a Go bot, where Jang observed agents were good at pursuing one experiment but bad at stopping at dead ends and working in extreme parallel. Dwarkesh notes that a Mathlib extender would have process supervision but no outcome, so it would likely need a supervisor model offering heuristics about usefulness. Sanderson cites a research project Terry Tao described that exhaustively searches possible axiom systems for algebras. Most collapse into nothing interesting, but occasionally an island appears that is rich in theorems and might later acquire motivation, the way group axioms look arbitrary until you realize they are about symmetry.

Dwarkesh points out that natural-language process supervision also seems to work. DeepSeek's DeepSeek Math paper describes a verifier trained by a meta-verifier. He suspects LLM-as-judge systems are behind coding agents writing cleaner code. Sanderson agrees that checking proof correctness lends itself to automated verification even in natural language, but still likes Lean's "tree of logic" for exploration disconnected from prior phrasing, invoking AlphaGo's move 37. And he names a second reason Lean matters: if AI mathematicians produce ten papers a day with any error rate, it becomes insufferable, a point he attributes to Alex Kontorovich. Even if 99 of 100 are right, finding the error is so laborious that you can't tell which are worth your time. A green checkmark guaranteeing correctness is something "every other field would kill for." So Lean may be overrated as an RL environment, but he wouldn't write it out of the story. Dwarkesh adds that an endlessly extended Mathlib is a metaphor for civilization: millennia of knowledge distilled into models that will then extend it.

Why AI Writing Lags

Dwarkesh offers two reasons writing progresses more slowly. Models are poor judges not just between A and B essays but get derailed by "B*," a bad essay that hits all the bells and whistles, inviting reward hacking. And writing isn't modular. A function or lemma can be written many ways and still work, but in writing the output is the substance, so every word matters. Sanderson asks why the progress from functional code to clean, mergeable code doesn't carry over, and whether writing has in fact improved. Dwarkesh admits his revealed preference is often to paste human writing into an LLM and ask for an explanation. Even on a live call with an expert, he'd like to pause and ask an LLM about basic concepts, saving the human for knowledge not in the distribution.

Sanderson distinguishes explanation, which models do well, from writing as insight. Good writing requires unpredictability, not just raised temperature, but knowing exactly when an unexpected move will be more insightful. The book being summarized was produced by an author who explored ideas and chose a coherent, motivated narrative, and a good author is worth reading directly.

Dwarkesh adds theory of mind. He cites a report by Andy Matuschak and a collaborator he couldn't name on teaching LLMs to write good spaced-repetition prompts. They tried RL on open-source models, chain of thought, and a large prompt to the best closed model. People talk about recursive self-improvement within a year, "and we can't get these things to write good flashcards." In his reading, the bottleneck is projecting a person's mind three months ahead. Writing similarly requires constantly asking what's happening in the reader's mind, perhaps a more diffusion-like, whole-text consideration.

Sanderson, cautioning that he may be misremembering, recalls a study of emotion recognition from faces. People tested before and after getting Botox were much worse afterward, supposedly because understanding an expression involves subconsciously mimicking it. By that lens, poor theory of mind in models isn't surprising. They know everything anyone wrote but have no face muscles and work completely differently, "like an alien trying to empathize." Humans have ready-made hardware for it.

How to Learn With LLMs

Dwarkesh finds LLMs helpful for well-known concepts but says they often become confused a few messages in, when the right human could clear things up in three minutes. Sanderson's pre-LLM principle is "who matters more than what." Pick courses by the teacher rather than your somewhat arbitrary current interests, and read more by an author you liked rather than more on a topic. He contrasts Wikipedia with the Stanford Encyclopedia of Philosophy or the Princeton Companion to Mathematics, whose single authors craft motivation and may deliberately say something slightly wrong and correct it later, which crowdsourcing edits out. LLM explanations feel like Wikipedia to him: amazing, but often most useful for the references. So he frequently asks an LLM whom to read. Once, looking for a well-visualized video on semiconductors, Claude recommended one supposedly by 3Blue1Brown. It was a real video, misattributed. Watching it beat continuing to question the model. He uses LLMs as a souped-up Google for finding the right human-written resource.

Dwarkesh agrees. His best sessions pair a human-made artifact that orders concepts and motivations with LLMs for "pruning" around it. Working through Steven Strogatz's Nonlinear Dynamics and Chaos, he split his screen three ways among Strogatz's lecture, the textbook, and an LLM, and was "bliss"-fully engaged, though he reflects that as a live student it would have gone over his head. He adds that LLMs rarely tell you that you're framing the topic wrong. They are too placating. Sanderson connects this to theory of mind: a question reveals the student's mental structure. Even good teachers find it hard to take a divergent approach seriously in the moment. The best ones "jujitsu" the student's creative framing into the lesson. He describes three levels, with LLMs at the first, good explainers at the second, and jujitsu explainers at the top, and allows that in five years LLMs might circle around to do it better.

Advice for Aspiring Mathematicians

To students asking whether to pursue math given AI progress, Sanderson says he wouldn't trust his own advice. He is a YouTuber and an outsider to academia. His general counsel, valid in 2016 or 2026, is to understand where the money comes from, what value you add, and how they connect. Students rewarded for clearing hoops often choose math as a place to keep doing that. Funding varies: a prestigious mathematician lends brand value to a university, NSF grants reflect belief in basic science as a public good and involve "a whole song and dance" of predicting progress, and teaching provides direct value. He admits he stumbled into monetizing math exploration as entertainment by accident, and could have done it more by design.

Even with near-automated theorem proving and good AI explanations, he thinks mathematicians' social role changes less than expected. Society trusts the community's judgment, prestige comes from peers, and the community's notion of valuable work may shift toward definitions or curation. In an abundant world, basic science may get more funding. He calls teaching one of the most stable post-AGI jobs, because it is relational, social, and coaching-oriented, and parents with abundant wealth will spend on it. He thinks more students should consider being math educators. Dwarkesh adds that if AIs within five or ten years produce new problems and fields, mathematics is where they'll have seen furthest, so distilling it will be in demand. And to whatever extent the math has economic use, human judgment about where to point "this behemoth of new math" becomes much more leveraged.

Will AI Math Be Good for Anything?

Finally, Dwarkesh asks whether a 10x or 100x acceleration of mathematics would matter or be bottlenecked elsewhere. Sanderson thinks it is spiky. Progress in algebraic number theory seems unlikely to unlock much. But he recalls a mathematician working on dynamics and PDEs whose group had insights that let Boeing do more in simulation instead of repeatedly disassembling and rebuilding planes after tests. Sanderson says it saved Boeing billions of dollars or something, after which Boeing funded the group. He expects incremental gains, such as faster CFD or better wing shapes, rather than, say, solving the Navier–Stokes problem and immediately unlocking massive economic value. He won't plant a flag, but he'd find it disappointing and somewhat surprising if the next five years brought no economically valuable improvements directly attributable to AI math, only a pile of toppled Erdős problems.

He closes with an uncomfortable possibility. Because AI forces people to ask what math is, one conclusion may be that much of it has become useless. If there's 10x progress and nothing visible elsewhere, people will ask why. Every grant proposal promising that elliptic curve progress would help cryptography may be exposed as possibly not true. That, he says, is one possibility.