Why Aliens Might Build a Different Tech Stack: Michael Nielsen on How Science Actually Progresses

Open on YouTube ↗
Overview

Dwarkesh Patel opens this conversation with Michael Nielsen, a pioneer of quantum computing, co-author of the field's standard textbook, and a longtime advocate of open science, by setting aside most of Nielsen's résumé. The question they pursue is narrower and stranger: how does a scientific community recognize progress? Dwarkesh frames it partly in terms of AI, since people are now trying to "close the RL verification loop" on scientific discovery. His preparation convinced him that even in human history, the way good theories win out is more mysterious than the textbook picture suggests. Nielsen agrees there is no crank-turning method behind it. The discussion runs from the ether and the muon to Darwin, AlphaFold, alien technology, the economics of scientific credit, and finally Dwarkesh's own problem of how to learn deeply as a podcaster.

39 min read

Michelson–Morley was not a test of whether "the ether" existed

Dwarkesh asks Nielsen to retell the Michelson–Morley story as it actually happened, rather than as popular accounts present it. The popular version says an 1880s experiment disproved the ether, created a crisis, and led Einstein to special relativity. Nielsen says there is a big gap between how Michelson, Morley and their contemporaries understood the experiment and how Einstein related to it. Einstein said later in life that he was not sure he had even known of the paper at the time. Nielsen thinks there is a lot of evidence that Einstein probably did know it, but that it was not decisive for his thinking at all.

Nielsen traces the ether back to Robert Boyle in the 1600s. Sound was understood as vibrations in air, so people wondered what light was a vibration of. Boyle found that light, unlike sound, could travel through a vacuum, and introduced the ether as the medium. For roughly two centuries afterward, people debated what the ether was like. Michelson and Morley were testing competing ether theories against each other, specifically whether there was an "ether wind" through which the Earth moved. Light sent with the wind should speed up slightly and light sent against it should slow down, and interference measurements should show the difference. They found no ether wind. According to Nielsen, this ruled out some ether theories but not others, and Michelson went on believing in the ether. Michelson ran the first version around 1881. Critics including, Nielsen thinks, Rayleigh pointed out problems, so the experiment was redone in 1887. Michelson was still experimenting on the question in the 1920s and, as far as Nielsen knows, still believed in the ether shortly before his death in the late 1920s.

Dwarkesh adds that the physicist Miller kept running these experiments in the 1920s. Miller reasoned that at a high enough altitude, such as Mount Wilson in California, the Earth would no longer drag the ether along, and he claimed to have measured an effect. Dwarkesh ties Einstein's famous line "Subtle is the Lord, but malicious He is not" to this episode. Both speakers draw a lesson about falsification. It is unclear what exactly gets falsified: perhaps only one version of the ether theory. You certainly cannot induce special relativity from the failure of one ether model. Nielsen is careful to say this does not show that falsification is wrong, only that "the most naive ideas" about it are. Even the phrase "the ether" misleads, since there were many theories and several leading contenders. The leading physicists of the day took the result as information about what the ether must be like, not as proof that it did not exist.

Lorentz, Poincaré, and theories that experiments could not yet separate

Dwarkesh notes that Lorentz worked out the transformations between reference frames, the mathematical core of special relativity, before Einstein did. Lorentz, however, read them as conversions between the privileged ether frame and moving frames, with length contraction and time dilation caused by motion through the ether. Dwarkesh claims the two interpretations could not be distinguished experimentally. Nielsen calls that a strong statement. He points out that Lorentz introduced "local time" without trying to give it a physical meaning, whereas Einstein would treat it simply as time in another inertial frame. Poincaré, in Nielsen's view, came much closer to seeing it as the time clocks actually register.

Nielsen then describes the muon experiments of around 1940, possibly published in 1941. Cosmic rays striking the upper atmosphere produce showers of muons, and you can count how many survive at different altitudes. Classically they should decay long before crossing the atmosphere. They decay slowly enough to survive, and the measured rates match special relativity, as later, more precise experiments confirmed. Nielsen speculates that Lorentz, who had died about ten years earlier, might have tried to patch his theory yet again, but that it would have been a major setback. Lorentz's mathematical convenience starts to look like what time actually is, at least for muons.

Dwarkesh restates his claim more carefully. The point is not that the interpretations could never be told apart. The scientific community adopted what we now regard as the better interpretation before experiments showed it to be preferable, so some "process" must be at work. Nielsen objects to the word "process," because it suggests something fixed in advance. Lorentz, whom Einstein admired enormously, Poincaré, "one of the greatest scientists who ever lived," and Michelson never reconciled themselves to the new view. There is no standard procedure and no central authority, and great scientists can stay wrong long after the community has moved on.

Asking the right question without clinching it

Dwarkesh raises a recurring pattern: someone poses the right question but does not reach the answer. Nielsen says you have to go case by case, since people may not be going wrong in the same way each time. Poincaré is the striking case. As Nielsen understands it, while adding that he doesn't read French, Poincaré seems to have grasped both the principle of relativity and the constancy of the speed of light across inertial frames, which are essentially the premises Einstein used. Yet Poincaré treated length contraction as a dynamical effect, with particles pushed together by some force, rather than as pure kinematics reflecting a new structure of space and time. Nielsen cites a paper from around 1909 in which Poincaré still holds the dynamical picture, which is unnecessary and a mistake from the modern view. Nielsen says he does not know why. His guess is that Poincaré "almost knew too much" and had too grand a vision, and that Einstein's move was to subtract. Einstein as a teenager in the 1890s also believed in the ether, Nielsen notes, but was less attached to it than older physicists, who may have been "prisoners of their own expertise." Nielsen labels this his own guess and says some historians would disagree. Dwarkesh adds that the older Einstein is often said to have resisted the accepted interpretations of quantum mechanics and cosmology for similar reasons.

Why pick Copernicus over Ptolemy?

Dwarkesh's larger puzzle is that progress seems to outrun its verification loops. Aristarchus proposed heliocentrism in antiquity. Athenians rejected it because a moving Earth implied the stars should shift, unless they were extremely far away. Stellar parallax was not measured until 1838. And when Copernicus proposed his model, Dwarkesh says, the Ptolemaic system was more accurate, after centuries of added epicycles. In a sense it was also simpler, because Copernicus's insistence on perfect circular motion at uniform speed forced him to use more epicycles. If Copernicus was neither more accurate nor simpler, how could anyone have known in advance that he was right?

Nielsen says he does not fully know. He offers a partial answer that he, looking back centuries later, finds compelling. Newton's theory of gravitation explained Kepler's planetary motions, projectiles moving in parabolas on Earth, and the tides as effects of the Moon and Sun. Three seemingly disconnected phenomena fell out of one set of ideas, and that kind of unification feels very persuasive to him.

Newton, the last of the magicians

This leads to Keynes's essay on Newton. Nielsen reads the famous passage: Newton "was not the first of the age of reason. He was the last of the magicians, the last great mind which looked out on the visible and intellectual world with the same eyes as those who began to build our intellectual inheritance rather less than ten thousand years ago." Nielsen sees Newton as a transitional hybrid, part superstitious and part modern. Dwarkesh quotes Keynes's observation that Newton's esoteric and theological manuscripts show the same careful method and sobriety as the Principia and were written in the same 25 years.

Dwarkesh asks whether the same heuristics, such as parsimony and aesthetics, work across eras and disciplines. If they did, one could perhaps encode that taste into AI systems as a substitute for a verification loop. Nielsen's reply is one of the central claims of the conversation. The places where science gets bottlenecked are, almost by definition, the places where previous heuristics fail. Smart people study what worked before and apply it, so they don't get stuck in the same places again; they get stuck somewhere new. If you try to reduce science to a method you can crank, you will stall wherever the method does not apply, and there "definitionally, there's no crank you can turn." What you need is many people trying different ideas. The harder an idea is to have, the bigger the bottleneck and the bigger the triumph. He cites quantum mechanics as a shocking theory. He also cites evolution, where the shock is less the principle of natural selection itself than how much it explains.

Why did natural selection take until 1859?

Dwarkesh contrasts the Principia (1687) with On the Origin of Species (1859). Natural selection seems conceptually easier, and Thomas Huxley reportedly reacted with "How extremely stupid not to have thought of this," whereas no one reads Newton and feels they should have beaten him to it. Nielsen answers that animal breeders must have known large parts of the idea of artificial selection for a long time. Darwin's genius was understanding how central it was to biology, and doing the hard work of making the case. The Origin is packed with evidence and examples and connects the idea to geology and everything else it can.

Dwarkesh had wondered why Lucretius, the first-century Roman poet, didn't set off the idea earlier. After looking into it, or "more accurately, asked LLMs," he concluded Lucretius's idea was very different. It involved a past generative period followed by a one-time filter, with no ongoing gradual process and no tree of life. Dwarkesh calls universal common ancestry an incredibly weird fact. Nielsen disagrees: if the origin of life was a hard bottleneck, a single common ancestor is not surprising.

Dwarkesh also argues that verification works differently in the two cases. Newton's theory, once formulated, can rack up striking confirmations: falling bodies, planetary periods, the Moon's orbit, the tides. Darwin had to assemble cumulative evidence with no single overwhelming piece, and he did not understand the mechanism of inheritance. Dwarkesh then points to Wallace's near-identical independent discovery, with Wallace sending Darwin his manuscript and the two presenting together. Such simultaneous discovery, he suggests, shows that certain building blocks had to be in place. Candidates include Lyell's deep time in the 1830s, paleontology and intermediate fossils, and biogeography from colonial-era voyages. Nielsen agrees that deep time seems to have been crucial, and notes Darwin was heavily influenced by Lyell. On a timescale of around 6,000 years, as in Bishop Ussher's chronology, evolution would have to be visible within human lifetimes, and it isn't. Whether there were other blockers, or how much earlier a much smarter person could have found the idea, Nielsen says he does not know.

AlphaFold: a data story, and a new kind of object?

Returning to AI, Nielsen offers AlphaFold as a cautionary example. He says it "really isn't about AI" in large part. Much of the success rests on the Protein Data Bank: X-ray diffraction, NMR, cryo-EM, several billion dollars, and roughly 180,000 structures gathered over decades. The model fitted at the end was a small fraction of the total investment. It is impressive, but mostly a story of data acquisition.

Dwarkesh asks whether AlphaFold counts as a scientific explanation, comparing its 100-million-plus uninterpretable parameters with general relativity, which reduces to a few equations and predicted things it was never designed for, such as Mercury's precession. Nielsen calls this possibly a pivotal question and gives three answers. The conservative answer is that science wants deep principles with few free parameters, so AlphaFold is a useful model but not an explanation. The second is that such models may contain many small explanations that interpretability work can dig out. Nielsen doesn't know how far that has gone for AlphaFold, but mentions chess. Some experts noticed that Magnus Carlsen changed his play substantially after public analyses of how AlphaZero worked, though Nielsen stresses there is no public confirmation. The third and, to him, most interesting answer is that such models are a new type of object that should be taken seriously, with new operations we can perform on them, such as merging and distilling. He sees a precedent in how mathematicians and physicists now use Mathematica. In 1920, a 100-page equation meant giving up. Now it is a workable intermediate object that sometimes yields a simple answer at the end. "We don't have the verbs yet," he says.

Could gradient descent have found general relativity?

Dwarkesh presses on the limits. Imagine deep learning existing in 1500, trained on astronomical observations. Interpretability would likely just reveal more epicycles, with this parameter range encoding one epicycle and that range another. Nielsen suggests you could constrain models toward the simplest explanation, asking for "the 90/10 explanation" and forcing them to boil things down, with the complicated model serving as early scaffolding. Dwarkesh objects that going from Ptolemy to Copernicus requires a global swap that doesn't improve local accuracy, and he doesn't think raw gradient descent would make that move.

Nielsen describes how he understands the shift from Newtonian gravity to general relativity. Once Einstein had special relativity, a conflict was obvious. Influences can't travel faster than light, yet Newtonian gravity acts instantaneously at a distance, so it could be used for faster-than-light signaling and even sending information backward in time. That was the forcing function. You begin with the simplest fixes, they fail, and you go through increasingly complicated and wrong intermediate stages. The final theory looks simple and beautiful, but it passed through ugly steps.

Dwarkesh proposes a split. For well-understood domains needing local answers, like protein folding, trained models work. For something like general relativity, which took decades, you would need independent research programs, perhaps different AI "thinkers" seeded with different biases, like Einstein's thought experiment about gravity versus acceleration, all kept alive for a long time. Nielsen calls the point about keeping diverse programs alive "very important and central." His example is that the same move can succeed in one case and fail in another. Uranus's anomalous orbit led to the prediction and discovery of Neptune, which Dwarkesh dates to Le Verrier in 1846, a triumph for Newtonian gravity. Mercury's orbit, whose ellipse precesses 43 arcseconds per century more than Newton predicts, led to the prediction of a planet called Vulcan, which wasn't there. The real answer was general relativity. Dwarkesh adds that a committed Newtonian can keep inventing escapes: dust, a planet too small to see, bigger telescopes, a magnetic field. Nielsen gives a modern case, the Pioneer anomaly of the 1990s. The spacecraft were slightly off course, raising hopes of new gravitational physics, but the accepted explanation today is a small asymmetry in thermal radiation producing a tiny acceleration toward the Sun. Nielsen puts it as "99.9% of the time" being something like this, with the history books suffering from selection bias. There is no advance heuristic telling you which case you are in.

Prout and a hostile verification loop

Dwarkesh explains why this matters for AI. Some expect AI to advance science quickly because experiments, like unit tests in coding, provide tight verification. But infinitely many theories fit any given experiment. From Lakatos he takes the case of William Prout, who in 1815 hypothesized that atomic weights are whole numbers because elements are built from hydrogen. Most measured weights did look like whole numbers, but chlorine came out at 35.5. Defenders proposed chemical impurities, which no reaction removed, and then half-integers, but more precise measurement gave 35.46, moving away from the fraction. The eventual explanation was isotopes, which cannot be separated chemically, only physically. Dwarkesh says that for about 85 years, the verification loop was "actively hostile" to the correct theory.

Nielsen reframes the issue as one of bottlenecks. Structural biologists he has talked to regard AlphaFold as an enormous, shocking advance, so AI does help with some bottlenecks, but not necessarily all. He points to his programmer friends, who are currently in "a state of shock and high excitement." Many now seem bottlenecked on having interesting ideas, especially design ideas, which have no verification loop. When a prototype took three weeks, there was time to think up the next idea. Now a prototype takes three hours, and the design ideas that follow are not as good.

Why aliens might have a different tech stack

Dwarkesh brings up a footnote from one of Nielsen's essays suggesting that aliens might have a completely different technology stack. This contradicted Dwarkesh's unexamined assumption that science is something a civilization finishes early, after a few centuries of working out the basics, with everyone converging on the same result.

Nielsen's view is that the science and technology tree is probably much larger than we realize. People sometimes treat a "theory of everything" as the end of physics, but that isn't how it works. Computer science got its theory of everything in the 1930s, when Turing, Church and others specified what computation is, and the ninety-odd years since have been spent exploring its consequences. Public-key cryptography, a deep and very non-obvious idea, was already latent in the 1930s. He considers such discoveries science, sometimes very fundamental science. Phases of matter are another example. School taught three, or four, or five, but physicists keep adding superconductors, superfluids, Bose–Einstein condensates, quantum Hall and fractional quantum Hall systems. Nielsen expects many more, and expects we will eventually design them within the laws of physics. To him this looks like being near the bottom of the tree. In programming too, the idea that all the deep ideas have been found "just seems obviously ludicrous," since we are "slightly jumped-up chimpanzees" and slow. He recalls Knuth's preface to The Art of Computer Programming: a mathematician once told Knuth to come back when computer science had a thousand deep theorems, and Knuth remarked decades later that there clearly are now.

If the tree is that large, choices about which direction to explore matter, and different civilizations could end up in different regions. Humans are highly visual, while other animals rely more on hearing, and Nielsen wonders whether such biases shape thought, let alone in far more exotic civilizations. He stresses it is all speculation. Even with numbers, he asks what fraction of intelligences have counting, which seems natural, versus a decimal place-value system. Maybe something far better exists. Maybe some went through intermediate stages, and some use two- or three-dimensional representations rather than linear ones. "It's a lot of design freedom."

Diminishing returns and the replenished dessert table

Dwarkesh wonders whether theoretical computer science is now just filling in a taxonomy, pointing to the Complexity Zoo, or whether the real point is that entirely new fields keep appearing, as computer science would have seemed unlikely in 1880. Nielsen responds with the low-hanging-fruit argument. At a wedding dessert buffet with thirty desserts, people have fairly similar preferences, so the best go first and it gets worse from there. That picture may hold for a static snapshot of science. But if someone behind the table keeps adding new desserts, better ones may appear later. Computer science arose as a side effect of abstruse questions in the philosophy of mathematics and logic, and diminishing returns didn't apply there. New fields keep opening where a 21-year-old can make major breakthroughs instead of spending 25 years mastering prior work. Nielsen says he is not sure anyone understands why knowledge is structured this way, but empirically it seems to be.

Dwarkesh counters that deep learning, about fifteen years into its revival, already needs billions to hundreds of billions of dollars at the frontier. He offers two readings: such work is intrinsically more intensive, or our civilization is so rich that it pours in resources immediately. He also notes there seem to be few such new fields. Nielsen attributes that impression to "the architecture of attention." There is always a biggest thing. Without deep learning, we might be talking about CRISPR, and protein structure prediction might have been solved by broader curve fitting and not seen as an AI success. The centralization is partly fashion, "but there is some dynamic there."

Gains from trade between civilizations

Dwarkesh draws out an implication. If branching is so wide and path-dependent that civilizations end up with different stacks, there could be large gains from trade even in the far future, which might shape how civilizations coordinate, pushing against a "go forth and exploit" posture. Nielsen partly agrees. Ideas spread fairly easily, so what matters is whether the hard part is "almost a Dan Wang kind of idea," about capacity: the right technologies and manufacturing base, which can't easily be rebuilt elsewhere. If so, comparative advantage offers large benefits in both directions, even to a civilization that is ahead, though he expects innovation to diffuse eventually.

Nielsen's thought experiment is "GitHub but for aliens": being handed an alien civilization's algorithm specifications, which would take humans forever to mine. The idea came to him from proteins. Biology has given us an immense library of machines we barely understand. There are hundreds of millions of known proteins, and we are still working out hemoglobin and insulin despite tens of thousands of papers. Dwarkesh adds kinesin walking along microtubules and the ribosome as a miniature factory. All of it grew from one particular chemistry, perhaps one among trillions of possible seeds, and an alien civilization would find it fascinating.

Dwarkesh suggests this makes friendliness more rewarding. Nielsen says he hadn't considered that and finds it a very interesting observation, but adds that comparative advantage is a limited model. We don't trade with chimpanzees, and he thinks part of the reason is power: with a large enough imbalance, groups often, though not always, switch to domination. Dwarkesh names two further limits, transaction costs and the fact that comparative advantage doesn't guarantee terms above subsistence. His analogy, from debates about human employment after AGI, is horses. Roads suited to both horses and cars are costly. AIs thinking a thousand times faster and exchanging latent states may find a human in the supply chain more costly than beneficial. And a horse's comparative advantage doesn't justify spending roughly $100,000 a year to keep one in San Francisco.

Nielsen says his intuition differs a lot from Dwarkesh's here. He expects most of the tech tree never to be explored by anyone, so exploration choices matter. He dislikes technological determinism, which he accepts only low in the tree. He sees early institutions for steering exploration in bans and restrictions on DDT and CFCs, limits on nuclear weapons, and the Non-Proliferation Treaty. These were not decided in advance, but in some cases we are getting close to deciding preemptively not to go down a path. Dwarkesh notes that trade gains are largest for pure information, which is expensive to produce but cheap to verify and send. Today much productivity lives as process knowledge, such as that held by the people in China's manufacturing sector, but this could change if AIs do the work. Nielsen asks why 3D printers, "the next big thing for at least 20 years," still aren't central to manufacturing, and contrasts them with the ribosome, which is central to biology. If fabrication becomes uniform, whether through bioreactors or 3D printers that actually work, the problem becomes much more one of pure information.

Infinitely many deep principles, and the "ideas getting harder to find" data

Dwarkesh asks whether principles as deep as Noether's theorem or the Church–Turing principle are themselves unlimited. Nielsen says he has only speculation and instinct, and his instinct is that we keep finding very fundamental things. The Church–Turing ideas turned out to contain public-key cryptography, which in turn contained the ideas behind cryptocurrency, the ability to collectively maintain an agreed-upon ledger, whose canonical form took years to work out. That pattern of new primitives has been an important intuition pump for him.

Dwarkesh cites the paper by Nicholas Bloom and co-authors. As Dwarkesh recalls it, transistor density under Moore's law grew about 40% a year while the number of semiconductor researchers had to grow about 9% a year, with similar findings across industries. Nielsen replies that the examples are narrow, a particular field and metric. GPUs and the new parallelism they enabled, for example, don't show up. He grants that some diminishing returns seem real but asks whether they are intrinsic. The individual minds doing the work haven't changed much. Before about 1700, progress was slow and irregular; the Ionians' achievements were lost and rediscovered repeatedly. Progress required both key ideas and institutions for training, capital allocation, and even basic security for researchers from things like the Inquisition. Solve those, and there is a burst. Stagnation may just mean some external condition needs to change again. AI may be one such change. Nielsen notes that modern instruments are arguably robots already. Calling the James Webb Space Telescope one is unconventional but not unreasonable, since it is highly automated, with electronic sensors and actuators and machine-learning data processing.

After the Intelligence Age?

Dwarkesh offers what he calls a "smoke a joint" thought. Transitions like the Stone Age, the Agricultural Revolution, and the Industrial Revolution came at accelerating intervals, and no one at the start of the Industrial Revolution would have predicted AI as the next one. So perhaps the "Intelligence Age" will also be short, followed by something we lack the concepts to describe. Nielsen finds that plausible but unprovable; you can't speculate with chimpanzees about what having language would be like. His own amusing speculation is that after AGI on classical computers, there may be a distinct transition with quantum computers, which can probably perform a strictly larger class of computations. That could make "AQGI" qualitatively different.

Dwarkesh pushes back: hasn't the field tightly bounded what quantum computers can do, with modest search speedups and Shor's algorithm mainly breaking encryption? Nielsen answers that we have only thought about it for about 40 years, not very hard as a civilization, and without working devices. It might turn out narrow or radically broad. Someone in the 1700s reasoning about AND and OR gates could hardly have anticipated Bitcoin or deep learning.

Why quantum computing arrived when it did, and how Nielsen found it

Asked why quantum computing wasn't born in the 1950s, Nielsen says it could have been. John von Neumann pioneered computing and wrote an important book on quantum mechanics. The founding papers came from Feynman and Deutsch in the 1980s, with earlier partial anticipations far less comprehensive. He suggests David Deutsch would know better. Nielsen offers two historically contingent factors that matured around 1980. Computation became far more salient because people could buy an Apple II or Commodore 64. And the Paul trap made it possible to trap and manipulate single ions, single quantum states. He adds a story that Feynman, excited about one of the first PCs around 1980–81, tripped and hurt himself badly while carrying it. A talented quantum physicist who was thrilled by the new machines made it unsurprising that he was thinking about the topic then, and Nielsen doubts a similar story could have been told ten years earlier.

Dwarkesh ties this to Nielsen's idea of a "market for follow-ups," how people realize which work to build on, as with Shannon's information theory. He asks how Nielsen chose quantum information. Nielsen was 11 when Deutsch's 1985 paper appeared. In 1992 he took an excellent quantum mechanics course from Gerard Milburn, asked for papers after about the fifth lecture, and was given a large stack including Feynman's and Deutsch's, when almost no one was working on the subject. Milburn was, and in Nielsen's account wrote the first paper proposing a practical approach in a real system, though "it wasn't very practical." So Nielsen says he benefited from someone else's taste. The papers asked fundamental questions where he could make progress. He singles out Deutsch's thesis that a universal quantum Turing machine could efficiently simulate any physical system. Deutsch more or less claimed to have proved it, which Nielsen is not sure everyone would accept, pointing to open questions about simulating quantum field theory. He also mentions Deutsch's ideas on the origin of quantum algorithms and the meaning of the wave function, still unsettled among physicists. The feeling was of being in contact with something deeply important that civilization did not yet have. Dwarkesh summarizes this as a low-hanging-fruit algorithm applied to one's own interests and abilities, and Nielsen stresses how unusual the choice was in 1992.

Open science and the constructed economy of credit

Asked whether the open science movement worked, Nielsen notes that Dwarkesh didn't need to define the term, which would have been necessary 20 years ago. Most people associate it with open-access papers and often with open code and data. Making these salient issues, captured in the meme "publicly funded science should be open science," is already a large success in his view, because it engages people with the political economy of science.

He compares it with an argument from three centuries ago about whether to disclose results at all. Galileo and Kepler sometimes published discoveries as scrambled anagrams, to be unscrambled if someone else later made the same claim. It took over a century, Nielsen thinks, to reach the modern norm: disclose in a paper, expect attribution, and build careers on reputation. That fit the printing press and journals. Now code, data and in-progress ideas can be shared, but there is no established credit for them, and how much there should be is socially constructed. His illustration is the long-standing difference in preprint culture. Biologists told him biology was too competitive to post preprints, since they had to protect priority through journals. Physicists told him physics was too competitive not to post immediately, to establish priority. To Nielsen this shows that the attribution economy is something we build by agreement, and changing it changes how knowledge gets made.

Collective discovery: the LHC

Asked for the best example of collective science outside mathematics, where no individual grasps all the levels involved, Nielsen declines to rank but describes sneaking into an accelerator physics conference years ago. The people there were experts in numerical inverse methods: reconstructing, from the final shower of decaying particles at a detector, what produced it. He thought you could spend a lifetime mastering that and know little about quantum field theory, detector physics, vacuum physics, or data processing, all essential to something like the Higgs discovery. He doesn't think anyone understands all of it in the depth actually used, which is why papers have over a thousand authors who can talk at a high level but not in each other's specialties.

Prolificness versus depth

Dwarkesh says he worries about being too slow, contrasting Darwin's decades of gestation with Einstein's productive 1905. Nielsen notes Darwin wrote an enormous number of letters, so he was prolific at something. He distinguishes routine work, where you should avoid procrastination, get good, outsource, and go fast, from high-variance work, where you must spend time, go different places, and talk to people, most of which won't pay off. People tend to prefer one mode as almost a personality trait. On 1905, Nielsen says you could delete special relativity and the Nobel-winning photoelectric effect and still have a plausibly multi-Nobel year, and maybe Einstein was just smarter, with luck too. For himself, speeding up routine tasks has paid off, and so has betting more on the variance side, though that is hard for productivity-driven people because it "doesn't feel right." In San Francisco he would take a beautiful 30-minute walk to work instead of the 15-minute one, partly as a reminder that inefficiency has real benefits. He admits this doesn't answer the question and that he struggles with it.

Dwarkesh cites Dean Keith Simonton's equal-odds rule: any given work has similar odds of being important, so the most productive periods are those with the most output, as with Shakespeare, though Gödel published little. Nielsen says he has met many brilliant people obsessed with a single great project who never produce anything, which he suspects is aversion to public judgment. He wishes there were many biographies of fantastically talented people who just missed, such as IMO gold medalists who failed as mathematicians, which he suspects would be more informative than success stories.

How to actually internalize what you learn

In the final section, Dwarkesh asks for advice on his own problem. He worries that his understanding of each topic is superficial and depreciates as he moves on, and notes that many podcasters interview far more experts without growing wiser. Nielsen frames the question as how to make the context more demanding. He used to ask students, and more often himself, whether they would have put in the same effort if a million dollars had been at stake when a week's goal went unmet. The answer was invariably no.

Dwarkesh says some topics have clear preparation. For an upcoming episode on chip design, he brainstormed five roofline analyses with the guest, a founder who wrote a textbook on the subject. When he interviewed Ilya Sutskever years ago, the task was to implement the transformer. Most fields lack such an exercise, and there is no curriculum for preparing to talk to Terence Tao. Nielsen observes that Dwarkesh can do good podcasting without this understanding, so getting it means changing the structure of the output. He rejects the idea that one should always be in flow. Athletes are in flow when competing, not when training, where they are stuck or doing things badly. Dwarkesh says he doesn't know what his "64 laps" would be. He considers choosing guests with legible curricula, which raises the question of whether to drop history episodes like the one with Ada Palmer. He also considers spacing episodes out and writing 2,000 words afterward on what he learned, and says he would pay enormously for someone who could design practice problems for each topic.

Nielsen says he only truly understands something by gradually converging on a problem, and that being stuck, once just annoying, now seems perhaps the most important part. Essays he wrote in a couple of days taught him little, while some that took three months he still remembers 15 years later. There is almost always a creative artifact involved: a class, a group project, an essay, a book. He says he accepted this interview partly because Dwarkesh asks demanding questions. Dwarkesh mentions working through three lectures of Susskind's special relativity book and hiring a physicist friend to write practice problems. Nielsen says the interview is not that high-stakes, compared with, for example, writing a book meant to replace the standard textbook. He notes that "going deep" means anything from reading a couple of blog posts to writing a book, and the standard you hold yourself to matters a lot.

Dwarkesh says AI helps him move faster but he's unsure he learns better. The most demanding thinking is aversive, and there is always a next question to ask a chatbot, which is entertaining and somewhat valuable. Nielsen agrees that this partial value is what makes it seductive: it can substitute for what you should be doing, though outsourcing routine low-value work is fine. He recalls Alan Kay, as best he remembers, calling Linux "a great big ball of mud" with a few ideas worth understanding and mostly non-transferable knowledge. Learning Linux is great if you want to be a sysadmin, but less clearly so if you want to understand computing fundamentals. For a certain kind of mind, Nielsen says, it is easy to confuse learning systems with understanding. Dwarkesh promises to report back within a month with a revamped learning system, noting that learning is the main input to the podcast, so even tiny improvements are worth a great deal.