Rebuilding AlphaGo from Scratch: What Search, Self-Play, and Supervision Reveal About RL for LLMs

Open on YouTube ↗
Overview

Eric Jang, most recently vice president of AI at 1X Technologies and earlier a senior research scientist at what is now Google DeepMind Robotics, spent part of a sabbatical rebuilding and hacking on AlphaGo with modern LLM coding tools. In this conversation with Dwarkesh Patel, Jang explains at a blackboard how AlphaGo works: the rules of Go, Monte Carlo tree search, the policy and value networks, and the self-play loop. They then use that foundation to ask why AlphaGo's style of reinforcement learning is so much more sample-efficient than the policy-gradient RL used on LLMs, and what Jang learned about automated AI research by running much of the project through Claude.

36 min read

Why rebuild AlphaGo now

Jang says AlphaGo is one of the things that drew them into AI. Their background is deep networks for robotics, where network decisions feel relatively intuitive. Go was long considered intractable for search, yet deep learning cracked it. Jang has always found it mysterious that a network of roughly ten layers can amortize the simulation of something so deep in a game tree.

There is also a practical reason. Jang points to KataGo, an open-source project released in 2020 by David Wu of Jane Street, which reduced the compute needed to train a strong Go bot from scratch by about 40x. Jang isn't sure whether KataGo is stronger than AlphaGo Zero, AlphaZero, or MuZero, but says it is very strong and is what most Go players now train against. With LLM coding assistance, Jang argues, work that once took a team of DeepMind research scientists and millions of dollars can now be done with a few thousand dollars of rented compute.

The rules, and why computers score differently

Go is simple to implement. Black moves first, and players place stones to claim territory. A stone is captured when all four orthogonal neighbors are taken by the opponent, cutting it off from "oxygen." Diagonals don't count. At the board, Jang showed how threatening a stone's last free neighbor works like check in chess: it forces a reply. They also noted that letting the opponent capture some stones can be the right choice if it sets up a bigger capture elsewhere. Jang calls this losing the battle but winning the war, and says these micro-versus-macro dynamics get richer as the board grows.

There are Chinese, Japanese, and Tromp-Taylor rule sets. Go AIs train on Tromp-Taylor because it is completely unambiguous. For example, a suicidal move that human rules forbid is legal under Tromp-Taylor and simply resolves to death immediately. The biggest difference is in scoring. Human players end a game by mutual agreement: both say the game is over and agree on which stones are dead. If they disagree, play continues. Jang frames this as two humans' "value functions" reaching consensus, an idea that returns later. Tromp-Taylor instead counts each player's stones plus the empty intersections that touch only that player's stones. Jang showed a position where a human would see that white's group is surrounded and lost, but Tromp-Taylor still credits white with those points. So under computer rules, games have to be played all the way out. A game ends when a player resigns or both players pass in succession.

Jang mentioned having had a bug in their own code around exactly this kind of corner-territory scoring.

The search problem: breadth and depth

Jang first presented the AI's view: a board encoded with at least three values (empty, black, white), and a choice among possible moves. Go gives no local reward, so you don't know which move was good until the game ends. On a 19x19 board there are about 361 legal moves early on, and games run roughly 250–300 moves. A naive tree without merging transpositions (the same position reached by different move orders) grows on the order of 361^300 nodes, far more than the number of atoms in the universe. The branching factor drops by one each move, and symmetries reduce the count, but the tree is still enormous. Jang says this is why computer scientists long doubted Go was solvable this century.

In principle Go is deterministic, so with enough compute you could search every future and always stay within the winning ones. Jang describes AlphaGo's core conceptual breakthrough as using neural networks to make that search tractable. Dwarkesh summarized the framing: there is a breadth problem and a depth problem, and AlphaGo shrinks both.

Tree nodes, UCB1, and PUCT

Before bringing in neural nets, Jang explained how to search a tree you build incrementally. The inspiration is the bandit algorithm UCB1, which isn't exactly suited to sequential games. It chooses the action that maximizes a mean value Q plus an exploration bonus. AlphaGo uses a variant called PUCT (Predicted Upper Confidence with Trees).

Each node stores a visit count, a mean action value Q, the probability of choosing that action from the parent, and a dictionary of children. Jang points out something that trips up people from robotics: nodes represent states, and since Go is deterministic, the action is implied by which child you move to. Jang says that if you ask an LLM to vibe-code MCTS it will probably design this structure correctly, and that this is what Claude 4.6 wrote for them. They call it a reasonable choice, though the layout is really up to the implementer.

Dwarkesh restated the selection rule in their own words, and Jang confirmed it. Q is the exploit term: the average win outcome over simulations beneath a node. The second term rewards actions that have been tried less than their siblings. Visit counts start at zero, so an action untried after many parent visits gets a large bonus, while an action chosen every time gets a quickly shrinking one. Because the logarithm grows slower than the count, selection shifts over time from being driven by exploration to being driven by Q. Jang says UCB was designed to bound regret when payoffs are unknown, but admits not knowing the proof or the exact bound, and suggests the PUCT form differs partly because Go has far more actions than a typical bandit problem.

Where probability and value come from

Dwarkesh asked where probability enters a deterministic game. Jang's answer: computer Go before AlphaGo already used Monte Carlo methods, averaging Q over randomly sampled trees. Q is the expected value under the distribution produced by a random search process, and that distribution is set by the per-action probabilities. With a uniform distribution over legal moves, you average over a very diffuse tree. The integral is valid but slow, because most paths are low-value and only a few matter. Jang compares it to importance sampling.

Values come from the leaves. At a terminal position you can score the game under Tromp-Taylor and get a definite win or loss. Parent nodes get the (possibly weighted) average of their children, propagated upward in what's called the backup step. Without neural networks, Jang says, this is still intractable, because you struggle to find the few actions worth sampling, especially when fighting out of a losing position.

The key intuition comes from how humans play. Strong players stop dozens or even a hundred moves before the end because they can glance at a board and judge who is winning. Jang describes this as an implicit neural network, a value function that takes a board and outputs the probability of winning, amortizing a huge number of possible playouts into a few seconds of judgment. Humans also judge at a glance which moves look promising. Those two abilities map onto the two networks AlphaGo uses: a value function that limits search depth and a policy that limits breadth.

Jang also stressed that MCTS runs from scratch on every move. The AI builds a tree from the current position, picks a move, and discards the tree. After the opponent moves, it starts over from a new root. Jang added that one thing is kept, which becomes important for training.

The network: two heads, and why ResNets still win at small scale

The value network does binary classification (win or lose), and the policy network outputs a categorical distribution over moves. Both can be trained with standard cross-entropy. Jang's setup encodes the board as a three-channel image, like RGB (black, white, and empty or masked), and feeds it into a ResNet with two heads: a single logit for value and 361 outputs for policy. The original AlphaGo used two separate networks, and later papers merged them into one network with two heads. Dwarkesh asked how much compute that saves. Jang expects the shared representations help, since policy and value should agree, but says answering it rigorously would take real work.

Jang tried both transformers and ResNets and found both work. For small data and low budgets, ResNets gave more for the money, which Jang attributes to the inductive bias of local convolutions. Jang tried hard to make transformers win, hoping they might remove the need for many of KataGo's tricks, but hasn't managed it. Jang also noted a KataGo finding: it helps to pool global features through the network, because convolutional receptive fields handle local fights well but struggle to connect a battle on one side of the board with another on the far side.

Dwarkesh asked whether games like poker or Diplomacy, where an old bluff matters now, would change the architectural choice. Jang said that Go is a perfect-information game, so a Nash equilibrium strategy exists that depends only on the current state. AlphaGo chose to rely on this, and in hindsight it worked because the equilibrium seems to be superhuman. Imperfect-information settings, such as 2v2 Go where you must model a partner's tendencies, need history and context. Jang called this an exciting research area and invited people to fork the repo.

Jang said architecture isn't very important overall, and that Karpathy-style AutoResearch hyperparameter tuning can make it good enough. What matters is having a clear optimization target.

Start from a good initialization

The original AlphaGo (AlphaGo Lee) initialized from supervised learning on expert human games. Later versions learned from scratch. Jang strongly recommends that practitioners start from something that works: "In deep learning, initialization is everything." Train the policy head on expert winners' moves and the value head on game outcomes from every position. Early positions will converge to about 0.5 because, across many games, openings split evenly, and predictions move toward 0 or 1 as games progress. This pattern appears with any training data, not only expert data.

Jang notes that the resulting raw policy, playing the argmax move with no search, is already a strong player, likely stronger than most humans. They find it remarkable that about ten layers and perhaps under 3 million parameters can do this. It also works as a checkpoint: confirm the rules are correct and simulation is fast before adding search.

The four steps of MCTS

Each move runs a fixed number of simulations, which Jang puts at somewhere between 200 and 2,048 in training. Jang believes the Lee Sedol match used tens of thousands per move on a large TPU setup. Dwarkesh remarked that this was a bit unfair to Lee, and Jang added that modern Go bots don't need much test-time compute, because training gradually pushes the work of search into the network.

Each simulation has four steps:

  1. Selection: walk down the tree by maximizing the PUCT score, Q(s,a) plus a constant times the policy prior times √N / (1 + Nₐ). At first all counts are zero, so selection favors the policy's highest-probability move.
  2. Expansion: when you reach a non-terminal node that hasn't been expanded, run the policy network on it and create its children, now from the opponent's point of view.
  3. Evaluation: score the new nodes with the value network. Because Go is zero-sum, one player's value is one minus the other's. Jang describes this as a shortcut for playing to the end.
  4. Backup: update running means of Q back up to the root.

AlphaGo Lee also averaged the value network's estimate with the result of a full rollout in which the policy network played both sides to a Tromp-Taylor finish, using α·Vθ + (1−α)·playout. Jang explains this grounded estimates in reality, especially in endgames. All later papers dropped it, and Jang did too, which sped things up considerably.

After many simulations, the tree is sparse. Most branches stop after a few visits, and visits concentrate along one line of play. This contrasts with tic-tac-toe's exhaustive 9-factorial tree. The final visit counts at the root become the vote for which move to play.

Jang posed a test question: could you drop the policy prior and just normalize the value estimates of the children? Jang says that would probably work, and the policy and value heads ought to agree anyway. Dwarkesh suggested the reason not to is avoiding 361 separate forward passes. Jang said batching would make that manageable, and the more important reason is that an explicit policy lets MCTS feed back into training, as described next.

MCTS as a policy improvement operator

After search, the visit distribution is usually sharper than the raw policy's prior. Jang calls this the MCTS-improved policy. You play a move from it (argmax or sample), discard the tree, and repeat until the game ends. The key idea is that the policy network is then trained to predict the search result directly: "why don't you just predict that from the get-go?"

Jang drew a test-time scaling picture. Win rate rises with the number of simulations. If the result of 1,000 simulations is distilled into the raw network, the next round of search starts from that higher point and goes further. The search effort is amortized into the network. Dwarkesh asked whether the diminishing returns shape holds for the distilled network too. Jang said they don't actually know how MCTS test-time scaling behaves, suspect it depends on network strength, and asked viewers not to read too much into the curve beyond it being monotonic.

Training uses every move from self-play, from both won and lost games, which Jang stresses is important. Jang compares this to DAgger from robotics and imitation learning. You may have played a trajectory that lost, but at every state search gives you a strictly better action. Retraining on those relabeled pairs improves the policy even though it doesn't guarantee a win, much like correcting a self-driving car that has drifted toward the edge of the road.

When search isn't better, and how to ground it

Dwarkesh asked whether MCTS is guaranteed to beat the policy. Jang said no, in practice it's a heuristic. Jang gave a failure mode. Suppose self-play games mostly end in resignation, so the replay buffer loses examples of late-game positions resolved under Tromp-Taylor. The value function then gets endgame positions wrong, those errors propagate up through backups and PUCT selection, and search visits a worse distribution than the policy would have chosen. With few simulations, variance adds more error. Convergence is only guaranteed as the number of simulations goes to infinity. Jang suspects this is why AlphaGo Lee kept full playouts, and suggests a fix: for about 10% of games, prohibit resignation and play to the end.

Dwarkesh summarized cold-start AlphaZero training: early on the policy is nearly useless, and training mainly teaches the value head who wins from each position. The policy starts improving once that is in place. Jang agreed and offered a tip, noting it is not peer-reviewed: make sure the value function is good before spending cycles on search, since "it doesn't really make a lot of sense to do search on garbage value predictions." Late-game positions are fairly easy to evaluate, since they approach a decidable problem. Games played to the end by reasonable players give good terminal value data, and search then backs good values into mid-game positions, which Jang calls the hard part. Dwarkesh noted the resemblance to TD learning.

For a tabula rasa start, Jang says that 50,000 random games on a small board already yield a decent value function. On 9x9, random play produces reasonable-looking final positions that are easy to score. Co-training on 9x9 and 19x19 with an architecture like one KataGo proposed transfers value knowledge between sizes. Dwarkesh pointed out that Go makes this possible because a bigger board doesn't introduce new kinds of pieces. Jang repeated that MCTS falls apart without a grounded value function.

Is AlphaGo less impressive than it looks?

Dwarkesh offered a contrarian view. The more you understand AlphaGo, the more it looks like hand-built structure (explicit tree search, tuned exploration), whereas LLM RL with verifiable rewards learns to build complex code repositories from a yes-or-no signal. Jang disagreed, said they don't understand LLM RL well enough to comment on that side, and gave their reasons for finding AlphaGo profound.

A ten-layer network can do at most about ten sequential steps of (parallel, distributed) computation, yet it approximates a nearly intractable search with high fidelity. Jang thinks most people still don't grasp how significant that is, and links it to AlphaFold and AlphaTensor, where a small network captures what looks like a massive simulation. Jang says this makes them wonder whether our understanding of computational hardness, P vs. NP, is incomplete, while stressing this is not a proof of anything. Worst-case hardness may matter less than the structure real problems have. Jang also speculates that simulating very complex systems, such as weather, might need far less compute than expected if the computation can be amortized into a forward pass.

Dwarkesh raised chaos: weather prediction gets much harder the further out you forecast. Jang said predicting the exact board state 100 moves ahead is chaotic in the same way, since one stone changes everything. Yet predicting who wins is feasible, because it is a macroscopic quantity averaged over many futures, like knowing where a hurricane goes rather than the wind speed at one point, or the overall shape of a Lorenz attractor. A hash function is equally sensitive to initial conditions but, ideally, has no such macrostructure. Jang noted that cryptography still hasn't proven fast approximations impossible, and said RSA has structure that quantum computers exploit. Dwarkesh mentioned a blog post by Reiner observing that cryptographic protocols and neural networks look alike, with layers mixing information together. Jang added that networks may be most powerful "at the edge of chaos," citing research by Jascha Sohl-Dickstein, and presented all of this as philosophy rather than expertise.

Naive self-play RL and the credit assignment problem

Jang then contrasted MCTS with a naive approach closer to current LLM RL: pit checkpoints from a league against each other and train the policy to imitate the winners. Jang set up a thought experiment. Two evenly matched policies play 100 games of 300 moves each, and policy A wins 51 to 49. Assume that in only one game did A play differently, making one better move by chance. Then there is one real supervision signal buried in about 99 × 300 moves that just reproduce the existing policy. The signal is extremely noisy. Dwarkesh connected this to Karpathy's description of RL as "sucking supervision through a straw," and noted that what would stall Go progress is the default for LLMs. Jang said it isn't that it doesn't work: with millions of games you can get meaningful signal, as long as you can mask out the neutral examples.

Jang then turned to the policy-gradient variance formula, where the return multiplies a sum of log-probabilities and the variance grows roughly quadratically with the horizon. Jang used this to explain why LLM RL typically treats a whole response as one action (T = 1): the sequence log-probability is the sum of token log-probabilities, but assigning rewards per token introduces interaction terms and a credit-assignment problem. Dwarkesh pushed back that per-step process rewards could simply be summed. Jang replied that decomposing into multiple steps introduces correlations between actions that increase variance, while the single-action version still carries the return term as a source of variance. (The episode description notes an erratum for this segment.)

The remedy in model-free RL is advantage estimation: subtract a baseline so that the multiplier is near zero for moves that didn't matter and positive only for ones that helped, pushing up actions that are better than average and down those that are worse. Jang recommended John Schulman's Generalized Advantage Estimation paper. Doing this needs a good estimate of average performance from each state, which brings back the value function. Jang's main point is that model-free RL is trying to solve credit assignment over wins, whereas MCTS isn't doing credit assignment at all. It improves the label on every action taken.

Neural fictitious self-play and the Q-learning connection

For games where tree search is hard, such as StarCraft, which lacks easy simulation and may not be deterministic, Jang described neural fictitious self-play, used in AlphaStar and OpenAI's Dota work. You freeze an opponent and train a best-response policy against it with any model-free algorithm (PPO, SAC, V-MPO), with return 1 for a win. You do this for each opponent in a league and distill the best responses into one mixed strategy that performs no worse than against an average league opponent. Jang says the model-free RL plays the role of the MCTS teacher, but the principle is the same: relabel states with better actions.

Jang then linked MCTS and Q-learning. In Q-learning, Q(s,a) is backed up as r plus a discount times the max of the next state's Q, a dynamic-programming consistency that networks can be trained to satisfy, and a policy can be recovered as the argmax over Q. Jang notes this is the "train only the value" approach Dwarkesh suggested earlier. The structural difference is that MCTS plans over trajectories the agent hasn't taken yet, while Q-learning propagates value backward over trajectories it already collected. Jang says Q-learning mattered historically because search wasn't feasible in high-dimensional problems like robotics without a good dynamics model.

Dwarkesh summarized the LLM parallel: LLM RL reinforces entire trajectories that passed a unit test, upweighting tokens whether or not they mattered, while MCTS improves each move locally because a value function truncates search before the trajectory finishes.

Why MCTS doesn't transfer easily to LLMs

Jang mentioned research from Google in 2023 or 2024 that tried applying tree structures to reasoning, and said the jury is still out. Two things make MCTS work well for Go: value estimation is concrete and can truncate depth, and the breadth of legal moves is fixed and suits PUCT's discrete exploration rule. For LLM reasoning, PUCT might be too greedy over local tokens and produce obvious but useless thoughts. The √N / (1 + Nₐ) term assumes children get revisited, but in language you will almost never sample the same continuation twice.

Dwarkesh asked whether LLMs already learn MCTS-like behavior, trying approaches and backtracking. Jang agreed they do something resembling human reasoning without an explicit tree, but thinks forward search and simulation may come back in some form. When Dwarkesh said there is "no way" to locally evaluate and improve the next move in LLM reasoning or robotics, Jang called that too strong. People are working on MCTS and MuZero-style methods for continuous control. Mathematics, with its rigid logical structure, may suit search better than something like a business negotiation.

Scaling laws, compute, and what has changed since 2017

Dwarkesh brought up Andy Jones's 2021 paper "Scaling Scaling Laws with Board Games," which showed that search compute and training compute can be traded off, anticipating LLM inference scaling. Jang also highlighted that the paper predicted the compute needed for larger board sizes, and suggested Go, which scales from 3x3 upward, would be a good place to test that again.

Jang started the project asking whether the Bitter Lesson and scaling laws would allow a strong Go bot without KataGo's tricks, and says they haven't succeeded at that yet. Early on, while they had bugs in MCTS labeling, they tried fitting scaling laws to supervised learning on expert data, but realized they might be studying scaling on bad data. Jang's advice: get a working, bug-free system first, then study it. "You don't necessarily want to jump into the science of studying your man-made artifact before your man-made artifact is interesting enough to be studied."

On cost, Dwarkesh noted that AlphaGo Zero, at 3E23 FLOPs, was a large outlier on the historical compute curve. Jang received a donation of about $10K in compute from Prime Intellect, spent about $4K on exploration and about $3K on the final run, and kept the rest for serving the model. Asked whether DeepMind was simply inefficient, Jang said that being first always costs far more than catching up, because followers can distill and bootstrap. Jang's hosted bot reached its strength through best-response training against KataGo models, and at recording time Jang was still verifying the tabula rasa version. The same pattern shows up in robotics, where frontier model compute is scattered because teams optimize for getting a capability to appear, not for FLOP efficiency, until spending reaches hundreds of millions of dollars.

Asked what actually changed since 2017, Jang offered an unreviewed "vibe guess." Architecture choices matter less. Infrastructure can be much simpler: a synchronous collect-train-collect loop instead of a distributed asynchronous system with replay buffers. Half as many desktop Blackwell GPUs can do work that needed V100s for KataGo. Some of KataGo's auxiliary objectives are unnecessary given a strong initialization, and none are needed when doing best-response training against KataGo itself. Reaching strong opponents quickly matters more than architectural novelty. Some multipliers still help: pretraining on 9x9, which Jang found useful for endgame values, cuts the roughly 30 hours AlphaGo Zero's plot shows being spent catching up to the supervised baseline. Varying the number of simulations between episodes turned out not to matter much.

Off-policy data and replay buffers

Dwarkesh asked why AlphaGo's replay buffer, where most training moves come from older models, is acceptable when researchers warn so much about off-policy training. Jang said the danger is relabeling states the current policy would never reach, which wastes capacity. In the extreme, a buffer made entirely of unreachable states would teach good moves in irrelevant positions. From the DAgger perspective, though, you want mostly on-trajectory states plus a "tube" of nearby off-trajectory states labeled with how to get back. Real environments push you off course (a gust of wind, uneven tire friction), and in games "the other player is always trying to do some shit," as Jang quoted.

Jang tried an experiment to keep the GPU fully busy. Instead of running MCTS move by move during games, they sampled random states from stored trajectories and relabeled them with MCTS under the current network, sometimes revisiting old states. It worked in practice. Jang connects this to off-policy robotics systems from the QT-Opt era at Google: a replay buffer, a "Bellman updater" that computes Q-targets by maximizing the next state's Q, and a trainer that regresses Q onto those targets, which Dwarkesh likened to daydreaming about past decisions. Jang's version replaced the Bellman updater with an MCTS relabeler that uses only the state and ignores the original action and outcome. Jang calls it moderately successful but too complex to open source. It stabilizes training when the states are reachable, but wastes capacity when they aren't. Jang observes that RL has mostly moved toward on-policy training because it is more stable, with off-policy estimates of Q used at most for variance reduction and advantage computation rather than as direct targets.

Bits per sample: why policy-gradient RL is even less efficient

Dwarkesh summarized a blog post of theirs. Learning efficiency is bits per FLOP, which equals samples per FLOP times bits per sample. Samples per FLOP fall as tasks get longer: an agent may need days of work before learning whether it succeeded. Bits per sample are also poor for RL. With a 100K-token vocabulary and an untrained model completing "The sky is…", supervised learning gets the label "blue" directly and learns −log(p) bits. RL has to guess "halycon," "told," and so on, getting almost nothing from each failure and needing on the order of 100K tries to hit "blue." Per-sample information in RL is roughly the entropy of a binary outcome, which peaks at a 50% pass rate. But training spends most of its time at very low pass rates, where on a log scale RL yields almost nothing and supervised learning yields the most. Jang added that if the policy can't sample "blue," it never gets any signal, which is why initializing to a nonzero pass rate matters.

Jang added that soft targets carry far more information per sample than one-hot labels, since the information is the entropy of the target distribution, which is why distillation is so effective. AlphaGo trains the policy on the full MCTS visit distribution rather than the chosen move. Jang suggests an experiment: train on the selected action alone to measure how much this "dark knowledge" matters.

Asked for the formal reason AlphaGo is so elegant, Jang said you never start from a 0% success rate and have to solve exploration. Every step is supervised learning: value classification plus KL minimization toward the search distribution, with no explicit TD learning or dynamic programming. Training is stable, networks can be as large as you like, and the infrastructure is simple. Because the MCTS policy stays above the raw network at each stage, the signal is always clean, unless the search distribution collapses to exactly what the policy already predicts.

What automated AI research can and can't do yet

Jang mostly used Claude Opus 4.6 and 4.7 for the project. They say the models are very good at hyperparameter optimization, and in a more open-ended way than grid or Bayesian search. The model can notice small gradients in one layer, rewrite code, or invent a data augmentation, grinding a metric like a grad student and substantially improving perplexity on a fixed dataset and time budget. Execution is also strong. Jang wrote a Claude Skill called Experiment: they describe the x-axis, the y-axis, and the question, and it runs the experiments, produces the plot, writes a report, and suggests causes.

The weakness is choosing what to do next. Jang's blog version includes a tree of all their experiments, each marked as a success, failure, or mixed result, organized into tracks, such as off-policy MCTS relabeling, that Jang eventually abandoned. Jang finds that publicly available closed models are not very good at picking the next experiment within a track, or at stepping back to say the whole track doesn't make sense and returning to first principles. Jang often had to catch infrastructure bugs by asking Claude the right question, after which it could answer. Jang allows that Mythos-class models might change this, but sees an opportunity for RL environments that reward that kind of lateral thinking.

Jang built the Go environment partly with this in mind. Go touches many research problems that overlap with LLMs and robotics but is quick to verify. The inner loop involves research engineering: distributed systems, predicting whether an idea will work, predicting how much a change will matter. The outer loop is simply checking a game result. Beyond win rate against KataGo, the goal could be predicting a bot's win rate or the scaling-law plots an idea will produce, with Go as a verification backstop against reward hacking. Jang hopes skills learned this way would transfer to biosciences, robotics, or, as Dwarkesh added, automating AI research itself.

Verifiability, stacking, and research taste

Dwarkesh raised Ilya Sutskever's view that good researchers hold strong beliefs about which ideas are right and can therefore tell a bug from a bad idea. How locally verifiable are good ideas? Jang pointed out that deep learning itself was a decades-long bet made against committees calling it a bad idea, a very long-horizon RL problem. Jang said they have no answer to how to design environments that give earlier feedback. A strong Go bot probably required discovering deep learning, so a hard-to-cheat game could in principle serve as the outer loop for such discoveries, but making it tractable depends on research taste about initialization. Jang speculates that LLMs, as a "universal grammar" able to work at any level of abstraction, might supply intermediate feedback and the lateral thinking to notice when the question itself is wrong.

On whether AI progress can be parallelized, Dwarkesh mentioned rumors that at some labs individually good ideas failed to combine and sank training runs. Jang said compute will ultimately determine everything, but with finite compute, researchers rely on heuristics that are probably partly redundant. That would explain why compute multipliers don't stack, and why they may stack even less as GPUs improve. Jang suspects their benefits are transitory, as seen with KataGo's tricks mattering less on Ada- and Blackwell-class hardware than on V100s. Knowing how much the Bitter Lesson can deliver at a given moment is a matter of taste.

Dwarkesh doubted how verifiable the outer loop really is. Before scaling laws were published, no automated process could have picked out the scaling-laws paper from any other plot, and for general AI we optimize what we can measure while caring about economically useful work that is hard to measure. Jang offered what they called a non-rigorous argument: DeepMind started with games and presumably gained transferable research skills that now help with LLMs, so automated researchers might likewise transfer from fast-verifying environments to ambitious ones like drug discovery. Dwarkesh countered that Google was long described as lagging in LLMs because it was attached to older approaches. Jang acknowledged that the games background could turn out to be a handicap, or that the apparent late start was really investment in TPUs, and concluded that even humans find it hard to reason about optimal research strategy with the data we have.

Where to go next

Jang's code is in the AutoGo repository on GitHub (username ericjang), and an interactive version of the tutorial is linked from evjang.com. Jang also recommended their essay "As Rocks May Think," about thinking as a primitive in computer science. Jang ended by encouraging people to explore the relationship between MCTS, search, and LLM reasoning. They don't claim LLMs should contain trees, but see a "very interesting duality" between the two that is underexplored because Go research has been overshadowed by LLMs, and that can be studied on very small budgets.