How Frontier LLMs Are Trained and Served: Reiner Pope's Blackboard Lecture

Open on YouTube ↗
Overview

Dwarkesh Patel's conversation with Reiner Pope, CEO of the chip startup MatX and previously a TPU architect at Google, was a blackboard lecture rather than a normal interview. The premise was that a few simple equations about memory, compute, and communication on a GPU cluster explain a lot about AI as it exists today: why API prices are structured as they are, why model architectures look the way they do, and why some kinds of progress have been slow. Pope worked from the physics of a single rack of chips up to guesses about how much data frontier models are trained on, and used public API prices to infer details the labs don't publish. The session closed with a discussion of the parallels between cryptography and neural networks.

28 min read

The Starting Question: What Is "Fast Mode" Buying You?

Patel opened with a concrete puzzle. Several products, including Claude Code, Codex, and Cursor, offer a "fast mode" that charges roughly 6x the price to stream tokens about 2.5x faster. He asked what is happening mechanically, whether you could pay 100x more for much higher speed, and whether a "slow mode" could make things much cheaper for users willing to wait minutes.

Pope's short answer was that the main lever is batch size. A second effect, speculative decoding or multi-token prediction, also exists, but he set it aside and did not come back to it. The rest of the first section quantified how batch size trades latency against cost.

A Roofline Model of Inference

The analysis used a "roofline" approach on a Blackwell NVL72 rack of 72 GPUs. Instead of predicting runtime exactly, Pope set lower bounds: a forward pass takes at least as long as its compute, and at least as long as its memory traffic.

Compute time is batch size times the number of active parameters, divided by the chip's FLOPs. Pope ignored the attention arithmetic because it is usually small compared with the weight matrix multiplies. He used DeepSeek V3 as an example: about 37 billion active parameters out of roughly 700 billion total. Only the active parameters count toward compute per token.

Memory time has two parts. The first is fetching the weights, and here the total parameter count matters, not the active count. The second is fetching the KV cache: for every sequence in the batch, the system reads a full context's worth of stored per-token state, at some number of bytes per token set by the model architecture. Both parts are divided by memory bandwidth.

Pope explained the KV cache using autoregressive decoding. To generate one new token, the model runs a full forward pass through every weight matrix, and the new token also attends to internal representations of all previous tokens. Those stored representations are the KV cache. He noted that attention during decode is dominated by memory fetches, not matrix multiplication.

Latency and Cost Curves as Batch Size Grows

Pope plotted each term against batch size. Compute time is a straight line through the origin. Weight-fetch time is a flat constant. KV-fetch time grows linearly with batch size. Latency is the maximum of compute time and total memory time. For small batches, latency barely depends on batch size and sits at a floor: the time to read all parameters from memory once. That floor answers part of the "pay 100x" question. On a given hardware configuration, you cannot go faster than reading every weight from memory at full bandwidth.

Patel asked whether compute always overtakes KV fetching at large batch sizes. Pope said it depends on context length. As context grows, the KV-fetch slope rises and the system moves from compute-bound to memory-bound. The context length where the slopes match is the ideal point, where you are equally limited by both. Patel asked whether doubling context from an optimal 100K to 200K would halve utilization. Pope agreed that it would under this model, but pointed out that the model assumes memory traffic grows linearly with context, which holds for dense attention. Sparse attention scales much better. He mentioned that some DeepSeek papers on sparse attention effectively put a square root into this term. He said he is excited about sparse attention but doesn't know what the labs actually use.

To turn latency into cost, Pope divided every curve by batch size, because a GPU rented for some milliseconds at some hourly rate processes B tokens in that time. Per-token compute cost becomes a constant. Per-token KV-fetch cost also becomes a constant. Per-token weight-fetch cost becomes a curve that starts near infinity at batch size one and falls toward zero as the weight reads are shared across more requests. The overall cost curve therefore drops steeply and then flattens at a floor set by compute.

This settles the slow-mode question. A hypothetical "Claude Code Slow" would sit on that floor and would not save much, because KV fetches and compute are unique to each sequence and cannot be amortized by batching more. Pope said that failing to batch can make serving economics "a thousand times worse" than batching well.

How Big a Batch Do You Need?

Pope next solved for the batch size at which weight-fetch time equals weight-compute time, leaving out KV traffic for simplicity. Rearranged, the condition puts hardware parameters on one side (FLOPs divided by memory bandwidth) and model parameters on the other (batch size times active parameters divided by total parameters). If FLOPs are counted as FP4 operations, each half a byte, the hardware ratio has no units, and Pope said it is about 300 on most GPUs. Patel asked if the ratio has shifted over time. Pope said FLOPs and bandwidth both rose a lot from A100 to H100 to B100, and the ratio stayed fairly stable.

The result is that batch size needs to be about 300 times the sparsity factor. DeepSeek activates 32 of 256 experts, a sparsity of 8, which gives roughly 2,400. Pope called this "remarkably accurate to practice." In reality, operators go two to three times higher because real efficiency falls short of the roofline, and adding KV traffic pushes the required batch higher still. He stressed that this counts sequences, not tokens: about 2,000 separate sequences each getting one more token per step. Notably, the number depends on sparsity and not on model size.

The Train That Leaves Every 20 Milliseconds

Patel asked whether requests wait for a batch to fill. Pope described it as a train schedule. The system launches a batch at fixed intervals, often around 20 milliseconds. Whoever is ready gets on. If the train is full, the rest wait for the next one. If it isn't full, it leaves anyway. The worst case is arriving just after a departure, which means waiting up to one interval and then riding for one, about 40 ms.

The 20 ms figure comes from memory capacity. A deployment wants to fill HBM with weights and KV cache and read all of it in each forward pass, so the natural step time is capacity divided by bandwidth. Pope said this has come out to around 20 ms across many HBM generations. For Rubin he estimated roughly 288 GB over 20 TB/s, about 15 ms. Taking longer than that just gives you time to read memory twice, and you don't need to read weights or KVs twice. Pope also noted that almost all HBM traffic is reads, since weights are read-only and KV cache access is mostly reads.

Converted to throughput, about 2,000 sequences per step at about 64 steps per second is roughly 128,000 tokens per second. Pope recalled announcements from last year putting Gemini's global traffic in the hundreds of millions of tokens per second, so the minimum efficient deployment is about one-thousandth of Gemini. Patel had wondered whether batching economies push toward centralization. Pope's framing was that competing at scale means serving at least that fraction of Gemini's traffic, which is a lot of traffic.

How Far Can Sparsity Go?

Because larger batches make compute the bottleneck, Patel asked how far sparsity can be pushed before model quality suffers. Pope said this can't be derived and must be measured, and pulled up the paper "Unified Scaling Laws for Routed Language Models." He noted it is somewhat old and that results depend heavily on the MoE design. MoE dates back to around 2017, with GShard and Switch Transformer as earlier approaches, and DeepSeek's version was a significant change.

In that paper's results, a 64-expert model with about 370 million active parameters matched a dense 1.3-billion-parameter model. Patel pointed out that this means about 64x more total parameters to get the equivalent of about 4x more active parameters, which is a modest return. Pope's view was that, from the bandwidth analysis alone, more sparsity is "a pure win," because the extra weight fetches are amortized by batch: raise the batch size until you run out of users. The catch is memory capacity. More total parameters and larger batches (with their KV caches) both consume HBM. Patel restated the chain: less compute leads to more sparsity, which requires bigger batches, which requires more memory capacity.

Laying a Mixture-of-Experts Layer Across a Rack

Pope drew an MoE layer. A router sends each token to a small subset of experts, perhaps 1 in 32. Each expert is an ordinary MLP with an up projection, a nonlinearity, and a down projection. The outputs are summed back together and added to the residual stream.

The standard mapping, which Pope called the best solution, is expert parallelism: different experts on different GPUs. DeepSeek's 256 experts on a 72-GPU Blackwell rack don't divide evenly, so for simplicity he used 64 GPUs with four experts each and left eight idle. The router is replicated on every GPU. Tokens go out to their chosen experts and come back, which creates an all-to-all communication pattern twice per layer. Pope said Blackwell's rack topology fits this pattern well.

The problem appears when a layer spans two racks. About half of each GPU's tokens then need to cross the rack boundary over a much slower network, and that becomes the bottleneck. In Pope's framing, one rack limits the size of an expert layer, and this has been a force behind larger and larger interconnect domains.

What a Rack Is and Why Scale-Up Domains Are Hard to Grow

Pope described a rack as a physical structure a few meters tall and a meter or two wide, usually holding about 64 accelerators, with its size limited by power delivery, weight, and cooling. In Nvidia's design, GPUs sit around the outside and NVSwitches in the middle, with cables from every GPU to the switches, so any GPU reaches any other in two hops. That is the scale-up network (NVLink). Traffic leaving the rack uses a scale-out network to data-center switches, which Pope said is typically about 8x slower. He noted that designs differ significantly between Nvidia, Google, and others including MatX.

Why not one huge switch? Pope said the hierarchy exists mainly to manage cable congestion. On the jump from Hopper (8 GPUs) to Blackwell (72), he said this was mainly a product decision to move from trays to whole racks as the form factor, not a major technical hurdle. From 64 to Rubin's roughly 500-something, he said there is "a bit of Jensen math," but also a real 4x increase that required a much more complex rack. Doubling GPUs per rack means doubling cable density, and the limits include connector density at the tray and backplane, cable bend radius, weight (more metal to prevent sagging adds more weight), power, and cooling. Pope said rack design isn't his specialty, but the people he talks to say modern racks push all of these to extreme limits.

Why Did Model Size Stall After GPT-4?

Patel observed that GPT-4, released in 2023, was rumored at over a trillion parameters, and only in the last six months or so have models clearly exceeded it. He asked whether the industry had been waiting for scale-up domains with enough memory for a multi-trillion-parameter model plus KV cache. He cited 8 Hoppers at about 640 GB versus Blackwell racks in the 10–20 TB range.

Pope agreed that larger scale-up domains were a major unlock, and said Google has had very large scale-up domains for a long time. Patel suggested this might be why Gemini seemed ahead in pre-training. Pope declined to credit that, saying he wasn't there. Higher sparsity could be part of it, but so could modeling choices, such as DeepSeek-style designs with more and finer-grained experts, and training data, and these are hard to separate. His general statement was that active parameters are limited by compute cost and total parameters are limited by scale-up size. Later in the talk he refined this by separating memory capacity from memory bandwidth, as covered below.

Pipeline Parallelism and When Crossing Racks Is Affordable

Pope said tensor parallelism has become much less relevant as experts have shrunk. The two remaining options that work across racks are data parallelism and pipeline parallelism, and he focused on the latter: put some layers on one rack and the next layers on another.

To check whether the handoff between racks becomes the bottleneck, he compared time spent on scale-up traffic with time spent on scale-out traffic. Scale-out is about 8x slower, which starts the ratio at 1/8. Inside a rack, though, each token is sent to many experts (the number of activated experts), there and back (a factor of two, which Patel caught), across several layers per pipeline stage. Crossing racks only needs one token sent once. As long as the product of those factors exceeds 8, the rack-to-rack link isn't the bottleneck. Pope called that easy to reach, since activated experts alone might be 8, and adding more layers per stage covers the rest. The result can be a chain of racks, each handling its own layers.

Patel found it notable that the best partitioning follows the model's structure: experts on GPUs, layers on racks. Pope's generalization was that a model grows along several dimensions (layers, model dimension, feed-forward dimension, expert count), and you can split along any of them. It pays off once a dimension is large enough. With current model shapes, two of the four are worth splitting.

"Pipelining Is Not Wise": Costs and Benefits

Patel brought up Ilya Sutskever's remark that "we now know" not to do pipeline parallelism, and Horace He's point that besides bubbles, pipelining limits architecture. Kimi's residuals that reach back several layers are hard to fit into a pipeline, for example. Pope agreed these costs are real and called pipelining "a massive hassle," but said the right question is what you get in return.

In inference, pipelining saves neither memory time nor compute time. It only moves the work to other chips. It does reduce the memory capacity needed per rack. Pope drew the timeline: with four racks, a single inference moving through the stages would leave each rack idle three-quarters of the time (the bubble), so you start the next inference right away and fill the pipeline. In inference this is natural, latency is unchanged compared with running everything on one rack, and the distinction between micro-batch and batch doesn't matter.

Training differs because there is a fixed batch that must be fully accumulated before the backward pass. Pope said that from a convergence standpoint smaller batches are always better, since each gradient step uses fresher information, while from a systems standpoint smaller batches are worse. The optimum balances the two. The forced stop between forward and backward creates bubbles, and there are more complex schedules, such as zero-bubble and one-forward-one-backward, that interleave passes. When Patel joked about mining Bitcoin in the bubble, Pope said the more useful option is computing weight gradients there.

Even in inference, Pope said pipelining isn't used much, because a Blackwell rack already has far more memory than a trillion-parameter model needs (about a terabyte). That led to a hardware point: if weights don't have to fit in one rack, you could build chips with less HBM. Patel contrasted this with the "memory wall," citing Dylan's claim that hyperscalers are spending around 50% of this year's CapEx on memory. Pope called that believable.

Why Pipelining Doesn't Help With the KV Cache

To address the memory question, Pope wrote per-GPU memory demand: total parameters plus KV cache (batch × context length × bytes per token), divided by the expert-parallel degree E and the number of pipeline stages P. Decode runs in a loop, and keeping every stage busy requires as many micro-batches in flight as there are stages. So the global batch is P times the micro-batch size, and the micro-batch size is the "300 × sparsity" figure.

Substituting that in, the weight term still shrinks as 1/(E·P), but the P cancels in the KV term. More pipeline stages keep reducing per-GPU weight memory, but per-GPU KV memory stays the same. After a little pipelining, sometimes only two stages, the KV cache dominates. Patel asked why storing fewer layers of KV per rack doesn't help. Pope said it does, but it is exactly cancelled by needing more sequences in flight to keep all racks busy. Patel summed it up: the KV cache can't be amortized over the batch and can't be sharded away through pipelining.

Pope said DeepSeek's paper reflects this in practice: use expert parallelism up to the size of the scale-up domain and very little pipelining, maybe none or two stages, just enough to manage weight storage. He said frontier inference therefore mostly runs within one scale-up domain, with some pipelining justified for a very large, very sparse model that exceeds one rack.

Latency and Bandwidth: The Real Case for Bigger Scale-Up Domains

Since pipelining solves weight capacity, Pope explained why bigger scale-up domains still matter. One reason is inter-rack hop latency, which he guessed is on the order of a few milliseconds but said could be wrong by an order of magnitude. The path runs from GPU to network card to top-of-rack switch, possibly through a data-center switch, and back out, and in decode these delays add up across stages. Patel estimated this could push per-token time from about 20 ms to about 30 ms.

The bigger reason is memory bandwidth. The weight-fetch time that sets the latency floor is total parameters divided by the aggregate bandwidth of the GPUs that can load weights at the same time, which means the GPUs in one scale-up domain, not across pipeline stages running at different moments. Aggregate bandwidth equals scale-up size times per-GPU bandwidth. Per-GPU bandwidth improves maybe 1.5–2x per generation, while scale-up size grew 8x from Hopper. Pope's summary: pipelining solves capacity, scale-up size solves bandwidth, and bandwidth mainly buys lower latency. He said a very sparse model on a small H100 box would have very high latency.

How Far Beyond Chinchilla Are Models Trained?

Patel asked how much frontier models are over-trained relative to Chinchilla-optimal, now that pre-training cost is spread across RL rollouts and user inference. Pope said labs don't publish updated scaling laws or traffic numbers, so this requires guessing.

He started with a heuristic that he stated as a conjecture: when minimizing a sum of costs where one falls and another rises, the minimum is often close to where they are equal, as with 1/x and x, and in many power-law setups. He applied this to three costs. Pre-training is about 6 × active parameters × tokens. RL is between 2 and 6 per parameter per token, depending on how many rollouts are trained on, and is further penalized because decode-heavy work runs at lower utilization. Inference is 2 × active parameters × tokens. Active parameters cancel. Patel suggested splitting 50% training and 50% inference instead of thirds, and Pope said that was equally valid because the heuristic can't tell them apart and the difference is small.

After some back-and-forth algebra (Pope reversed a factor at one point and joked about billions of dollars flowing the wrong way), the result was that inference tokens, pre-training tokens, and RL tokens should all be roughly equal within factors they couldn't pin down, with somewhat fewer RL tokens because RL is less efficient per token. Patel put it this way: all users of a model should, combined, produce about as many tokens as the model was pre-trained on, "the sum of human knowledge." Pope added that uncertainty, including the chance a model isn't frontier and gets retired early, should reduce the expected inference tokens.

Then the numbers. Pope assumed perhaps 500 million tokens per second worldwide (he said he didn't really know), over a two-month deployment, which gives about 2.6 × 10^15. He divided by 5–10 to account for multiple models in a family, getting about 50 million tokens per second per model, or around 200 trillion tokens over its life. Patel mentioned that someone had told him a frontier model was pre-trained on 150 trillion tokens, which roughly matches. Pope guessed, admitting it was poorly sourced, about 100 billion active parameters. Chinchilla's ~20 tokens per parameter would then suggest about 2 trillion tokens, so frontier models would be about 100x over-trained. Pope emphasized the large error bars but called it empowering to "just set A equal to B."

Reading Architecture Out of API Prices

Long-context surcharge. Patel noted that Gemini 3.1 charges 50% more above 200K tokens. Pope plotted per-token cost against context length. Compute is almost flat. Memory cost starts at the weight-fetch baseline and rises with context. Where they cross, costs start to grow, and a two-tier price schedule stays profitable across the whole range. He assumed the 200K threshold roughly matches that crossover, took batch size large enough that weight fetches are negligible, and solved for KV bytes per token. The inputs were the hardware ratio of 1/300, a guessed 100 billion active parameters, and 200K context. The answer was about 1,667 bytes, roughly 2 KB per token.

Pope found that plausible, perhaps slightly small. One way to reach it is dense attention with heavy KV sharing across layers, as in Character AI's published approach (alternating long and short context, with global context shared across layers) that also appeared in Gemma models. KV bytes are roughly unique KV layers × 2 × d_head × number of KV heads. With one shared layer, d_head of 128, and 8 KV heads (typical values run from 1 to 8), you get about 2 KB. Another path is sparse attention with larger dimensions offset by a 1/sparsity factor. Pope's explanation for why prices reveal this much: providers have an incentive to price close to cost, or someone will undercut them.

Output vs. input prices. Pope recalled output tokens costing 3–5x more than input tokens. Prefill processes many tokens in one pass. Its compute scales with the number of tokens, but the memory traffic (weights plus existing KV cache) is roughly the same as a decode step, and the new tokens' attention can stay on-chip with techniques like flash attention. Dividing by pass length, per-token memory cost drops sharply for prefill while per-token compute stays constant. Patel's conclusion, which Pope confirmed: a roughly 5x premium for output means decode is heavily memory-bandwidth-bound, and the gap between memory and compute time in decode is about that large.

Why Context Windows Have Stopped Growing

Patel asked about the quadratic cost of attention. Pope said the attention FLOPs per token grow linearly with context, but with such a small slope that they only matter around millions of tokens. The real limits on long context are memory bandwidth and memory capacity. Patel cited Dario Amodei's argument that in-context learning might replace continual learning, and said that would need something like 100 million tokens of context to match a month of working with a colleague.

Pope pointed out that context lengths jumped from about 8K to 100–200K between GPT-3-era and GPT-4-era models and have held there for a year or two. He reads that as the cost-balanced point, with going far beyond it prohibitive because of memory bandwidth, not compute. He said he doesn't see a good way around it, since HBM isn't improving dramatically. Sparse attention is a big improvement and may already be reflected in current pricing, but it has limits: too sparse, and the model attends to too few tokens and quality falls. His conclusion was that there is no solution to the memory wall here.

Prompt-Cache Pricing and Memory Tiers

Cache hits cost about a tenth of normal input, and writing to cache costs more. Pope framed this as two ways to get a token's KV cache: recompute it from token IDs ("rematerialization"), or keep it stored somewhere. He listed tiers (HBM, host DDR, flash, and later spinning disk), each with a cost to hold data over time and a cost to retrieve it. Rematerialization costs nothing to hold but a full forward pass to retrieve. HBM costs the most to hold, because a GPU with HBM full of idle KV caches can't do other work. Slower tiers are cheaper to hold but cost more to retrieve, since data must be copied back to HBM. Pope said Nvidia deploys systems with both DDR and flash.

Patel read out the pricing: base input $5 per million tokens (the rematerialization price), a five-minute cache write at $6.25, plus a separate one-hour option. Pope proposed that the right tier for a given hold time is the one whose "drain time" (capacity divided by bandwidth) roughly equals that hold time, because that balances the fraction of the device used for holding against the fraction used for retrieval. HBM drains in about 20 ms, far too short. He estimated DDR at about 1–10 seconds, flash around a minute, and spinning disk around an hour, noting he didn't have exact figures memorized. On that basis he suggested the five-minute and one-hour tiers may correspond to flash and spinning disk, not HBM and DDR as Patel first guessed. He was surprised spinning disk would be used at all, calling it an unattractive technology that is still useful in some places.

Convergent Evolution: Ciphers and Neural Networks

In a seated segment, Patel raised Pope's blog post on how cryptographic constructions and neural networks both depend on thorough mixing of information across inputs, despite opposite goals. Ciphers make structured data look random, and neural nets pull structure out of data that looks random. Pope compared it to stirring cake batter one way and then another.

He located the difference in gradient descent. A randomly initialized network might work as a reasonable cipher, but training makes it interpretable, and architects work to keep derivatives simple with residual connections and LayerNorm. One major attack on ciphers is also differentiation, done over binary values: differential cryptanalysis. A good cipher makes small input changes produce large output changes. Patel brought up backdoors. Pope pointed to adversarial examples for image classifiers as the clearest overlap: a tiny perturbation that completely changes the output is the "avalanche" property ciphers want and neural nets don't.

On using neural nets as ciphers, Pope was skeptical. A new cipher without about ten years of scrutiny is probably broken, and he estimated 99% are. The transfer in the other direction has worked. The Feistel construction takes a non-invertible function f (for example an MLP) and builds an invertible layer: from inputs (x, y), output x alongside y + f(x). Then y can be recovered as the second output minus f(x). The RevNets paper (Pope gave its date as 2017 at one point and 2018–2019 at another) applies this to network layers, which amounts to a residual connection reaching back two layers. Because the whole network becomes invertible, activations don't have to be stored for the backward pass. They can be rebuilt by running the network backward during backpropagation, which removes what can be the largest memory cost in training.

Patel noted that this trades compute for memory, the opposite of the KV cache, which trades memory for compute. Pope's last point was that with current hardware, spending memory to save compute is generally the more profitable direction.