Dylan Patel on Why Chips, Not Power, Will Cap AI Compute This Decade

Open on YouTube ↗
Overview

In this long conversation, Dwarkesh Patel asks SemiAnalysis CEO Dylan Patel a series of questions about the AI compute buildout: what the hundreds of billions in hyperscaler CapEx is actually buying, why some labs are scrambling for capacity, and what will ultimately limit how much AI compute the world can deploy. Patel's position is that the binding constraint has moved back from power and data centers to semiconductor manufacturing itself. By the end of the decade, they argue, the lowest rung of that supply chain, ASML's EUV lithography tools, sets a hard ceiling. Power, labor, and permitting, by contrast, are problems Patel expects capitalism to solve.

31 min read

What $600 Billion of CapEx Is Buying

Dwarkesh opens with a puzzle. The combined forecast CapEx of Amazon, Meta, Google, and Microsoft for this year is about $600 billion. At current rental prices, that would pay for something like 50 gigawatts of compute, and nobody is bringing 50 gigawatts online this year. Meanwhile OpenAI just raised $110 billion and Anthropic $30 billion, and renting a gigawatt costs roughly $10–13 billion a year. So what is all that money for?

Patel's answer is that much of the spending is setup for future years. Across the whole supply chain the figure approaches a trillion dollars. Part of it pays for chips going online this year, but a large share is advance spending. Of Google's roughly $180 billion, Patel says, a big chunk goes to turbine deposits for 2028 and 2029, data center construction for 2027, and power purchase agreements and down payments, all so the company can scale quickly later. Conversely, some of the roughly 20 gigawatts being added in the US this year was paid for in prior years.

The biggest customers of all these providers are Anthropic and OpenAI, which Patel puts at roughly two to two-and-a-half gigawatts right now. They draw a straight line through Anthropic's recent growth. The company added $4 billion of revenue in January and $6 billion in February. At $6 billion a month, it would add about $60 billion over the next ten months, and Patel notes some would call that assumption bearish. At the gross margins last reported by the media, that implies about $40 billion of inference compute. At roughly $10 billion per gigawatt in rental cost, that means about four more gigawatts just to serve new revenue, even if the research and training fleet stays flat. Anthropic therefore needs to get well above five gigawatts by year end. Patel calls this "really tough" but possible.

Anthropic's Conservatism and the Cost of Buying Compute Late

Patel is blunt about the contrast between the two leading labs. Dario Amodei had said on Dwarkesh's podcast that Anthropic would not "go crazy" on compute, to avoid bankruptcy if revenue inflected differently than expected. Patel's view is that this approach has left Anthropic worse off than OpenAI, which signed aggressive deals and will have considerably more compute by the end of the year.

Buying compute in a pinch, Patel explains, means going to providers Anthropic would not previously have used. Historically Anthropic worked with the highest-quality providers, Google and Amazon. OpenAI spread itself widely: Microsoft, Google, Amazon, CoreWeave, Oracle, and less obvious names such as SoftBank Energy, which had never built a data center, and Nscale. Patel recalls that in the second half of last year the financial "freakouts" around Oracle and CoreWeave, and the turmoil in credit markets, came from doubts that OpenAI could pay for its commitments. For about four months, Patel says, counterparties refused to sign with OpenAI. After its raises, the attitude flipped to "we believed you the whole time."

Where does last-minute capacity come from? Not every compute contract is a five-year deal. H100s from 2023–2025 were often signed for six months to three years, and as those contracts roll off they go to whoever pays most. Patel says H100 prices have "inflected a lot." They have seen AI labs sign two-to-three-year deals at as much as $2.40 an hour, against a build cost of about $1.40 an hour over five years. Neoclouds held a higher share of Hopper and tended to sign shorter deals, so that capacity is available. Some neoclouds and hyperscalers also have unsold capacity under construction, or capacity earmarked for less AGI-focused internal uses that they may now sell. Finally, Anthropic doesn't need to own all its compute: Amazon Bedrock, Google Vertex, or Microsoft Foundry can serve Claude under a revenue share.

Dwarkesh summarizes this as Anthropic paying either an effective markup through revenue share or spot premiums it would have avoided by committing early. Patel agrees it is a trade-off. They add that there are not yet many other large incremental buyers, because Anthropic reached the capability tier where revenue is "mooning" first. Patel's estimate is that Anthropic will reach five to six gigawatts by year end, counting capacity serving its models through the cloud platforms, well above its original plans. They put OpenAI at roughly the same level, slightly higher by SemiAnalysis's numbers.

Dwarkesh also clarifies what he was trying to say to Amodei. He was not claiming the singularity is two years away. His point was that Amodei's own stated timelines, a "data center of geniuses" in about two years and no more than five, sit uneasily with a conservative compute posture. Patel adds that a meme has formed about Anthropic having "commitment issues."

Why an H100 Is Worth More Now Than Three Years Ago

The discussion turns to GPU depreciation. Bears such as Michael Burry have argued GPUs should be depreciated over three years or less, which would make cloud economics look much worse. Patel offers two lenses.

The first is a mechanical total-cost-of-ownership model. It adds up data center, networking, on-site staff, spare parts, chip, and server costs, plus depreciation and financing. By this model an H100 costs about $1.40 an hour to deploy at volume over five years, so a five-year deal at around $1.90–2.00 yields roughly a 35% gross margin. The bear case says that because Nvidia triples or quadruples performance every two years while raising prices only 50% to 2x, an H100 worth $2 in 2024 falls to about $1 an hour once Blackwell ships in the millions in 2026, and to about $0.70 once Rubin reaches high volume in 2027.

The second lens is the value you can extract from the chip. That lens dominates, Patel argues, because supply is constrained. With infinite new chips, older ones would be priced against newer ones. With chips scarce, they are priced by what they can produce today. Patel's example: GPT-5.4 is a sparser mixture-of-experts model with fewer active parameters than GPT-4, cheaper to serve, and much better thanks to advances in training, RL, architecture, and data. An H100 therefore produces more tokens of a far more valuable model. Patel guesses GPT-4's maximum token market was a few billion to tens of billions of dollars, and GPT-5.4's probably above $100 billion, subject to adoption lag and competition. Hence the "crazy" conclusion that an H100 is worth more today than three years ago. Dwarkesh extends the point: if an H100, which some estimate at about the brain's 1e15 FLOPS, could run something like a human knowledge worker earning six figures, it would pay for itself in months.

Dwarkesh raises the Alchian-Allen effect: a fixed cost added to both a high- and low-quality good shrinks their price ratio and pushes buyers toward quality. Applied to AI, if a Hopper rises from $2 to $3 an hour and can produce one million tokens of Opus or two million of Sonnet, the relative gap between the two models narrows. Patel agrees and notes that volumes and revenue are already concentrated on the best models.

Who Captures the Margin

In a compute-limited world, Patel says, companies that locked in five-year contracts have a large margin advantage over those buying at today's value-based prices. Long-term contracts also make up far more of the market than flexible short-term capacity. But because the fleet grows so fast, most spending goes to newly added compute priced at new rates. CoreWeave's average contract term is over three years for 98%+ of its compute, so it can't reprice existing capacity, but it adds large amounts each year. Patel says Meta alone is adding this year as much capacity as its entire 2022 fleet serving WhatsApp, Instagram, and Facebook. OpenAI, in the example Patel gives, went from 600 megawatts to two gigawatts last year, to six-plus this year, and perhaps twelve next year.

Upstream, the leverage sits with whoever secured memory and logic capacity: mostly Nvidia, which Patel says holds about $90 billion in long-term contracts and is negotiating three-year deals with memory vendors, plus Google and Amazon via Broadcom, Amazon directly, and AMD. TSMC is not raising prices. Memory vendors are, and Patel expects them to double or triple prices again while signing long-term deals. This year, Patel expects model vendors' margins to rise a lot too, because they must "destroy demand." There is no way, Patel says, for Anthropic to keep its current pace otherwise.

How Nvidia Locked Up TSMC and Google Got Squeezed

Dwarkesh asks why TSMC would let Nvidia hold something like 70% of N3 capacity by 2027 instead of fragmenting the market the way Nvidia fragments the neocloud market. Patel's explanation has several parts. Last year most N3 went to Apple, which is moving to N2 and may cut volumes as memory costs rise. TSMC earns higher margins on high-performance computing than on mobile. It also actually prefers CPU customers, such as Amazon's Graviton or AMD's CPUs, over AI accelerators like Trainium, because it sees CPU demand as more stable. A conservative company, Patel says, allocates to steady markets first.

Nvidia still got the majority because it sent the market signal earliest: non-cancelable, non-returnable orders, sometimes with deposits, while Google and Amazon hit stumbling blocks such as chip delays of a couple of quarters. TSMC then checked the rest of the supply chain, including PCB suppliers like China's Victory Giant and memory vendors, and found Nvidia had capacity lined up. Patel doesn't think Jensen Huang is "AGI-pilled" like Amodei or Sam Altman, but says Nvidia was far more so than Google or Amazon in Q3 of last year, partly because it could see the data center construction. As a result, Google has to deploy many GPUs even though TPUs suit it better, because it cannot get enough TPUs fabbed.

On why Google sold about a million TPUv7 (Ironwood) chips to Anthropic instead of keeping them for DeepMind, Patel offers what they call a narrative they have "spun" themselves from supply chain data, not confirmed fact. Anthropic's compute leads both came from Google and spotted the dislocation. Over six weeks in early Q3, SemiAnalysis saw TPU capacity requests rise several times, so abruptly that Google had to explain the increase to TSMC. Much of it was for Anthropic. Then Nano Banana and Gemini 3 sent Google's user metrics up, leadership "woke up" and talked about doubling compute every six months, and TSMC replied that it could offer perhaps 5–10% more for 2026 and would work on 2027. Patel points to Gemini's ARR, near zero through most of the year and about $5 billion on an ARR basis by Q4, as the reason Google hesitated. Since then, Patel says, Google has become "absurdly AGI-pilled": buying an energy company, putting down turbine deposits, and buying powered land.

The Bottleneck Moves to ASML

Patel traces the shifting bottlenecks: CoWoS packaging, then power, then data centers. All were short-lead-time items. Data centers take under a year (Amazon has built one in eight months), while fabs take two to three years and tools have long lead times. Until now, capacity could also slide from mobile and PC chips to AI. That reserve is exhausted: Nvidia is now the largest customer at both TSMC and SK Hynix. By 2028–2029, Patel expects the constraint to fall to ASML.

The arithmetic runs as follows. ASML can make about 70 EUV tools now, 80 next year, and, even under aggressive expansion, just over 100 a year by decade's end, at $300–400 million each. A gigawatt of Nvidia's Rubin needs about 55,000 N3 wafers, 6,000 N5 wafers, and 170,000 DRAM wafers. A leading-edge 3 nm wafer has about 70 lithography layers, around 20 of them EUV, so N3 alone takes about 1.1 million EUV passes per gigawatt, and roughly 2 million once N5 and memory are included. At about 75 wafers an hour and 90% uptime, that works out to about 3.5 EUV tools per gigawatt. Patel points out the contrast: a gigawatt costs roughly $50 billion, while 3.5 EUV tools cost about $1.2 billion.

With about 250–300 EUV tools already installed and the planned additions, Patel gets to roughly 700 tools by 2030. If all were devoted to AI, which they won't be, that supports about 200 gigawatts of AI chips a year. Altman's stated goal of a gigawatt a week, about 52 GW a year, would then be roughly 25% of the total, which Patel finds reasonable given that OpenAI will have access to about a quarter of this year's Blackwell deployments.

Dwarkesh is surprised that 2030 will still depend on tools first shipped around 2020. Patel explains that EUV entered high-volume production around 2020, but the tools have improved in throughput and in overlay, the precision with which successive layers align. Prices rose from about $150 million to an expected $400 million by 2028 while capabilities more than doubled. Patel calls ASML perhaps "one of the most generous companies in the world": despite having no competitor, it has never raised prices faster than capability, unlike Nvidia or the memory makers. For simplicity, the tool-count math ignores these per-tool gains.

Why ASML Can't Just Build More

Patel attributes the slow ramp partly to mindset. The semiconductor supply chain has lived through booms and busts and doesn't believe in demand for 200 GW a year. SemiAnalysis is constantly told its numbers are too high. The rest is sheer complexity. The tool has four major subsystems. The source, made by ASML-owned Cymer in San Diego, hits tin droplets three times with a laser to produce 13.5 nm light. The reticle stage is made in Wilmington, Connecticut; the wafer stage and the optics are made in Europe. Carl Zeiss's optics are multilayer mirrors, about 18 per tool, where any deposition defect or curvature error ruins them. Production is "artisanal," on the order of a thousand per year. The reticle and wafer stages accelerate at about nine Gs in opposite directions while scanning 26×33 mm fields, with overlay held to about 3 nm, which requires sub-nanometer accuracy in each component. Tools are built in Eindhoven, disassembled, flown to customers on many planes, reassembled, and retested over months. ASML says its supply chain includes over 10,000 suppliers.

Patel contrasts this with US power, where moving from 0% to 2% growth was hard despite a relatively simple supply chain with perhaps 100,000 workers. Zeiss probably has fewer than a thousand people on this work, all highly specialized. Everyone downstream of the labs builds "X − 1," sometimes "X ÷ 2," and the whip takes a long time to crack. Adding up Elon Musk's 100 GW a year in space, Altman's 52 GW, and similar needs from Anthropic and Google, Patel says the supply chain cannot possibly satisfy everyone.

Why Not Fall Back to 7 nm?

Dwarkesh suggests the semiconductor equivalent of behind-the-meter power. Holding FP16 constant, the A100 (312 TFLOPS) to B100 (a little over one petaflop) gain is only about 3x, and much of it came from design rather than process. So why not build on 7 nm with DUV multi-patterning, as China does, and accept a haircut?

Patel says it could happen if demand gets desperate enough, but the comparison misleads. Each generation targets different numerics: A100 for FP16/BF16, Hopper for FP8, Rubin for FP4/FP6. More importantly, models run across many chips. DeepSeek's production deployment, over a year old, used 160 GPUs. Bandwidth falls roughly an order of magnitude at each boundary: tens to hundreds of TB/s on-chip, around a terabyte a second between chips in a rack, around 100 GB/s between racks. Serving DeepSeek or Kimi K2.5, both eight-bit models, at 100 tokens a second, Patel says Blackwell outperforms Hopper by about 20x despite both being on essentially the same process. The gain comes from networking and system design, not FLOPS. Some of these improvements can be ported back to 7 nm and some cannot, and the gaps compound across FLOPS, networking, and memory bandwidth.

On putting more dies in a package: it helps, and Patel cites Tesla's wafer-sized Dojo with 25 chips, which they call probably still the best chip for CNNs but ill-suited to transformers and unable to use HBM. Huawei's Ascend 910C went from one to two dies. But larger packages bring networking, memory bandwidth, and cooling limits, and anything done in packaging on 7 nm can also be done on 3 nm.

When Could China Outscale the West?

Asked when China's scale might overtake the West's process lead, Patel notes that China's 7 nm and 14 nm capacity still relies on ASML DUV tools, and the scale advantage remains with the West plus Taiwan, Japan, and Korea. Patel is "quite bullish" that China will scale over five to ten years. They expect fully indigenized DUV by 2030 "for sure," and working EUV tools, but not volume production, because of "production hell": ASML had EUV working in the early 2010s and needed another five to seven years to reach mass production. On Chinese DUV output in 2030, Patel calls any estimate "a shot in the dark," but thinks about 100 tools a year is plausible, versus ASML's hundreds.

Patel frames the outcome as depending on timelines. Frontier US models such as Opus 4.6 and GPT-5.4 have recently widened the gap. Distillation will get harder as labs sell completed work, not visible reasoning chains. And US labs are scaling compute far faster: OpenAI exited last year around two gigawatts, and Patel expects both OpenAI and Anthropic at about ten gigawatts by the end of next year. Anthropic's $20 billion ARR at sub-50% margins implies $13–14 billion of rental compute, or about $50 billion of CapEx someone laid out, and China has not built that. If returns stay high, the US and China diverge. If AI infrastructure yields middling returns, if Google is wrong to push free cash flow toward zero for $300 billion of CapEx next year, then China's vertically integrated supply chain could let it scale past a fragmented Western coalition. Dwarkesh summarizes: fast timelines, the US wins; long timelines, China wins. Patel adds that one needn't believe in AGI to be in the US-wins scenario.

The Memory Crunch

Dwarkesh asks whether accelerators could use commodity DRAM instead of HBM, which yields three to four times fewer bits per wafer, and serve slower agentic workloads, a "Claude Slow." Patel argues the highest-paying customers are the least price-sensitive and want speed. Anthropic could already offer a slow mode, perhaps cutting Opus 4.6's price 4–5x for about a 2x speed loss on existing HBM, but doesn't because nobody wants slow models. Agentic tasks that take hours would take a day.

Technically, the relevant metric is bandwidth per wafer, not bits per wafer. I/O escapes only from the chip's edges, the "shoreline." An HBM4 stack is 2,048 bits wide across about 11–13 mm at about 10 GT/s, roughly 2.5 TB/s. DDR5 in the same edge space is 64–128 bits at up to about 8 GT/s, roughly 64–128 GB/s, an order of magnitude less. Switching to DDR would leave FLOPS idle waiting on memory. Patel notes that SemiAnalysis's open-source InferenceX searches this design space across chips and models.

Patel confirms that about 30% of Big Tech's 2026 CapEx goes to memory, before accounting for Nvidia's margin stacking. The consequence is that consumer devices get worse and people "hate AI more." An iPhone has about 12 GB of DRAM. At $3–4 per GB that was about $50; with prices tripled to around $12 per GB it's about $150. Including NAND, Patel estimates a $150 cost increase and roughly $250 more for consumers once Apple's margin is applied, though long-term contracts delay the effect until the next iPhone. The low end is hit far harder: bigger memory share of the bill of materials, thinner margins, fewer long-term agreements. Smartphone sales, once 1.4 billion a year and now about 1.1 billion, could fall to 800 million this year and 500–600 million next, by SemiAnalysis projections. Its analysts in Asia see Xiaomi and Oppo cutting low- and mid-range volumes by half. Because consumer devices are over half of memory demand, this frees DRAM for AI. NAND prices are rising too, but less, because destroyed consumer demand frees proportionally more NAND.

Why not just make more memory? For now, Patel says, the constraint is fabs, not EUV. Memory makers lost money in 2023 and stopped building. SemiAnalysis argued from 2024 that reasoning means long context, long context means a large KV cache, and that means memory demand. It took about a year to show in prices, another three to six months for vendors to start building, and fabs take two years, so meaningful new space arrives in late 2027 or 2028. Meanwhile Micron bought a lagging-edge fab in Taiwan, and Hynix and Samsung are squeezing existing fabs. EUV is 28% of N3 wafer cost but in the teens for DRAM. Other toolmakers such as Applied Materials and Lam are expanding, but there is nowhere to put the tools.

Clean Rooms, TeraFab, and Wild Cards

On Musk's TeraFab plan, Patel thinks Musk can recruit great people with an audacious goal, a million wafers a month, far beyond any existing fab, and can probably build the clean room in a year or two. Patel is "100%" sure the idea that the fab can be dirty is wrong: fab air is replaced every three seconds. The hard part is process technology, which only TSMC, Intel, and Samsung integrate, and "these two other companies aren't even that great at it." Clean rooms are the main fab bottleneck this year and next, with tools taking over later.

Patel assigns a very low probability to some simple lithography breakthrough. There are companies building particle-accelerator or synchrotron light sources for 13.5 nm or X-ray-scale lithography, which could disrupt the industry but are themselves very complex. 3D DRAM, following the current path from 6F to 4F cells, is likely by the end of the decade or early next. It would still use EUV but greatly raise bits per EUV pass, though it requires heavy retooling. Lithography's share of wafer cost has risen from about 17% around 2014 to about 30% in logic, and toward the high teens in DRAM.

Patel suggests someone could arbitrage EUV the way early movers bought turbine slots from Siemens, Mitsubishi, and GE Vernova: pay ASML a billion dollars for the right to buy ten tools in two years, then resell the option once everyone realizes the shortage. Patel doubts ASML or TSMC would ever agree.

Power Is Solvable

Patel stresses that their gigawatt figures are critical IT load. Transmission, conversion, and cooling losses add 20–30%, and grid operators like PJM plan about 20% reserve with turbines derated to around 90%, so nameplate capacity must be much higher. Still, Patel sees many paths beyond the three combined-cycle turbine makers: aeroderivatives (including new entrants like Boom Supersonic working with Crusoe), medium-speed reciprocating engines from makers like Cummins with spare automotive capacity, ship engines (Nebius uses them for a Microsoft data center in New Jersey), Bloom Energy fuel cells, solar and wind with batteries. Utility-scale batteries or peaker plants covering the grid's 10–20% summer peak could, in Patel's framing, unlock around 20% of a terawatt-scale grid. Data centers are 3–4% of US power today and will be about 10% by 2028, Patel says. SemiAnalysis tracks over 16 gas-generation vendors with hundreds of gigawatts of orders and expects about half of added capacity by decade's end to be behind the meter, each source contributing tens of gigawatts.

Cost is no obstacle in Patel's view. Even at $3,500 per kilowatt against combined-cycle's $1,500, a Hopper's $1.40 hourly cost rises only to about $1.50, trivial next to model value.

Labor, Modularization, and Permitting

Dwarkesh extrapolates from Crusoe's 1.2 GW Abilene site for OpenAI, which Patel recalls peaking at about 5,000 workers, to 400,000 people for 100 GW. Patel agrees labor is "a humongous constraint." Electrician wages may double or triple again and skilled workers may be imported. The main relief, Patel expects, is modularization: factory-integrated blocks from Korea, Southeast Asia, and China, such as two-megawatt power blocks, integrated cooling units, and full rows of servers on skids as racks approach a megawatt with Nvidia's Kyber. Crusoe, Google, and Meta are pursuing this. Early adopters risk delays; laggards face labor shortages.

On Musk's claim that permitting will block hundreds of gigawatts on Earth, Patel says data centers use little land, the Trump administration eased air permits, and places like Texas cut red tape. Musk's Memphis choice likely reflected grid, gas, and water access at an idled appliance factory; Patel isn't sure why that site was chosen but bets Musk would pick Texas in hindsight. Workers can be housed temporarily and paid well, since labor is cheap relative to GPUs. Data centers are also going up in Australia, Malaysia, Indonesia, and India, though over 70% of AI data centers remain in the US.

Space GPUs Aren't Happening This Decade

Power is nearly free in space, Patel acknowledges, but energy is a small share of cost. Through ClusterMAX, which rates over 40 clouds, SemiAnalysis sees failure management as the main differentiator: about 15% of deployed Blackwells need RMA. Testing on Earth, disassembling, launching, and bringing chips back online could add six months, about 10% of a five-year life, and the earliest months are the most valuable while compute is scarce. Networking is another issue. Starlink's 100 Gbps links compare poorly with per-GPU InfiniBand bandwidth multiplied by 72 per rack and doubling each generation. Sparse models with hundreds or a thousand experts spread across many chips, networking is 15–20% of cluster cost, and space lasers would replace cheap, mass-produced pluggable transceivers that are already less reliable than GPUs.

Dwarkesh suggests hotter chips radiate heat better. Patel corrects this: chips can't run hotter, only denser. Raising power density from about one to two watts per square millimeter of die might give roughly 20% more tokens per wafer but requires exotic liquid or immersion cooling, which is easier on Earth. Patel's conclusion: space data centers share the same contended resource, chips. They may become a 10x win once Earth's resources grow contentious, perhaps once the semiconductor industry catches up around 2035, but not this decade.

Scale-Up Domains and Model Size

Patel explains scale-up domains. Nvidia went from eight all-to-all GPUs per H100 server to 72 in Blackwell NVL72. Google's TPU pods number in the thousands (about 4,000 for v4, eight to nine thousand for v7/v8) in a torus where each chip links to six neighbors and traffic must hop. Amazon sits in between, and all three are moving toward dragonfly topologies.

Asked whether small Nvidia scale-up memory explains slow parameter growth since GPT-4, Patel says partly, but the bigger factor is RL speed and research. Most compute should go to research, since efficiency gains make models about ten times cheaper per year at fixed capability. A 5T-parameter model has rollouts five times larger than a 1T model; even if twice as sample-efficient, it needs 2.5x the RL time. The smaller model gets done sooner and feeds back into research faster. Google deploys the largest production model, Gemini Pro, larger than GPT-5.4 or Opus, because its fleet is almost all TPU. Anthropic juggles H100s, H200s, Blackwell, Trainium, and TPUs.

Why More Funds Aren't Making the AGI Trade

Dwarkesh asks why Leopold Aschenbrenner seems uniquely to profit from SemiAnalysis data. Patel says about 60% of revenue comes from industry and 40% from hedge funds. Leopold is the only client who says the numbers are too low; everyone else says too high, and hyperscalers sometimes take six months to a year to accept them. Many funds use the data and did well; Leopold simply has "the most conviction." A year ago, predicting quadrupled memory prices and a 40% smartphone decline sounded crazy. It took belief in AI takeoff to trade on it.

Apple, N2, and Huawei

Patel doesn't expect TSMC to kick Apple off N2. More likely, AI customers prepay for expansions and Apple gets its order trimmed to "X − 1," losing its traditional 10% flex buffer. Apple has most of N2 this year (with some AMD), about half the next, then far less. N2 is the first node where Apple isn't first, besides Huawei in 2020, since AMD is risking a CPU and GPU chiplet there, and A16's first customer is AI. Apple becomes "any old customer."

Patel thinks Huawei with 3 nm could potentially beat Rubin. Huawei had the first 7 nm AI chip, two months before TPU and four before A100 by Patel's recollection, and uniquely has software, networking, AI talent, its own fabs, and its own token business. Had it not been cut off from TSMC in 2019, Patel believes Huawei would likely now be TSMC's biggest customer.

Robots and Taiwan Risk

On humanoids, Patel expects heavy cloud offloading. A large batched model plans and identifies objects, while a smaller onboard model handles force and motion at perhaps one to ten updates per second. Onboard-only processing would be more expensive, less intelligent, and would consume leading-edge chips that would otherwise go to data centers. Patel reads Musk's Samsung deal to make robot chips in Texas as hedging Taiwan risk and avoiding competition on TSMC, where Nvidia's new LPU is the only other notable AI chip on Samsung.

On airlifting TSMC's engineers in a crisis, Patel says the know-how would survive, but the tools themselves depend on chips made in Taiwan, "a dragon eating its tail." Without Taiwan's fabs, China would have the more vertically integrated supply chain. Rebuilding in Arizona would take years, global GDP would shrink, and incremental compute additions would fall from hundreds of gigawatts a year to perhaps 10–20 gigawatts across Intel and Samsung, which Patel calls "nothing" next to the expansion now underway.