Dylan Patel on Why Two AI Labs May Soon Control Most of the World's Compute and Workforce

Open on YouTube ↗
Overview

Dwarkesh Patel's yearly conversation with Dylan Patel, founder of SemiAnalysis (the two are not related, despite the running joke), starts from one premise: the world economy is increasingly a function of AI lab economics and the compute market. Across the session they project lab compute and revenue a few years out and ask who captures the value. They then follow the capital requirements into a speculative scenario of rising interest rates, sovereign defaults, and collapsing non-AI equities. Dylan's position throughout is that almost every force in the industry points toward concentration in OpenAI and Anthropic, and that only regulation, credit markets, and politics are slowing that down. Neither speaker could say what would reverse it.

26 min read

Where lab compute and revenue stand today

Dylan began with the macro picture. By the end of last year, he said, most US GDP growth was AI infrastructure. This year, about a third of the compute coming online is ultimately for OpenAI and Anthropic, even when other companies build it and rent it to them. Total capex is a little over a trillion dollars this year, and he expects it to exceed $2 trillion by 2028. The labs are moving from spending tens of billions a year to hundreds of billions, and some of the contracts they have signed with partners imply trillions a year toward the end of the decade.

That requires a change in their economics. Until recently, both labs mostly lost money on venture funding. Dylan said Anthropic began turning a profit in Q2. He said OpenAI is believed to be capable of turning a profit sometime in Q3, helped by Codex and GPT-5.6. Both still raise capital to grow faster, but more of the business is now funded from revenue.

The key change in his account is gross margin. Compute costs roughly $10–15 million per megawatt. When OpenAI served GPT-4 on Nvidia Hopper GPUs, it ran at negative gross margin. Now, serving GPT-5.6, Opus 5, or Fable 5, revenue per megawatt is well above that cost. For Anthropic, Dylan said it has reached as high as $50 million per megawatt. His summary of the new logic: spend $10 on inference capacity, earn $50 in revenue, and spend the profit on training.

How fast the labs are absorbing the world's compute

Dwarkesh asked when more than half of incremental compute would go to the labs. Dylan's numbers:

  • OpenAI started this year at about 2 gigawatts and Anthropic at under 2. Both end the year above 5, roughly a 3–4x increase.
  • That growth is about 30% of all compute added this year.
  • Based on contracts already signed, the two labs take 40–50% of new compute next year.
  • By the end of next year, half of incremental compute goes to them. Because compute is growing so fast, incremental compute soon makes up most of the total stock.

The builders will change. Dylan named SpaceX as a big new entrant next year. It is building a lot of compute and, he expects, will likely lease much of it to Anthropic and OpenAI, because they can pay the highest price. The labs are also building their own: OpenAI with its own chips, and Anthropic with TPUs bought from Google and deployed with Fluidstack. He added a counting convention: when Amazon serves Anthropic models through Bedrock, SemiAnalysis counts that as Anthropic compute, since it ends up as Anthropic revenue despite revenue-share arrangements.

Dwarkesh noted that frontier-lab compute seems to triple yearly while world compute roughly doubles. Continuing that trend gives about 6 gigawatts per lab at the end of this year, 18 at the end of 2027, and 54 at the end of 2028. Dylan estimated world incremental additions at about 30 gigawatts this year, 50 next year, and roughly 70–80 in 2028. He called the 2028 figure his "so fucking bullish" upper bound. Adding it up gives over 200 gigawatts globally by the end of 2028.

Watts also understate the concentration. New chips such as GB300s, TPUv7s, and Trainium3s deliver 3–5x more performance per watt than the previous generation, and they are better suited to AI workloads. If the labs take about half of new compute by December 2027, that half is also the most capable hardware. Dylan said that if the trend continues, and he sees nothing stopping it, by late 2028 the two labs control most of the world's usable FLOPs.

$6 billion of fab capex, a trillion dollars of revenue

Dwarkesh questioned why world compute would grow only by tens of gigawatts a year if compute is so valuable. He built a rough calculation from figures Dylan gave in an earlier interview: a gigawatt of Vera Rubin requires about 55,000 N3 wafers, 6,000 N5 wafers, and 170,000 DRAM wafers. He had an LLM run Dylan's wafer fab equipment model. It estimated $3–4 billion of tooling to produce a gigawatt of compute per year, or about $6 billion including cleanrooms and fab shells. If a gigawatt generates about $100 billion a year, each year's output of that fab keeps earning for years. Over five years, the $6 billion yields over a trillion dollars of end revenue. Dylan pointed to the other costs along the way: operating costs, data centers, power, installation, and lab R&D. Dwarkesh halved the figure for all the middlemen and still got a 100x gap between fab capex and end revenue. Dylan said the true gap is larger and the calculation was conservative.

Dwarkesh asked why capitalism wouldn't close the gap by making more ASML mirrors. Dylan agreed it eventually would, but said the supply chain responds like a whip: the signal takes a long time to reach the far end. Arbitrage is already happening. People buy turbines to resell because turbines bottleneck data centers. Dylan said anyone with $400 million who could persuade ASML to sell them an EUV tool should buy one and later sell it for over a billion dollars.

On Carl Zeiss, which makes the mirrors, Dylan said that earlier this year the company did not think it needed mirror capacity for 100 EUV tools a year by the end of the decade. It now accepts that target, but Dylan thinks the economics justify even more. At the current rate of expansion, he considers about 100 tools by 2030 still the right number. Handing Zeiss $10 billion would change that, but the same would have to happen at every company in the chain.

Dwarkesh asked whether the labs will soon have enough cash flow to fund that expansion themselves. Dylan did not expect it this year, next year, or the year after, because the world is capital-constrained. He expects the labs to generate hundreds of billions in revenue next year against about $2 trillion of total capex. That total includes roughly $200 billion in wafer fab equipment plus larger amounts for data centers, accelerators, and energy. He added that the labs will never fully fund capex from cash flow, since the goal is always to invest more than current returns.

Compute prices have to rise for the labs to win

Reaching 100 gigawatts between them by 2028 would mean the labs taking 70–80% of incremental compute. Dylan said that would disrupt the market, because at today's prices almost anyone can make money on compute. His example: buy a GB300 rack, download Kimi weights, set up vLLM or SGLang (Codex and Fable can help), and list it on OpenRouter. He said revenue will exceed the compute cost. This has already started pushing prices above $10–15 million per megawatt. For the labs to take most of the market, they would have to pay $25, $30, or $50 million per megawatt.

Dwarkesh argued that the labs' lead in revenue per megawatt should keep growing, especially if they have unreleased internal models helping them build the next one. Players slightly behind, such as SpaceX, would then rationally sell to the highest bidder. Dylan agreed with that view, then raised the main caveat: regulation and safety constraints, which he said already slow the US labs more than Chinese open models. His examples were OpenAI not releasing Astra, OpenAI pausing training for two weeks, and Anthropic not releasing what its safety assessment calls "Model 2," widely believed to be the next Mythos. If labs cannot ship their best models, revenue per megawatt stalls or falls as other models catch up, and so does their ability to outbid everyone. In a world where safety didn't matter, Dylan said, the labs could earn $100 million per megawatt or more and pay $50 million. Everyone else would just ask Dario to take their compute.

Dwarkesh offered an intuition pump. If a gigawatt could sustain about a million fully automated white-collar workers at $100,000 each, that would be $100 billion. He found that surprisingly low and said full AGI would mean many hundreds of billions per gigawatt.

Which layer captures the surplus

Dylan said most of the value AI creates currently goes to users, not the labs. His examples were Jane Street, with an exclusive OpenAI contract for GPT-5.6 Ultrafast mode and a spot among Anthropic's biggest customers, and Meta, once rumored to be up to 10% of Anthropic's business. Both, he said, earn far more from the tokens than Anthropic earns in profit, through trading or through ad-algorithm and engagement gains.

Dwarkesh asked whether the market would reach an equilibrium where compute prices approach what the labs can earn from it, given the current 4x-or-more gap. Dylan described value capture moving through the stack. A year ago, the models ran at negative gross margins on VC money, hyperscalers built without knowing if it would pay off, and the hardware chain took the margin. In 2023, memory makers earned almost nothing on HBM despite its value. Now memory captures more than TSMC. The model layer has recently moved to large positive margins.

Elon Musk, in Dylan's telling, showed that the labs won't necessarily take everything. SpaceX sold compute to Anthropic and Google at $25–40 million per megawatt, enough to recoup its capex in about a year. Dylan still predicted that most compute will trade below $20 billion per gigawatt through the end of next year, because most compute is contracted and financed before it is built. A typical cloud needs a customer commitment to raise capital from credit markets. Meta and SpaceX are different: they have balance sheets to build without a signed customer, which lets them hoard compute and later decide whether to use it internally or rent it at high margins. Dylan called them the only plausible #3 players.

Revenue per megawatt and the regulatory brake

For the end of 2027, Dylan estimated lab revenue of at least $50 million per megawatt, possibly $70–80 million blended across the company, provided labs can keep releasing their best models. Dwarkesh said that seemed low given how far models have come in the past year and a half. Dylan predicted a bullwhip effect in prices. If Anthropic pays SpaceX $40 million per megawatt, Nvidia can raise prices, and so can SK Hynix, Micron, and Samsung. In his account, memory and substrate suppliers are raising prices quickly and TSMC slowly.

Dwarkesh argued that a leap like GPT-4o to Mythos 2 by the end of 2027 should push revenue per gigawatt far higher. Dylan kept returning to deployment. By his account, the best model in the world was trained in February, Mythos 2 isn't out, Mythos itself has been "neutered" so they can't use it to optimize inference, and Astra isn't widely deployed even inside OpenAI. He said regulation is also spreading to physical supply: New York banning data centers, Texas holding moratoriums, and Ohio trying to require operators to pay property taxes within a radius. That raises costs, which get passed on, and slows external progress even as internal models improve. Dwarkesh added that in a takeoff, a lab would have competitive as well as regulatory reasons to keep its best model six months ahead of what it releases. With accelerating progress, that gap matters more and caps revenue-per-megawatt growth.

Why the labs will shift compute from inference to R&D

Dwarkesh posed a hypothetical. Suppose a public lab with about 20 gigawatts wants to move from 60% to 70% of compute on training, giving up roughly $200 billion of revenue at $100 billion per gigawatt. How would investors react? Dylan's answer, which he called very non-consensus, is that labs will allocate less and less compute to inference, against the common view that inference will dominate. He expects more compute to go to forward passes for training than to revenue-generating inference.

His reasoning: at $30–40 million per megawatt, a lab might put 40% of compute on inference. At $60–70 million, keeping 40% would produce huge profits to spend on dividends and buybacks, or the lab could build AGI instead. He thinks executives and boards will choose AGI as the more profitable path. Ultrafast modes will serve internal researchers as well as customers, because the internal value is higher. On this view, the main purpose of inference revenue is to fund the training fleet.

Dylan said this is already visible. Anthropic adds more compute nearly every month, but after new revenue surged early in the year, it stopped adding around $25 billion of ARR each month. So, by his reading, the marginal megawatt is going more to R&D than inference, which he called self-evident to anyone watching closely.

China: under 10% of new compute, but its labs need less

Dylan said that in 2022 the US added about 45–50% of new compute and China 30–35%. Since then, export controls and the US buildout have left America with about 70% of new watts and China under 10%. China's domestic production and Nvidia purchases remain small, and some purchased chips end up elsewhere, such as Malaysia. He put China at 30 gigawatts or less by 2028.

He expects 2026 to depend on smuggled chips, TSMC chips made for companies later revealed as Huawei fronts, and HBM shipped by Samsung. SMIC and CXMT fabs start ramping in 2027, and especially 2028, reaching many millions of units a year and adding 5–10 gigawatts of domestic chips in 2028 alone. Those chips will be worse than Nvidia's, Google's, or OpenAI's 2028 chips. How steep China's hockey stick gets depends on whether the US passes the MATCH Act, whether tool export controls continue, and how fast China builds its own equipment. Dylan is sure it will come, since scaling manufacturing is what China does best. He called 50 gigawatts in 2029 "completely reasonable," partly from foreign purchases, though mostly domestic chips might make that worth about 20 American-chip gigawatts.

Dwarkesh concluded that the leading US lab in 2028 might have more quality-weighted compute than all of China in 2029 or 2030. Dylan qualified this: it assumes nothing slows the US labs, while politicians are already trying to, and China will only accelerate. Dwarkesh said that when he interviewed Jensen Huang he had steelmanned cooperation with China, partly because China controls supply chains robotics will need. He hadn't realized how lopsided the compute situation was and now sees export controls as a notable success. Dylan added that the gap also reflects finance: American markets are more willing to "YOLO" into startups, but once China's system targets an industry it subsidizes heavily. He said Chinese semiconductor subsidies exceed those of the rest of the world combined. If takeoff is slower, he expects China to catch up drastically on chips.

How labs actually spend their compute

Dylan also noted that Chinese models are not far behind in public perception given their compute. Leading Chinese labs have at most 100–200 megawatts, with ByteDance Seed the outlier, and Kimi is nowhere near a gigawatt. Anthropic will have more than 5 gigawatts by year end. He thinks the difference doesn't matter much yet, because of how labs split their budgets.

Labs have spent about 60% on training and 40% on inference, but the training share splits into roughly 50% research and 10% development. Research means testing ideas, architectures, data mixes, hyperparameters, and attention techniques. Development means the training run itself. He said Anthropic's Mythos pre-training used under 200 megawatts for about two months, and RL used less at any single site, though total compute was probably higher because the phases ran sequentially. Most of a lab's multiple gigawatts go to research. That is partly because coordinating clusters, multi-site training, and RL at scale are hard, and more RL rollouts don't necessarily help. As automated coding, automated research, and continual learning arrive, he expects the research/training line to blur and training's share to rise.

The scale of capex, and who pays

Dylan said 100 gigawatts a year at current prices would mean about $5 trillion of capex a year. Power plants are 30-year assets and data centers 15–20-year assets, and both must be built before the chips arrive. Commonly cited per-gigawatt figures of $40–50 billion cover only critical IT: servers, networking, fiber, transceivers, and optics. Accounting for the buildings and power for next year's larger buildout, he put the annual total closer to $7–10 trillion. Dwarkesh noted that $10 trillion a year by 2030 would be about a tenth of the world economy, and at today's US GDP, a quarter to a third of it. Saying that aloud, he wondered whether society simply won't allow it.

Paring back to 2028, Dylan estimated $3–4 trillion: over $2.5 trillion of IT capex plus $1–2 trillion for data centers, energy, and upstream supply chains. Nobody generates that much cash. The hyperscalers (Google, Microsoft, Amazon, Meta), which funded most growth so far, now spend all their cash flow on capex and borrow on top. Meta, Amazon, and Google already do, and Microsoft soon will. Stopping buybacks had some market effect, but by 2028 hyperscalers and their suppliers would be raising hundreds of billions in debt. He named the likely funders: chip and memory companies like Nvidia and Broadcom financing capex, infrastructure investors putting money into data centers instead of bridges, and everyone else. Individuals might skip buying a home or mortgage credit or government bonds to buy hyperscaler, data center, or Anthropic debt. Anthropic might pay 20% rates for an extra billion because that still beats renting from SpaceX at $50 billion per gigawatt.

Could AI trigger a sovereign debt crisis?

Dwarkesh laid out the argument they had debated off air. When a little investment produces a lot of return, the rate of interest rises. Heavy borrowing for data centers competes with governments, companies, and mortgage borrowers, raising everyone's costs. He expects the US to be fine if it builds data centers domestically and taxes them. But corporate income is under 10% of federal revenue, and over 80% comes from payroll and income taxes that automation will shrink. About 20% of tax revenue already goes to interest, and much of the debt rolls over within about five years. In his rough figures, a 1-point rise in rates takes interest to about 25% of revenue over five years. A 5-point rise takes it past 40%, and past 60% once you include the roughly $2 trillion borrowed each year. He expects countries with heavy debt, low tax revenue, and frequently rolled debt, such as Pakistan or Nigeria, to be hit very hard.

Dylan said this crowding-out is why the world won't build unlimited gigawatts. Consumer packaged goods, telecom, and banks all rely on debt. What matters is less the Fed rate than the spread markets charge as, for example, Amazon borrows heavily. SemiAnalysis models about $11 trillion of capex from 2024 to 2029: about $6 trillion from cash and $5 trillion from debt. That debt pushes rates up. Even so, he considers it not enough compute for the demand, so revenue per megawatt keeps rising. Regulation, angry consumers and politicians, and withheld models all push the other way, bending the buildout below what pure economics would want.

Asked for a rate, Dylan gave what he called a heavily vibed guess. Meta recently borrowed at 5–6% and would gladly pay 8% given the returns on compute. A roughly 250-basis-point rise would spread to everyone else. Banks would suffer because their debt reprices faster than their assets. Dwarkesh added that a higher discount rate hits equities built on long, steady cash flows. The index might be fine, but most individual stocks would fall, with Johnson & Johnson and railways as the examples. Dwarkesh cited economist Basil Halperin's prediction of a "second Volcker shock." In the 1980s, Paul Volcker's rate hikes, which Dwarkesh described as reaching about an 8% real rate, were followed by about 40 countries defaulting, mostly in Latin America. Dwarkesh expects something similar, and Dylan said all of this happens before any singularity.

Explosive growth and the reallocation of capital

Dwarkesh then went further. Citing researcher Damon Binder's input-output work, he argued that if the labor force can double every year, the whole economy could eventually double yearly, or at least grow tens of percent a year. Interest rates should roughly track growth, so in the 2030s rates might be tens of percent, possibly hundreds. In that world, he said, every country not producing AI defaults, non-AI stocks are worth almost nothing, the federal government can't service its debt unless it taxes AI, and mortgages become unobtainable. The underlying cause is the opportunity cost of capital: money used to pay pensions could instead build robot factories that build more robot factories.

Dylan applied this to markets. People ask why memory makers like Micron, Hynix, or Kioxia trade at 2–3x earnings. His answer: if you're truly AI-pilled, everything should trade at 2–3x, and the market should crash. So he thinks memory will do great but its stocks shouldn't 10x again. Conversely, he called Meta at about $1.5 trillion "silly," given the compute it is hoarding and could monetize through its own lab or by selling to Anthropic and OpenAI. The reallocation happens by pricing everyone else out. So the limit on AGI, he said, is not how fast researchers like their roommate Sholto work, but how much the rest of the world allows: regulation, interest rates, data center and fab opposition, and falling equity values that make other companies less able to buy AI. That pushes the labs to build their own chips and infrastructure. Dylan believes the models could support a fast takeoff but hopes a slow one is possible because of these frictions.

Dwarkesh worried more about something else. Restricting external deployment is counterproductive, he argued, because a rule like a six-month wait before public release would let recursive self-improvement happen inside the labs while the public uses models years behind. Slowing compute by a year also counts for little if RSI then yields 3–6 years of progress in one year. Dylan expects governments to restrict internal use too. He cited Anthropic's claim that it stopped giving Mythos to foreign employees for a while, and predicted the US government won't let Anthropic freely use "Mythos 4" internally. His reasoning is political: elected officials and their constituents already hate AI. Society might tear itself apart before AI arrives.

Most of the world's labor inside two companies

In the final segment, Dwarkesh raised what he finds most striking. Frontier compute is growing 4–5x a year in FLOPs, and the compute needed for a given capability falls about 3x a year. Together, the effective AI population at frontier labs grows about 10x a year. That matters little now, since AIs can't do full jobs autonomously. But if trends hold, OpenAI could go from roughly 10 million AI workers this year to 100 million next year and a billion after that. Dylan called it very plausible that by the end of the decade a single lab has more effective labor than there are people on Earth. Dwarkesh noted this holds without RSI, and with RSI growth might be 100x or 1,000x a year, or intelligence might rise instead of headcount. If these AIs are misaligned, most of the world's minds are; if not, very few companies still hold enormous influence.

Dylan brought up the recent dispute where Gavin Baker said Dario believes there will be only one company in the world, and Sholto and Dario denied it. Dylan's view: if labs use compute most productively and RSI is real, centralization follows. He asked Dwarkesh what non-centralized world he could imagine, calling the direction "scary as hell." Dwarkesh named the structural drivers:

  • Training has huge economies of scale, because skills trained once are amortized over billions of sessions.
  • Under compute scarcity, the leader can charge a higher markup.
  • Models learning from deployment favor the most widely used one.

He called designing a decentralized, broadly empowered future that takes these economies of scale seriously a major intellectual project. The alternative, government control, doesn't reassure him. He said he trusts neither the government nor Dario nor Sam Altman.

Dylan contrasted this with capitalism's success through decentralized decisions, which AI may overturn by making a centralized AI economy grow faster. Dwarkesh noted the property may stay private, but AI is already roughly 2% of the economy by his rough math, concentrated in Nvidia, Anthropic, OpenAI, and the hyperscalers. Asked what could prevent concentration, Dylan said "I don't know." Unless progress slows or governments regulate heavily, he sees two paths: extreme concentration where we hope one company gets everything right, or a slowdown that preserves more balance of power.

Dylan's one hopeful point was that Anthropic doesn't capture most of the value. It still pays about $13 million per megawatt for much of its compute, and it can charge $100 million because customers like Jane Street capture $300–500 million per megawatt. Dwarkesh countered with Dylan's own earlier argument: shifting compute from inference to R&D makes sense only if labor is worth more inside the labs than outside. Dylan conceded that this was his "cope." If Anthropic can earn hundreds of millions per megawatt using compute internally, it has no reason to let Jane Street keep that value, and in his view that is already happening. The conversation ended there, with no answer to what could counter the pull toward centralization.