Why Smarter AI Could Make Compute Up to 10x More Expensive

Open on YouTube ↗
Overview

In this narrated version of a blog post, Dwarkesh asks what the compute situation for frontier AI labs will look like over the next few years if current revenue trends continue. The argument is that lab revenue is growing much faster than lab compute. If that continues, the difference has to show up somewhere, and the most plausible place, in the speaker's view, is a sharp rise in the price of compute itself.

10 min read

The gap between 10x revenue and 3x compute

The starting point is Anthropic's revenue. According to the speaker, it has grown 10x year over year for three consecutive years and is likely to do so again this year. Anthropic ended last year with $9 billion in revenue, and the speaker expects it to end this year somewhere between $100 billion and $150 billion. If the trend held, Anthropic would need $1 trillion in revenue by the end of next year.

Dwarkesh stresses that nothing guarantees this. The conclusion is "very wild," and whether it happens depends on whether AI actually becomes that useful by the end of next year. The essay is a thought experiment: suppose the trend does continue, and work out what follows.

The second trend is that lab compute grows only about 3x per year. For revenue to keep growing 10x while compute grows 3x, the speaker says at least one of three things has to happen:

  1. Lab margins increase.
  2. The price of compute increases.
  3. Labs spend a larger share of their compute on inference rather than training.

All three are already happening

The speaker believes all three are already underway and gives an example for each.

  • Margins: Anthropic's inference margins reportedly rose from about 40% in the middle of last year to upwards of 80% now for Fable.
  • Compute prices: spot prices are more than 40% higher than at the trough in February of this year.
  • Inference share: according to Epoch, OpenAI spent only about a quarter of its compute on inference in 2024. The speaker thinks that figure is now likely closer to 50%, if not higher.

Why labs resist shifting compute to inference

Dwarkesh argues that labs would rather not pull the third lever. As the labs see it, inference revenue exists mainly to convince investors to fund training of the next, bigger model. A lab that spends most of its compute on inference is effectively declaring that AI progress has stalled and that it has become a cloud provider. That is a less compelling business than building AGI.

The labs also don't believe they are in that world. They expect that within a year they will have models that make today's models look "extremely shitty." Building those models takes the majority of their compute for training and experiments. The speaker therefore sets this option aside, which leaves two ways to close the gap. Either lab margins rise and the labs capture the surplus, or compute prices rise and everyone in the stack below the labs captures it.

Why margins above 90% seem implausible

The speaker says it is unclear which world we end up in. If the margin effect dominates, top models would go from roughly 80% margins to above 90%. Dwarkesh doubts this. In a market economy, margins exist because what you sell is much better than what a customer could buy from someone else. Margins above 90% would require the leading model to be very far ahead of the competition. The speaker finds it "really wild" to imagine margins on something like intelligence staying above 90% without being competed away.

That leaves rising compute prices as the main escape valve. Dwarkesh notes that this is already happening, and that the effect is stronger for the kind of compute frontier labs actually need. They can't just buy spot instances. They need enough scale for good efficiency and flexibility, and they need compute that meets the security requirements for their own model weights and their customers' data.

Case study: the SpaceX compute deals

The speaker points to compute that Google and Anthropic are renting from SpaceX. According to the essay, Google pays $900 million a month for about 110,000 GPUs, a mix of GB200s and GB300s. Dwarkesh says this works out to roughly 2x the spot price per GPU-hour. That spot price is itself more than 40% above its February level.

The core claim: smarter models monetize the same compute better

Dwarkesh calls this the key conclusion: as AI models get smarter, they can extract more revenue from the same amount of compute. The illustration is a truly human-level software engineer running on one H100-equivalent. At today's software engineer salaries, that H100 should rent for over $250,000 a year. The speaker says that is more than 15x the current H100 spot price, before counting the fact that an AI can work nights and weekends.

The obvious objection is that ten million extra software engineers appearing in the economy would lower the marginal value of a software engineer, so the H100's earnings would not really be 15x higher. Dwarkesh isn't sure this holds. Applied to people, the argument would be the classic lump-of-labor fallacy. The speaker notes that economists generally believe high-skill immigration does not lower wages in the long run, because innovation and specialization raise the value of labor. This supply shock might be so big and so fast that the heuristic breaks down. But if standard economics holds, the speaker argues, the marginal value of labor, and therefore of compute, should stay "astonishingly high."

What changes in that world

The speaker draws three implications.

Competition gets harder. If the top labs keep getting better at monetizing compute while compute gets more expensive, other players have to bid for the same resource against someone who can use it more productively. That makes it harder to compete with the leaders.

Efficient models command a premium. Dwarkesh calls this the most interesting implication of the exercise and ties it to the Alchian-Allen effect. If renting an H100 costs $20 an hour, using a weaker, less efficient model would be "extremely stupid," because it burns more tokens on expensive compute to reach the same result. A lab that trains a model that uses this scarce input more sparingly can charge a much larger premium. Getting the same result with less compute is, in a sense, creating more compute, and that compute is becoming more valuable.

Cheap uses get priced out. The speaker suggests many popular AI applications will likely become unaffordable. AI is relatively cheap today because it can't yet do much of what top humans can do. Once that changes, Google, Anthropic, or OpenAI will pay more for tokens to automate AI research than ordinary users will pay to produce what Dwarkesh calls "AI slop."

Is this just another wrong scarcity prediction?

Dwarkesh admits the analysis resembles past predictions of scarcity that turned out wrong. The example is the Simon–Ehrlich bet. Paul Ehrlich, a well-known pessimist about population growth, bet that a basket of commodities would rise in price over the decade before 1990. The bet is usually cited to show that Ehrlich's Malthusian view underestimated how market signals and human ingenuity find ways to economize on scarce inputs.

The speaker thinks the analogy probably doesn't hold, for two reasons. First, other analysis has shown that Ehrlich might have won if the bet had covered a different decade. Second, and more generally, the speaker argues that compute supply is much less elastic than metal extraction. It is less able to absorb large demand shocks and less able to be replaced with substitutes.

Why 3x compute growth is hard to speed up, or even sustain

To explain why the 3x annual growth in compute is hard to raise, Dwarkesh breaks it into three components and argues none can be much accelerated:

  • About 1.4x from Moore's Law. Far from speeding up, the speaker says it will be "a miracle" if it keeps going for a few more years.
  • About 1.2x from new fabs. This is bottlenecked, until 2030 and possibly beyond, by how fast ASML can build new EUV machines. The speaker points to a podcast conversation with Dylan a few months earlier that covered this in detail.
  • About 1.8x from AI taking wafer allocation from smartphones and PCs. The speaker expects this to hit a wall by the end of next year. At TSMC's leading-edge N3 nodes, AI's share will have gone from 60% to 86%. Once AI has absorbed essentially all leading-edge wafer capacity, this number can't keep rising.

Given these limits, Dwarkesh says it is unclear how compute scaling can even continue at 3x per year for the next few years, let alone go faster.

A pre-singularity regime

The speaker clarifies that compute will eventually get cheap again. At some point, robots may be able to turn silica sand and copper directly into chips, and the price of compute will then be roughly the cost of raw materials and processing tools. The argument applies only to the current "pre-singularity regime," in which AI compute grows merely 3x per year. In the speaker's view, that is not enough to offset how quickly AI is becoming more valuable.

Economies of scale and a worry about concentration

Dwarkesh ends with one more observation. The fact that Anthropic's revenue has grown 10x a year while its compute has grown only 3x shows, in the speaker's view, how strong the economies of scale are in the model business. That makes sense to the speaker: training a model is a one-time cost of learning many skills, which are then shared across all users. Human labor is different, because each person has to be trained from scratch.

The speaker says they wish intelligence did not have such strong economies of scale, because of concerns about power concentration, but concludes that it seems to.