Jensen Huang on Nvidia's Moat: Supply Chains, TPUs, CUDA, and the Case for Selling Chips to China

Open on YouTube ↗
Overview

In this long interview, Dwarkesh presses Nvidia CEO Jensen Huang on whether Nvidia's dominance can last. The questions cover whether Nvidia is really a software company that others manufacture for, whether its hold on scarce components is the true moat, whether Google's TPUs and other custom accelerators threaten it, why Nvidia doesn't become a cloud provider, and, at greatest length, whether the United States should allow Nvidia to sell AI chips to China. Huang's position throughout is that Nvidia's advantage is not any single factor. In his account it comes from a combination of programmability, ecosystem, install base, supply-chain reach, and an annual cadence of large performance gains. On China, the two sides of the conversation never fully converge.

33 min read

Is Nvidia a Software Company That Can Be Commoditized?

Dwarkesh opens with a deliberately naive framing. Software company valuations have fallen because people expect AI to commoditize software. Nvidia sends a GDS2 design file to TSMC, which builds the logic dies and switches and packages them with HBM from SK Hynix, Micron, and Samsung. An ODM in Taiwan then assembles the racks. If Nvidia is fundamentally making software that other people manufacture, does it get commoditized too?

Huang says the framing matches his own mental model of the company: electrons go in, tokens come out, and Nvidia sits in the middle. He argues that turning electrons into tokens, and making tokens more valuable over time, is hard to commoditize. He compares it to making one molecule more valuable than another, and he says the artistry, engineering, and science involved are "far from deeply understood" and "the journey is far from over." The company's rule is to do "as much as necessary and as little as possible." Whatever Nvidia doesn't need to do, it hands to partners. He describes AI as a "five-layer cake" and says Nvidia has partner ecosystems across all five layers. The part Nvidia must do itself, he says, "is insanely hard," and he doesn't think it gets commoditized.

He also disputes the premise that software companies will be commoditized. He sees many of them as toolmakers, citing Excel, PowerPoint, Cadence, and Synopsys. His prediction runs the other way: the number of agents will grow exponentially, and so will the number of tool users. He expects the number of instances of tools like Synopsys Design Compiler, floor planners, layout tools, and design rule checkers to "skyrocket," because engineers today are limited by headcount and will soon be supported by many agents exploring the design space. In his view this hasn't happened yet only because agents aren't good enough at using tools. He expects it to come through some combination of software companies building their own agents and agents getting better.

Purchase Commitments and "Informing" the Supply Chain

Dwarkesh cites Nvidia's filings showing close to $100 billion in purchase commitments with foundries, memory makers, and packaging suppliers. He adds SemiAnalysis's report that the figure will reach $250 billion. Could the real moat be that Nvidia has locked up years of scarce components, so a rival with a good accelerator can't get the memory or logic to build it?

Huang accepts this as one thing Nvidia can do that others find hard, but he describes it as more than contracts. Some commitments are explicit. Others are implicit: suppliers make investments because he has personally explained to their CEOs how big the industry will be and why. They are willing to invest on Nvidia's behalf rather than someone else's, he says, because they know Nvidia has the downstream demand to buy their supply and sell it through. He compares this to cash flow. There is also "supply chain flow," and nobody builds a supply chain for an architecture whose business churns slowly.

He presents GTC, Nvidia's conference, as part of this effort. It puts upstream suppliers, downstream customers, and AI startups in one place so they can see each other and see firsthand what he has been telling them. He acknowledges that his keynotes can feel "a little torturous," closer to education than a string of announcements, and says that is deliberate. He wants the whole ecosystem to understand what is coming, when, and how big it will be, and to "reason about it systematically, just like I reason about it." If Nvidia's next several years reach a trillion dollars in scale, he says, the supply chain is there to support it.

Can the Upstream Keep Up With Doubling?

Dwarkesh pushes on the physical limits. Nvidia has roughly doubled revenue year over year and more than tripled the flops it provides annually. Citing SemiAnalysis, he notes that AI will account for about 60% of TSMC's N3 node this year and 86% next year. How do you double when you're already the majority, and how do you get twice as many EUV machines each year?

Huang argues that instantaneous demand exceeding supply is a healthy condition for an industry. When one component falls far behind, the industry "swarms it." His example is CoWoS packaging. People no longer talk about it much, he says, because for two years Nvidia "swarmed the living daylights out of it," doubling capacity several times over. TSMC now scales CoWoS and future packaging technologies in step with logic, and CoWoS and HBM, once specialty technologies, are now mainstream.

He says he was making today's predictions five years ago, and that some suppliers believed him early. He singles out Sanjay and the Micron team, recalling a meeting where he laid out what was going to happen. Micron doubled down across LPDDR and HBM, which he says has been "tremendous for the company." Others came later. Nvidia now tries to "prefetch" bottlenecks years in advance. Examples include investments with Lumentum, Coherent, and the silicon photonics ecosystem, which he says reshaped the supply chain around TSMC; partnering with TSMC on COUPE and licensing the resulting patents to keep the supply chain open; and developing new testing equipment such as double-sided probing.

Asked directly about EUV, Huang says none of this is impossible to scale quickly and all of it is "easy to do within two or three years" given a demand signal: "Once you can build one, you can build ten, and once you can build ten, you can build a million." He doesn't always go to ASML directly. "If I can convince TSMC, ASML will be convinced." His claim is that no bottleneck lasts longer than two or three years. Meanwhile Nvidia improves computing efficiency by 10x to 20x per generation, and in his account by 30x to 50x from Hopper to Blackwell.

What worries him is downstream: energy, and energy policy in particular. You can't reindustrialize the United States, bring back chip and packaging manufacturing, or build EVs, robots, and AI factories without energy, and those things take a long time. Asked which bottleneck is hardest, he answers "plumbers and electricians." This leads into a criticism of "doomers" who predict the end of work. If people are discouraged from becoming software engineers, he warns, there will be a shortage of software engineers. He notes that people were told a decade ago not to become radiologists, and "guess what we're short of? Radiologists." Dwarkesh remarks that other guests have told him the opposite about which bottlenecks are hard, and that he lacks the technical knowledge to adjudicate.

TPUs and the Case for General Programmability

Dwarkesh points out that arguably two of the top three models, Claude and Gemini, were trained on TPUs. Huang's first answer is that Nvidia builds something different: accelerated computing, not a tensor processing unit. Its GPUs are used for molecular dynamics, quantum chromodynamics, data processing, fluid dynamics, and particle physics as well as AI, so its market reach is "far greater than any TPU or ASIC can possibly have." Because Nvidia's systems are designed to be operated by others, they run in every cloud, including Google, Amazon, Azure, and OCI. They can also be run by a company for itself, as with Elon Musk's xAI, or used for a drug-discovery supercomputer at Lilly. By contrast, he says, most "home-built systems" require the builder to be the operator.

Dwarkesh presses the point. Nvidia isn't earning $60 billion a quarter from pharma and quantum. It's earning it from AI, and his AI researcher friends describe the TPU as a big systolic array ideal for matrix multiplies. GPUs spend die area on flexibility, such as warp schedulers and switches between threads and memory banks, that AI's predictable workloads may not need.

Huang replies that matrix multiplies are important but not the whole story. Inventing a new attention mechanism, disaggregating computation differently, building a hybrid SSM, or fusing diffusion with autoregressive techniques all call for a generally programmable architecture, and he argues that the ability to invent new algorithms is what makes AI advance so quickly. TPUs, like everything else, are subject to Moore's Law, which he puts at roughly 25% improvement per year. Getting 10x or 100x leaps requires changing the algorithm and how it is computed. He recounts announcing that Blackwell would be 35x more energy efficient than Hopper. Nobody believed it, and then Dylan (of SemiAnalysis) wrote that he had "sandbagged" and the real figure was 50x. That result, he says, came from models such as mixture-of-experts parallelized and distributed across the system, from new CUDA kernels, and from "extreme co-design," including offloading computation into the NVLink fabric or the Spectrum-X network. "Without CUDA to do that, I wouldn't even know where to start."

Does CUDA Matter to Customers Who Write Their Own Kernels?

Dwarkesh notes that about 60% of Nvidia's revenue comes from five hyperscalers. Unlike professors running experiments, these customers can write their own kernels. OpenAI has Triton and its own stack rather than relying on cuBLAS and NCCL, and that stack compiles to other accelerators. If the biggest customers can replace CUDA, how much does it matter?

Huang gives three reasons CUDA remains valuable. First, the ecosystem is rich and well tested. Nvidia supports every framework and contributes heavily to Triton, whose back end he says contains "huge amounts of Nvidia technology." He also lists vLLM, SGLang, and new reinforcement learning frameworks such as verl and NeMo RL. Because the foundation is so well "wrung out," a developer who hits a bug can reasonably assume it is in their own code rather than in the mountain of code underneath. Second, the install base: several hundred million GPUs across generations and form factors, including in robots, so software written for CUDA runs nearly everywhere. Third, Nvidia is in every cloud and on-prem, so an AI company that hasn't settled on a cloud partner can run anywhere.

Margins, Kernel Engineering, and Performance per TCO

Dwarkesh sharpens the question. If AI gets especially good at tasks with tight verification loops, and writing an efficient attention or MLP kernel is exactly that kind of task, hyperscalers could write their own kernels. Competition would then reduce to flops and memory bandwidth per dollar. Could Nvidia sustain margins above 70%?

Huang describes CPUs as a Cadillac, a comfortable cruiser anyone can drive, and Nvidia's accelerators as F1 cars that anyone can drive at 100 mph but that take real expertise to push to the limit. Nvidia assigns an "insane" number of engineers to AI labs to optimize their stacks, and it uses "a ton of AI" to create its own kernels. He says it is not unusual for this work to speed up a lab's model by 50%, 2x, or 3x. Across a lab's installed fleet of Hoppers and Blackwells, a 2x gain translates directly into doubled revenue.

He then makes his strongest competitive claim: Nvidia offers "the best performance per TCO in the world, bar none." He challenges competitors to prove otherwise on public benchmarks, naming Dylan's InferenceMAX and MLPerf. "TPU won't come, Trainium won't come," he says, and he invites Trainium to demonstrate the 40% advantage he says it "claim[s] all the time," and Google to demonstrate the claimed cost advantage of TPUs. "On first principles, it makes no sense."

He also disputes the idea that the hyperscaler concentration means a few internal buyers. Most Nvidia capacity at AWS serves external customers, he says, and at Azure and OCI all of it does. The clouds favor Nvidia because it brings them the most customers. He describes the flywheel this way. Among tens of thousands of AI startups, each would choose the most abundant architecture, the largest install base, and the richest ecosystem. Nvidia also offers the lowest-cost tokens through performance per dollar and the most tokens per watt, which matters because a one-gigawatt data center's revenue depends on how many tokens it produces. And for anyone renting out infrastructure, Nvidia has the most customers.

Anthropic, ASIC Margins, and a Missed Investment

Dwarkesh replies that the compute on these clouds is actually used by a few labs, such as Anthropic and OpenAI, that can make other accelerators work. Huang says the premise is wrong and asks to come back to it, because it is "too important to AI." Dwarkesh then asks about Anthropic's recently announced multi-gigawatt deal with Broadcom and Google for TPUs, which would make up the majority of its compute.

Huang calls Anthropic "a unique instance, not a trend." Without Anthropic, he asks, why would there be any TPU growth? "It's 100% Anthropic." He says the same about Trainium. When Dwarkesh raises OpenAI's deals with AMD and its own accelerator effort, Huang says OpenAI is still "vastly Nvidia." He says he isn't offended when customers try alternatives, since that is how they learn how good Nvidia is, and Nvidia must keep earning its position. He points to the number of canceled ASIC projects as evidence that building something better than Nvidia is hard.

Dwarkesh suggests the logic of an ASIC is that it only needs to be no more than 70% worse, given Nvidia's margins. Huang counters that ASIC vendors' margins are also high, around 65% by his estimate, so "what are you really saving?" He adds that ASIC makers are "quite proud of their incredible ASIC margins."

He then offers his own explanation for Anthropic's path, framed as a personal mistake. Early on he did not "deeply internalize" how hard it would be to build a foundation lab, or that labs like Anthropic needed huge investments from their suppliers. Google and AWS made large early investments, and Anthropic used their compute in return. Nvidia had never invested outside the company at that scale, and he assumed the labs could raise from VCs "like all companies do." He now recognizes that no VC would have put $5–10 billion into a lab on the hope it would become Anthropic. Even if he had understood this, he doesn't think Nvidia could have done it at the time. "I'm not going to make that same mistake again," he says, citing Nvidia's investment in OpenAI and later in Anthropic once it was in a position to invest. Though Nvidia's absence pushed Anthropic elsewhere, he says Anthropic's existence "is great for the world."

Why Nvidia Won't Become a Hyperscaler

Dwarkesh cites reports of Nvidia investing up to $30 billion in OpenAI and $10 billion in Anthropic, and backstopping CoreWeave for up to $6.3 billion on top of a $2 billion investment. With so much cash, why not become a cloud and rent compute directly?

Huang returns to "as much as needed, as little as possible." Nvidia should do the work that nobody else would do: building NVLink, building the whole stack, and "20 years of CUDA while losing money most of that time." He also points to the domain-specific CUDA-X libraries Nvidia began building about 15 years ago, for ray tracing, image generation, early AI, and structured and vector data processing, and to cuLitho for computational lithography. "If we didn't create it, nobody would have. I am completely certain of that." The world already has many clouds, however, and if Nvidia didn't build one, someone else would.

Dwarkesh points out an apparent contradiction: Huang says CoreWeave, Nscale, and Nebius wouldn't exist without Nvidia, which sounds like propping them up. Huang explains that these companies first had to want to exist and bring a business plan, expertise, and their own capabilities. If they needed investment to get off the ground, Nvidia would help, but it does not want to be in the financing business and prefers to work with financiers. Large investments like the one in OpenAI, which he describes as needed at "$30 billion scale" before its IPO, are made because the company needs them, not because Nvidia is trying to do more.

He adds that Nvidia deliberately avoids picking winners among model companies: "when I invest in one of them, I invest in all of them." His reason is humility drawn from history. When Nvidia started there were about 60 3D graphics companies, and Nvidia would have been at the top of the list of those not expected to survive. Its early graphics architecture was "precisely wrong," reasoned from good first principles but impossible for developers to support. Yet it was the only one that survived.

How Scarce GPUs Are Allocated

Dwarkesh describes Nvidia's reputation for splitting scarce supply among neoclouds like CoreWeave, Crusoe, and Lambda rather than selling to the highest bidder. Huang rejects the characterization. The first step is joint forecasting with customers, because chips and data centers take a long time to build. After that, "you still have to place an order," and allocation is essentially first in, first out. The main exception is that if a customer's data center or other components aren't ready, Nvidia may serve another customer first to maximize the throughput of its own factory.

He says the story that Larry Ellison and Elon Musk begged him for GPUs over dinner "never happened." They did have dinner, and it was a wonderful one, but "they just had to place an order."

Asked why Nvidia doesn't simply sell to the highest bidder, Huang calls it "a bad business practice." Nvidia sets a price and customers decide whether to buy, even if demand "goes through the roof." He acknowledges that others in the chip industry raise prices when demand spikes and says Nvidia never has. He ties this to dependability and compares it to Nvidia's nearly 30-year relationship with TSMC, which he says involves no legal contract: "sometimes I got a better deal, sometimes I got a worse deal," but he can completely trust and depend on them.

He extends the dependability claim to Nvidia's roadmap. Vera Rubin arrives this year, Vera Rubin Ultra next year, Feynman the year after, and then an as-yet-unnamed generation. He says customers can count on token cost decreasing "by an order of magnitude every single year," as reliably as a clock, and that no ASIC team can make that promise. Customers can order one graphics card or $100 billion of AI factory capacity, which he says only Nvidia (and, at the foundry level, TSMC) can offer. He says this position took "a couple of decades" of commitment to reach.

China, Part One: Mythos and the Cyber-Offense Argument

Dwarkesh says he is unsure whether selling chips to China is good and is playing devil's advocate, having argued the opposite side with Anthropic's Dario Amodei. He cites Anthropic's recently announced Mythos Preview. Anthropic is withholding it from public release because of its cyber-offensive capability, saying it found thousands of high-severity vulnerabilities across every major operating system and browser, including a 27-year-old bug in OpenBSD. If Chinese labs had the chips to train such a model and run millions of instances, would that threaten American security?

Huang says Mythos was trained on "fairly mundane capacity," in a fairly mundane amount, and that this kind of compute is "abundantly available in China." He asserts that China manufactures 60% or more of the world's mainstream chips, has abundant energy, and has about half of the world's AI researchers. He also says most researchers in American AI labs are Chinese. Given those assets, he argues that "victimizing them, turning them into an enemy" is probably not the safest approach. China is an adversary and the US should win, but he calls research dialogue "glaringly missing" and says American and Chinese researchers must talk and try to agree on what AI should not be used for.

On vulnerability discovery itself, he says finding bugs "is what AI is supposed to do." He highlights an ecosystem of AI security, privacy, and safety startups working toward a future in which one powerful agent is surrounded by thousands of agents keeping it safe. The idea of an agent "running around with nobody watching after it is kind of insane." That ecosystem, he argues, needs open source and open models, much of which comes from China, and the US "ought to not suffocate that." He considers it "extremely foolish" to end up with an open-source ecosystem running on a foreign tech stack and a closed one running on the American stack.

China, Part Two: Flops, 7nm, and Whether Energy Substitutes for Chips

Dwarkesh responds with estimates that, lacking EUV, China has about one-tenth the flops of the US. More compute lets American labs reach capabilities first. He notes that Anthropic is holding Mythos back while American companies patch vulnerabilities, and that inference compute determines whether an attacker runs a thousand hackers or a million. He adds that leaders at DeepSeek and Qwen have said they are bottlenecked on compute, so China's strong researchers make additional compute more dangerous, not less.

Huang agrees the US should always be first and have more, but says Dwarkesh's argument only works if China has no compute at all. China is the second-largest computing market in the world. AI is a parallel problem, and with abundant, essentially free energy, China can gang together 4x or 10x as many 7nm chips. He claims China has "ghost datacenters" sitting fully powered and empty, and overcapacity in mainstream chip manufacturing. The US, being energy-scarce, needs performance per watt, but "if your amount of watts is completely abundant… what do you care about performance per watt for?" He says 7nm chips are "essentially Hopper," and that today's models were largely trained on the Hopper generation. As evidence of volume, he cites Huawei's record year and shipments of "millions" of chips, "way more than Anthropic has."

Dwarkesh raises memory bandwidth: HBM2 versus the newest memory could mean close to an order-of-magnitude gap, and he believes advanced HBM requires EUV. Huang calls this "not at all true." He says Huawei, as a networking company, can gang chips together the way Nvidia does with NVL72, and has demonstrated silicon photonics connecting compute into a single supercomputer. He adds that compute-limited researchers invent smarter algorithms, and argues that most AI advances, including MoE and efficient attention mechanisms, came from algorithms rather than raw hardware. "DeepSeek is not an inconsequential advance."

He then states what he sees as the real danger: "The day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation." Dwarkesh asks why, since open models can run on any accelerator. Huang says models optimized for one architecture run worse on others: "I am the evidence." Nvidia's success shows that models created on its stack run best on its stack. If the world's AI models come out of the box running best on non-American hardware, particularly in the Global South and the Middle East, he calls that bad for the US. Dwarkesh presses that the hypothetical still has China finding American vulnerabilities first, merely on Nvidia hardware. Huang agrees that would not be good, "so let's not let it happen." He argues that if the next few years are critical, all the world's AI models should be built on the American tech stack during those years, while acknowledging there is "no guarantee either way."

China, Part Three: Analogies, Lock-In, and "Conceding the Market"

The exchange becomes more heated. Huang asks why one layer of the AI industry should lose an entire market to benefit another layer, and why Dwarkesh is "so fixated" on the model layer. He says the US has 100x more compute than anywhere else, that Nvidia gives US labs first access to new technology, and that Vera Rubin is for the United States. Nvidia is an American company, he argues, so why not write a more balanced regulation that lets it win around the world instead of "giving up the world"?

Dwarkesh brings up Dario Amodei's analogy comparing chip sales to China with Boeing boasting that it supplied the missile casings for North Korean nukes, and then compares compute to enriched uranium. Huang calls this "lunacy," a "lousy analogy," and "illogical." He says the answer to misuse is dialogue with researchers, with China, and with all countries, combined with ensuring the US has "mountains" of Vera Rubin and Blackwell. He says the same logic could be applied to microprocessors, DRAM, or electricity. Dwarkesh replies that the US does in fact export-control advanced DRAM-making technology and chipmaking equipment.

Dwarkesh then raises lock-in. Tesla and Apple sold excellent products in China without preventing Chinese EVs and smartphones from coming to dominate. Huang answers, "We're not a car." Switching car brands is easy, but computing ecosystems like x86 and ARM are sticky and costly to replace. He says 50% of AI developers are in China and the US should not give them up. When Dwarkesh argues that American labs already use multiple accelerators and Chinese labs could too, Huang says Nvidia's share is growing, not shrinking, and rejects the premise that it would lose China anyway: "You're not talking to somebody who woke up a loser."

Dwarkesh identifies what looks like a contradiction: Huang claims Nvidia would beat Huawei because its chips are better, and also that China would do the same things without Nvidia. Huang says both are true, because "in the absence of a better choice, you'll take the only choice you have." When Dwarkesh equates "better" with more compute, Huang says it is better because it is easier to program and has a better ecosystem. "And of course we're going to send them compute. So what?" He lists the benefits as American technology leadership, developers working on the American stack, and that stack being the best choice as models diffuse worldwide. He points to the American telecommunications industry, which he says was "policied out of basically the world" to the point that the US no longer controls its own telecommunications, and calls the current approach "a little narrow-minded." He says China accounts for about 40% of the world's technology industry and that Nvidia once had a large share of that market and no longer does.

Dwarkesh asks Huang to acknowledge a cost: compute feeds powerful models, and if China had reached Mythos-level capability first and deployed it widely, that would have been very bad. Huang responds with a cost of his own. Conceding the second-largest market lets China build scale and its own ecosystem, so future models are optimized for a different stack. Because Chinese models are open, their standards could become superior as AI spreads. Dwarkesh says he trusts Nvidia's kernel engineers and techniques like distillation to prevent long-term lock-in. Huang replies that AI is "more than kernel optimization," and lists what he calls facts: China is the largest contributor to open-source software and to open models, and today those are built on Nvidia's stack.

He broadens the argument into a critique of fear-driven framing. He says scaring the country into treating AI as a nuclear bomb, or scaring people away from software engineering or radiology, does the US a disservice. On radiology he distinguishes a job from a task: "The job of a radiologist is patient care. The task is to read a scan." He makes an explicit prediction. When the US wants to export its technology and standards to India, the Middle East, Africa, and Southeast Asia, he wants to revisit this conversation, and he expects to show how this policy "caused the United States to concede the second largest market in the world for no good reason at all." He says nobody is advocating shipping everything to China. The US should always have the best and most technology first while also competing worldwide, which he says "requires some amount of nuance, some amount of maturity instead of absolutes."

In a final round, Dwarkesh argues that for Chinese chips to set global standards, 7nm chips would have to beat Nvidia's future 1.6nm chips on exported models. Huang answers with his own numbers. The transistor improvement from Hopper to Blackwell, three years apart, was "call it 75%," while Blackwell delivers 50x Hopper. "Moore's Law is dead," and architecture, networking, energy, and computer science matter most, which he says is why Nvidia bought Mellanox. He says export restrictions have already backfired by accelerating China's chip industry and pushing its AI ecosystem toward domestic architectures. It is "not too late," but it has already happened. He adds that China will not stay stuck at 7nm, and that 7nm versus 5nm is not a 10x difference.

Would Nvidia Go Back to Older Process Nodes?

Returning to supply constraints, Dwarkesh asks whether Nvidia might, before 2030, use spare N7 capacity to build a Hopper- or Ampere-class chip updated with modern numerics. Huang says it is not necessary. Each generation involves so much packaging, stacking, numerics, and system architecture work beyond the transistors that porting back to an older node is "a level of R&D that no one could afford." Nvidia can afford to lean forward, not back. Only in a thought experiment where the world would never gain more capacity would he go back to 7nm, and then "in a heartbeat."

Why Not Pursue Several Radically Different Architectures?

Dwarkesh asks why Nvidia doesn't hedge by running parallel projects, such as Cerebras-style wafer-scale chips, Dojo-style large packages, or designs without CUDA. Huang says Nvidia could, but has no better idea. Nvidia simulates these alternatives and finds them "provably worse." It would add a different accelerator only if the workload changed dramatically, meaning the shape of the market rather than the algorithms.

He says that is what has now happened with Groq, which Nvidia recently added and will fold into the CUDA ecosystem. His reasoning is that tokens have become valuable enough to support different price tiers. A couple of years ago tokens were nearly free. Now some customers, such as highly paid software engineers, would pay for much more responsive tokens. Nvidia therefore wants to extend the Pareto frontier with a faster-response, lower-throughput inference segment, betting that high average selling prices per token can make up for lower factory throughput. Apart from that, he says, more money would go into Nvidia's existing architecture.

What Nvidia Would Be Without Deep Learning

In the closing question, Dwarkesh asks what Nvidia would be doing if the deep learning revolution hadn't happened. Huang answers: accelerated computing, as always. The company's premise is that general-purpose computing has largely run its course for many workloads, and that pairing a GPU running CUDA with a CPU to offload kernels can speed applications by 100x to 200x. Early domains included computer graphics, molecular dynamics, seismic processing for energy discovery, and image processing, and he lists particle physics, fluids, and structured data processing among others. Even without AI, he says, Nvidia "would be very, very large."

He adds that without AI he "would be very sad," but that Nvidia's computing advances are what democratized deep learning by letting any researcher or student do science on a PC or a GeForce card. He notes that the opening part of every GTC keynote, covering computational lithography, quantum chemistry, and data processing, has nothing to do with AI. He ends by saying that much important work is not AI-related and that "tensors are not the only way that you compute."