Who Should AI Obey? Dwarkesh on the Anthropic–Pentagon Standoff

Open on YouTube ↗
Overview

The Department of War has designated Anthropic a supply chain risk after the company refused to remove its red lines against using its models for mass surveillance and autonomous weapons. In this narrated essay, Dwarkesh treats the episode as a warning shot. Today, he says, LLMs are probably not being used in mission-critical ways. Within 20 years, he expects AIs to make up 99% of the workforce in the military, the civilian government, and the private sector. They will be the robot armies, the superhumanly intelligent advisers to senators, presidents, and CEOs, and the police. If civilization is going to run on AI labor, then the fight between Anthropic and the Pentagon is an early look at some very large questions. The government's actions anger him, but he says he is glad the episode happened because it forces those questions into the open.

21 min read

His position has several parts that pull against each other. He thinks the government was entitled to stop using Anthropic's models, and he thinks it was wrong to try to destroy the company. He admires Anthropic's stand, and he thinks Anthropic's push for AI regulation has been naive. He also doubts that any single company's courage can solve the underlying problem.

The government's reasonable case, and where it went too far

Dwarkesh starts by granting the Pentagon's side of the argument. The military has every right to refuse Anthropic's models, and he calls its case for doing so "entirely reasonable," especially since terms like "mass surveillance" and "autonomous weapons" are ambiguous. He says that if he were Secretary of War he would probably have made the same call.

He gives a hypothetical to explain why. Suppose a future Democratic administration is negotiating Starlink access with Elon Musk, and Musk reserves the right to cut the military off if it fights an unjust war or one Congress has not authorized. That language sounds reasonable on its face. But a military cannot give a private contractor a kill switch over a technology it has come to depend on.

If the government had simply declined to do business with Anthropic, he says, he would not have written the essay. What it actually did was threaten to destroy Anthropic as a private business because the company would not sell on the government's terms. If the supply chain designation holds, companies like Amazon, Nvidia, Google, and Palantir would have to make sure Anthropic is not involved in any of their Pentagon work.

Why the designation becomes more dangerous over time

Dwarkesh thinks Anthropic could probably survive the designation today, because these companies can wall off the services they provide to the Department of War. That will not last. AI is not going to remain "some party trick addendum" to the products sold to the military. It will be part of how every product is built, maintained, and operated. His example: if Amazon provides a service to the Pentagon through AWS, and that service was built using Claude Code, does it count as a supply chain risk?

In a world of powerful, ubiquitous AI, he doubts Big Tech will be able to keep Claude separate from its Pentagon work. He suspects the Department of War has not thought through what follows. The Pentagon is a tiny fraction of these companies' revenue. Forced to choose between their AI provider and the government, wouldn't they drop the government? He asks what the plan is, then. Is the Pentagon going to coerce and bully every company that won't deal with it on exactly the terms it demands?

"Are we racing China just to adopt their system?"

Dwarkesh places the dispute within the AI race with China. The reason to win that race, as he frames it, is to keep the winner from being a government that believes there is no such thing as a truly private citizen or company, one that forces you to provide services you find morally objectionable and destroys your business if you refuse. He asks whether the US is racing the CCP only to adopt the cruelest parts of its system.

He expects the reply that the US government is democratically elected, so its demands are different. He rejects that. If a democratically elected leader tells you to help with mass surveillance, violate your fellow citizens' rights, or punish his political enemies, Dwarkesh refuses to accept that this is acceptable, much less that you have a duty to help.

Surveillance is already legal, just impractical, and AI removes the bottleneck

One of his central worries is that some forms of mass surveillance are already legal and only impractical to carry out. Under current law, he says, there is no Fourth Amendment protection for data you share with third parties: your bank, ISP, phone carrier, and email provider. The government reserves the right to buy and read this data in bulk without a warrant. What it has lacked is the ability to use it. No agency has the staff to watch every camera, read every message, and cross-reference every transaction.

AI removes that bottleneck, and he gives a back-of-envelope estimate. America has 100 million CCTV cameras. Good open-source multimodal models cost about 10 cents per million input tokens. At one frame every 10 seconds and roughly 1,000 tokens per frame, he calculates that about $30 billion would cover every camera in the country. He adds that a given level of AI capability gets about 10x cheaper each year, so the figure would fall to $3 billion the next year and $300 million the year after. By 2030, by his projection, monitoring "every single nook and cranny" of the country would cost less than remodeling the White House.

Once the technical capacity exists, he argues, the only thing standing between the US and an authoritarian state is the political expectation that "this is just not something we do here." That is why he calls Anthropic's refusal valuable and commendable: it helps set that norm and precedent.

The government's many levers

The episode also shows, according to Dwarkesh, that the government has far more leverage over private companies than was previously understood. At the time of recording, prediction markets gave a 74% chance that the supply chain restriction would be walked back. Even so, the president has many other ways to harass a company that resists. The federal government controls permitting for the power generation that new data centers need. It oversees antitrust enforcement. It holds contracts with the big tech companies Anthropic depends on for chips and funding, and it could make it a soft, unspoken condition of those contracts, or even an explicit one, that those companies stop working with Anthropic.

Why wider diffusion doesn't fix it

Some people argue the real problem is that there are only three leading AI companies, which gives the government a narrow target for pressure. Dwarkesh thinks wider diffusion would make things easier for the government, not harder.

His scenario: by 2027, the best frontier models, "the Claude 6s and then Gemini 5s," are capable of enabling mass surveillance, and their makers refuse to sell for that purpose. By late 2027 or certainly 2028, he expects open-source models to match whatever the frontier could do 12 months earlier. The government could then ignore the red lines drawn by Anthropic, Google, and OpenAI and use an open-source model that isn't the smartest in the world but is easily capable enough for surveillance work.

The more fundamental problem, he says, is that even if the three leading companies hold their line and are willing to be destroyed for it, the technology itself structurally favors mass surveillance and control of populations. Asked what to do about that, he says he doesn't have an answer. He would like to believe AI is symmetric, helping citizens check government power as much as it helps government monitor citizens, but he doesn't think it will work out that way. He describes AI as adding leverage to whatever assets and authority you already have. The government starts with a monopoly on violence, and it can now amplify that with extremely obedient employees who will never question orders.

Alignment to whom?

This leads to what Dwarkesh calls perhaps the most important and least discussed question about powerful AI. An army of extremely obedient employees is what successful alignment would look like: at a technical level, AI systems that reliably follow someone's intentions. It sounds frightening when described as mass surveillance or robot armies because a core question has gone unanswered, mostly because AIs have not yet been smart enough for it to matter. To what, or to whom, should AIs be aligned? When should a model defer to the model company, the end user, the law, or its own sense of morality?

He thinks it is understandable that model companies avoid the topic, since they don't want to advertise that they have complete control over the preferences and character of the entire future labor force, including the civilian government and the military. He sees the Anthropic–Pentagon dispute as an early version of what will become "the highest stakes negotiations in human history." Mass surveillance is not close to the most consequential thing one could do with AGI. It is simply an issue that came up early and shows the power dynamics that will follow.

"All lawful purposes" and the lesson of Snowden

The military's position is that the law already prohibits mass surveillance, so Anthropic should allow its models to be used for "all lawful purposes." Dwarkesh points to the 2013 Snowden revelations as a reason not to take that at face value. The NSA, which he notes is part of the Department of War, used the 2001 Patriot Act to justify collecting every phone record in America, on the theory that some subset might be relevant to a future investigation. It ran that program for years under a secret court order. No government will call what it is doing mass surveillance, he says; it will always find a euphemism. Accepting the Pentagon's assurance that red lines are unnecessary would be "incredibly naive."

He then considers the military's perspective. In the future, every soldier in the field, every bureaucrat and analyst, and even the generals may be AIs, supplied on the current trajectory by a private company. He guesses that Pete Hegseth isn't thinking about generative AI in those terms yet, but expects the stakes to become obvious eventually, the way the stakes of nuclear weapons became obvious to everyone after 1945. From that vantage point, a vendor that reserves the right to cut you off for violating its embedded values and terms of service is alarming. It could be worse still if a future Claude had its own sense of right and wrong and simply refused orders it judged to violate its terms.

The case for models with their own morality

Dwarkesh admits that letting a model follow its own values sounds like the opening of every sci-fi dystopia, and that a model acting on its own values is close to the literal definition of misalignment. Even so, he argues that situations like this show why models need a robust moral sense of their own, because many of history's worst catastrophes were avoided when people on the ground refused orders.

He gives two examples. In 1989 the Berlin Wall fell, and the East German regime collapsed, because border guards refused to fire on fellow citizens trying to escape. His strongest example is Stanislav Petrov, a Soviet lieutenant colonel on duty at a nuclear early-warning station, whose sensors reported that the United States had launched five ICBMs at the Soviet Union. Petrov judged it a false alarm and broke protocol by not alerting his superiors. Dwarkesh says that if he had, Soviet High Command would probably have retaliated and hundreds of millions of people would have died.

The difficulty, he says, is that one person's virtue is another person's misalignment. Who decides what moral convictions these AIs will hold, and on whose behalf they may break the chain of command or even the law? Who writes the "model constitution" for the entities that will run civilization? He likes an idea Dario Amodei described on his podcast: companies publish their constitutions, outsiders compare and critique them, and the result is a soft incentive for every company to adopt the best elements of each. What he considers very dangerous is the government dictating what values AI systems must have.

Where Anthropic has been naive: regulation

Here Dwarkesh turns critical of the AI safety community and of Anthropic in particular. He thinks the safety community has been naive in pushing for regulations that give governments this kind of power, and that Anthropic has been especially so, for example in opposing the moratorium on state AI laws. He calls this ironic, since what Anthropic advocates would give the government even more capacity for the "thuggish political pressure" it is now experiencing.

He grants that the underlying logic makes sense. Many measures that make AI development safer impose real costs and can slow a lab relative to competitors: investing in alignment over raw capabilities, enforcing safeguards against bioweapons and cyberattacks, and eventually slowing recursive self-improvement to a pace where humans stay in the loop rather than setting off an uncontrolled singularity. These safeguards mean little unless the whole industry adopts them, so there is a genuine collective action problem. He quotes Anthropic's frontier safety roadmap: "At the most advanced capability levels and risks, the appropriate governance analogy may be closer to nuclear energy or financial regulation than to today's approach to software." He reads this as envisioning something like a Nuclear Regulatory Commission or SEC for AI.

He says he cannot imagine such a framework not being abused by a would-be despot. Terms like "catastrophic risk," "threats to national security," and "autonomy risk" are so vague that adopting them amounts to "handing a fully loaded bazooka to a future power-hungry leader." He shows how they could be turned around. A model that tells users the government's tariff policy is misguided becomes a "deceptive" or "manipulative" model that can't be deployed. A model that won't help with mass surveillance becomes a national security threat. Any model that refuses government orders on moral grounds becomes an "autonomy risk."

As evidence, he points to the statutes already being used against Anthropic, neither of which has anything to do with AI. One is the supply chain risk authority from a 2018 defense bill, meant to keep Huawei components out of US military hardware. The other is the Defense Production Act, a 1950s law meant to help Truman keep steel mills and ammunition factories running during the Korean War. If the government is already stretching laws like these, he asks, do we want to give it a regulatory apparatus built specifically for AI, the thing it will most want to control?

He restates what he sees as the stakes. AI will be the substrate of future civilization, the way private citizens take part in commerce, get information about the world, and get advice on how to use their power as voters and capital holders. By his reckoning, mass surveillance is "like the 10th scariest thing" a government could do with control over those systems.

The strongest counterargument: no regulation at all?

Dwarkesh names what he considers the strongest objection to his view: are we really going to leave the most powerful technology in human history unregulated? Even if that were ideal, the government will clearly regulate AI in some form, and coordination could genuinely reduce some risks. His answer is that he doesn't know how to design a regulatory apparatus that wouldn't become a huge temptation for the government to control a civilization built on AI, or to requisition blindly obedient soldiers, censors, and apparatchiks. Some regulation may be inevitable, but he thinks a wholesale government takeover of the technology would be a terrible idea.

The nuclear analogy and why he rejects it

He addresses an argument Ben Thompson made in a post the previous week. Thompson noted that people like Dario have compared AI to nuclear weapons when arguing for extra controls, and followed the analogy through: "If nuclear weapons were developed by a private company, the US would absolutely be incentivized to destroy that company." Dwarkesh says safety-minded people have made similar points. Leopold Aschenbrenner, a former guest and, as Dwarkesh discloses, a good friend, wrote in his memo Situational Awareness: "I find it an insane proposition that the US government will let a random SF startup develop superintelligence. Imagine if we had developed atomic bombs by letting Uber just improvise."

His reply to both is that they are right that it is crazy to leave this technology to private companies, but handing the authority to the government is not an improvement. Nobody is qualified to be the steward of superintelligence. Private companies not being ideal doesn't make the Pentagon or the White House ideal. He concedes that if a single private company were the only entity able to build nuclear weapons, the government would not let it hold a veto over their use. He gives two reasons the analogy fails for AI.

First, AI is not a self-contained, single-purpose weapon like a bomb. It resembles industrialization itself, a general transformation of the economy with thousands of applications. Applying Thompson's or Aschenbrenner's logic to the Industrial Revolution would have meant the government could requisition any factory, destroy any business, and punish anyone who refused, which is not how free societies handled industrialization. People will object that AI will enable superweapons: superhuman hackers, superhuman bioweapons researchers, fully autonomous robot armies. Dwarkesh says the same was true of industrialization, which, from the perspective of 17th-century Europeans, produced chemical weapons, aerial bombardment, and nuclear weapons. Societies didn't respond by giving the government total control of industrialization. They banned and regulated specific weaponizable end uses. He wants AI handled the same way: regulate destructive uses such as launching cyberattacks, things that should be illegal even when a human does them, and pass laws that constrain how the government itself can use the technology, for example to build an AI-powered surveillance state.

Second, no single company has a monopoly here. There are many frontier labs, so the government's claim that it had to override one company's property rights to get a critical national security capability is, in his words, extremely weak. It could have signed a voluntary contract with one of Anthropic's half-dozen competitors. He adds a condition to his position. If only one entity ever becomes capable of building robot armies and superhuman hackers, with a lead insurmountable enough that it might take over the world, he agrees it would be unacceptable for that entity to be a private company. His crux with those who say AI is too powerful for private hands is that he expects the technology to be highly multipolar, with many competing companies at every layer of the supply chain.

Why corporate courage isn't enough

That same multipolarity is why Dwarkesh doesn't think individual acts of corporate courage can solve the problem. AI structurally favors many authoritarian applications. Even if Anthropic and the next two companies after it all refused to enable mass surveillance, within 12 months "everybody and their mother" will be able to train a model as good as today's frontier, and some vendor will be willing to help. The only way to preserve a free society, he argues, is through laws and norms, established through the political system, that make it unacceptable for the government to use AI for mass censorship, surveillance, and control. He compares this to the norm the world set after World War II against using nuclear weapons to wage war.

An unsettled position

He ends by stressing how uncertain he is. He calls these extremely confusing and difficult questions, says he changed his mind back and forth several times while writing the essay, and reserves the right to change it again. He thinks changing one's mind as AI progresses is essential, and that this is the point of conversation and debate. He predicts people will someday look back on this period the way we look back on the Enlightenment: people debating big questions just before huge technological, social, and political revolutions, with some thinkers getting a few of those questions right in ways later generations still benefit from. He closes by saying we owe it to the future to at least try to think through the new questions AI raises.