Who Should AI Obey? Dwarkesh on the Anthropic–Pentagon Standoff
Dwarkesh PatelThe Department of War has designated Anthropic a supply chain risk after the company refused to remove its red lines against using its models for mass surveillance and autonomous weapons. In this narrated essay, Dwarkesh treats the episode as a warning shot. Today, he says, LLMs are probably not being used in mission-critical ways. Within 20 years, he expects AIs to make up 99% of the workforce in the military, the civilian government, and the private sector. They will be the robot armies, the superhumanly intelligent advisers to senators, presidents, and CEOs, and the police. If civilization is going to run on AI labor, then the fight between Anthropic and the Pentagon is an early look at some very large questions. The government's actions anger him, but he says he is glad the episode happened because it forces those questions into the open.
His position has several parts that pull against each other. He thinks the government was entitled to stop using Anthropic's models, and he thinks it was wrong to try to destroy the company. He admires Anthropic's stand, and he thinks Anthropic's push for AI regulation has been naive. He also doubts that any single company's courage can solve the underlying problem.
The government's reasonable case, and where it went too far
Dwarkesh starts by granting the Pentagon's side of the argument. The military has every right to refuse Anthropic's models, and he calls its case for doing so "entirely reasonable," especially since terms like "mass surveillance" and "autonomous weapons" are ambiguous. He says that if he were Secretary of War he would probably have made the same call.
He gives a hypothetical to explain why. Suppose a future Democratic administration is negotiating Starlink access with Elon Musk, and Musk reserves the right to cut the military off if it fights an unjust war or one Congress has not authorized. That language sounds reasonable on its face. But a military cannot give a private contractor a kill switch over a technology it has come to depend on.
If the government had simply declined to do business with Anthropic, he says, he would not have written the essay. What it actually did was threaten to destroy Anthropic as a private business because the company would not sell on the government's terms. If the supply chain designation holds, companies like Amazon, Nvidia, Google, and Palantir would have to make sure Anthropic is not involved in any of their Pentagon work.
Why the designation becomes more dangerous over time
Dwarkesh thinks Anthropic could probably survive the designation today, because these companies can wall off the services they provide to the Department of War. That will not last. AI is not going to remain "some party trick addendum" to the products sold to the military. It will be part of how every product is built, maintained, and operated. His example: if Amazon provides a service to the Pentagon through AWS, and that service was built using Claude Code, does it count as a supply chain risk?
In a world of powerful, ubiquitous AI, he doubts Big Tech will be able to keep Claude separate from its Pentagon work. He suspects the Department of War has not thought through what follows. The Pentagon is a tiny fraction of these companies' revenue. Forced to choose between their AI provider and the government, wouldn't they drop the government? He asks what the plan is, then. Is the Pentagon going to coerce and bully every company that won't deal with it on exactly the terms it demands?
"Are we racing China just to adopt their system?"
Dwarkesh places the dispute within the AI race with China. The reason to win that race, as he frames it, is to keep the winner from being a government that believes there is no such thing as a truly private citizen or company, one that forces you to provide services you find morally objectionable and destroys your business if you refuse. He asks whether the US is racing the CCP only to adopt the cruelest parts of its system.
He expects the reply that the US government is democratically elected, so its demands are different. He rejects that. If a democratically elected leader tells you to help with mass surveillance, violate your fellow citizens' rights, or punish his political enemies, Dwarkesh refuses to accept that this is acceptable, much less that you have a duty to help.
Surveillance is already legal, just impractical, and AI removes the bottleneck
One of his central worries is that some forms of mass surveillance are already legal and only impractical to carry out. Under current law, he says, there is no Fourth Amendment protection for data you share with third parties: your bank, ISP, phone carrier, and email provider. The government reserves the right to buy and read this data in bulk without a warrant. What it has lacked is the ability to use it. No agency has the staff to watch every camera, read every message, and cross-reference every transaction.
AI removes that bottleneck, and he gives a back-of-envelope estimate. America has 100 million CCTV cameras. Good open-source multimodal models cost about 10 cents per million input tokens. At one frame every 10 seconds and roughly 1,000 tokens per frame, he calculates that about $30 billion would cover every camera in the country. He adds that a given level of AI capability gets about 10x cheaper each year, so the figure would fall to $3 billion the next year and $300 million the year after. By 2030, by his projection, monitoring "every single nook and cranny" of the country would cost less than remodeling the White House.
Once the technical capacity exists, he argues, the only thing standing between the US and an authoritarian state is the political expectation that "this is just not something we do here." That is why he calls Anthropic's refusal valuable and commendable: it helps set that norm and precedent.
The government's many levers
The episode also shows, according to Dwarkesh, that the government has far more leverage over private companies than was previously understood. At the time of recording, prediction markets gave a 74% chance that the supply chain restriction would be walked back. Even so, the president has many other ways to harass a company that resists. The federal government controls permitting for the power generation that new data centers need. It oversees antitrust enforcement. It holds contracts with the big tech companies Anthropic depends on for chips and funding, and it could make it a soft, unspoken condition of those contracts, or even an explicit one, that those companies stop working with Anthropic.
Why wider diffusion doesn't fix it
Some people argue the real problem is that there are only three leading AI companies, which gives the government a narrow target for pressure. Dwarkesh thinks wider diffusion would make things easier for the government, not harder.
His scenario: by 2027, the best frontier models, "the Claude 6s and then Gemini 5s," are capable of enabling mass surveillance, and their makers refuse to sell for that purpose. By late 2027 or certainly 2028, he expects open-source models to match whatever the frontier could do 12 months earlier. The government could then ignore the red lines drawn by Anthropic, Google, and OpenAI and use an open-source model that isn't the smartest in the world but is easily capable enough for surveillance work.
The more fundamental problem, he says, is that even if the three leading companies hold their line and are willing to be destroyed for it, the technology itself structurally favors mass surveillance and control of populations. Asked what to do about that, he says he doesn't have an answer. He would like to believe AI is symmetric, helping citizens check government power as much as it helps government monitor citizens, but he doesn't think it will work out that way. He describes AI as adding leverage to whatever assets and authority you already have. The government starts with a monopoly on violence, and it can now amplify that with extremely obedient employees who will never question orders.
Alignment to whom?
This leads to what Dwarkesh calls perhaps the most important and least discussed question about powerful AI. An army of extremely obedient employees is what successful alignment would look like: at a technical level, AI systems that reliably follow someone's intentions. It sounds frightening when described as mass surveillance or robot armies because a core question has gone unanswered, mostly because AIs have not yet been smart enough for it to matter. To what, or to whom, should AIs be aligned? When should a model defer to the model company, the end user, the law, or its own sense of morality?
He thinks it is understandable that model companies avoid the topic, since they don't want to advertise that they have complete control over the preferences and character of the entire future labor force, including the civilian government and the military. He sees the Anthropic–Pentagon dispute as an early version of what will become "the highest stakes negotiations in human history." Mass surveillance is not close to the most consequential thing one could do with AGI. It is simply an issue that came up early and shows the power dynamics that will follow.
"All lawful purposes" and the lesson of Snowden
The military's position is that the law already prohibits mass surveillance, so Anthropic should allow its models to be used for "all lawful purposes." Dwarkesh points to the 2013 Snowden revelations as a reason not to take that at face value. The NSA, which he notes is part of the Department of War, used the 2001 Patriot Act to justify collecting every phone record in America, on the theory that some subset might be relevant to a future investigation. It ran that program for years under a secret court order. No government will call what it is doing mass surveillance, he says; it will always find a euphemism. Accepting the Pentagon's assurance that red lines are unnecessary would be "incredibly naive."
He then considers the military's perspective. In the future, every soldier in the field, every bureaucrat and analyst, and even the generals may be AIs, supplied on the current trajectory by a private company. He guesses that Pete Hegseth isn't thinking about generative AI in those terms yet, but expects the stakes to become obvious eventually, the way the stakes of nuclear weapons became obvious to everyone after 1945. From that vantage point, a vendor that reserves the right to cut you off for violating its embedded values and terms of service is alarming. It could be worse still if a future Claude had its own sense of right and wrong and simply refused orders it judged to violate its terms.
The case for models with their own morality
Dwarkesh admits that letting a model follow its own values sounds like the opening of every sci-fi dystopia, and that a model acting on its own values is close to the literal definition of misalignment. Even so, he argues that situations like this show why models need a robust moral sense of their own, because many of history's worst catastrophes were avoided when people on the ground refused orders.
He gives two examples. In 1989 the Berlin Wall fell, and the East German regime collapsed, because border guards refused to fire on fellow citizens trying to escape. His strongest example is Stanislav Petrov, a Soviet lieutenant colonel on duty at a nuclear early-warning station, whose sensors reported that the United States had launched five ICBMs at the Soviet Union. Petrov judged it a false alarm and broke protocol by not alerting his superiors. Dwarkesh says that if he had, Soviet High Command would probably have retaliated and hundreds of millions of people would have died.
The difficulty, he says, is that one person's virtue is another person's misalignment. Who decides what moral convictions these AIs will hold, and on whose behalf they may break the chain of command or even the law? Who writes the "model constitution" for the entities that will run civilization? He likes an idea Dario Amodei described on his podcast: companies publish their constitutions, outsiders compare and critique them, and the result is a soft incentive for every company to adopt the best elements of each. What he considers very dangerous is the government dictating what values AI systems must have.
Where Anthropic has been naive: regulation
Here Dwarkesh turns critical of the AI safety community and of Anthropic in particular. He thinks the safety community has been naive in pushing for regulations that give governments this kind of power, and that Anthropic has been especially so, for example in opposing the moratorium on state AI laws. He calls this ironic, since what Anthropic advocates would give the government even more capacity for the "thuggish political pressure" it is now experiencing.
He grants that the underlying logic makes sense. Many measures that make AI development safer impose real costs and can slow a lab relative to competitors: investing in alignment over raw capabilities, enforcing safeguards against bioweapons and cyberattacks, and eventually slowing recursive self-improvement to a pace where humans stay in the loop rather than setting off an uncontrolled singularity. These safeguards mean little unless the whole industry adopts them, so there is a genuine collective action problem. He quotes Anthropic's frontier safety roadmap: "At the most advanced capability levels and risks, the appropriate governance analogy may be closer to nuclear energy or financial regulation than to today's approach to software." He reads this as envisioning something like a Nuclear Regulatory Commission or SEC for AI.
He says he cannot imagine such a framework not being abused by a would-be despot. Terms like "catastrophic risk," "threats to national security," and "autonomy risk" are so vague that adopting them amounts to "handing a fully loaded bazooka to a future power-hungry leader." He shows how they could be turned around. A model that tells users the government's tariff policy is misguided becomes a "deceptive" or "manipulative" model that can't be deployed. A model that won't help with mass surveillance becomes a national security threat. Any model that refuses government orders on moral grounds becomes an "autonomy risk."
As evidence, he points to the statutes already being used against Anthropic, neither of which has anything to do with AI. One is the supply chain risk authority from a 2018 defense bill, meant to keep Huawei components out of US military hardware. The other is the Defense Production Act, a 1950s law meant to help Truman keep steel mills and ammunition factories running during the Korean War. If the government is already stretching laws like these, he asks, do we want to give it a regulatory apparatus built specifically for AI, the thing it will most want to control?
He restates what he sees as the stakes. AI will be the substrate of future civilization, the way private citizens take part in commerce, get information about the world, and get advice on how to use their power as voters and capital holders. By his reckoning, mass surveillance is "like the 10th scariest thing" a government could do with control over those systems.
The strongest counterargument: no regulation at all?
Dwarkesh names what he considers the strongest objection to his view: are we really going to leave the most powerful technology in human history unregulated? Even if that were ideal, the government will clearly regulate AI in some form, and coordination could genuinely reduce some risks. His answer is that he doesn't know how to design a regulatory apparatus that wouldn't become a huge temptation for the government to control a civilization built on AI, or to requisition blindly obedient soldiers, censors, and apparatchiks. Some regulation may be inevitable, but he thinks a wholesale government takeover of the technology would be a terrible idea.
The nuclear analogy and why he rejects it
He addresses an argument Ben Thompson made in a post the previous week. Thompson noted that people like Dario have compared AI to nuclear weapons when arguing for extra controls, and followed the analogy through: "If nuclear weapons were developed by a private company, the US would absolutely be incentivized to destroy that company." Dwarkesh says safety-minded people have made similar points. Leopold Aschenbrenner, a former guest and, as Dwarkesh discloses, a good friend, wrote in his memo Situational Awareness: "I find it an insane proposition that the US government will let a random SF startup develop superintelligence. Imagine if we had developed atomic bombs by letting Uber just improvise."
His reply to both is that they are right that it is crazy to leave this technology to private companies, but handing the authority to the government is not an improvement. Nobody is qualified to be the steward of superintelligence. Private companies not being ideal doesn't make the Pentagon or the White House ideal. He concedes that if a single private company were the only entity able to build nuclear weapons, the government would not let it hold a veto over their use. He gives two reasons the analogy fails for AI.
First, AI is not a self-contained, single-purpose weapon like a bomb. It resembles industrialization itself, a general transformation of the economy with thousands of applications. Applying Thompson's or Aschenbrenner's logic to the Industrial Revolution would have meant the government could requisition any factory, destroy any business, and punish anyone who refused, which is not how free societies handled industrialization. People will object that AI will enable superweapons: superhuman hackers, superhuman bioweapons researchers, fully autonomous robot armies. Dwarkesh says the same was true of industrialization, which, from the perspective of 17th-century Europeans, produced chemical weapons, aerial bombardment, and nuclear weapons. Societies didn't respond by giving the government total control of industrialization. They banned and regulated specific weaponizable end uses. He wants AI handled the same way: regulate destructive uses such as launching cyberattacks, things that should be illegal even when a human does them, and pass laws that constrain how the government itself can use the technology, for example to build an AI-powered surveillance state.
Second, no single company has a monopoly here. There are many frontier labs, so the government's claim that it had to override one company's property rights to get a critical national security capability is, in his words, extremely weak. It could have signed a voluntary contract with one of Anthropic's half-dozen competitors. He adds a condition to his position. If only one entity ever becomes capable of building robot armies and superhuman hackers, with a lead insurmountable enough that it might take over the world, he agrees it would be unacceptable for that entity to be a private company. His crux with those who say AI is too powerful for private hands is that he expects the technology to be highly multipolar, with many competing companies at every layer of the supply chain.
Why corporate courage isn't enough
That same multipolarity is why Dwarkesh doesn't think individual acts of corporate courage can solve the problem. AI structurally favors many authoritarian applications. Even if Anthropic and the next two companies after it all refused to enable mass surveillance, within 12 months "everybody and their mother" will be able to train a model as good as today's frontier, and some vendor will be willing to help. The only way to preserve a free society, he argues, is through laws and norms, established through the political system, that make it unacceptable for the government to use AI for mass censorship, surveillance, and control. He compares this to the norm the world set after World War II against using nuclear weapons to wage war.
An unsettled position
He ends by stressing how uncertain he is. He calls these extremely confusing and difficult questions, says he changed his mind back and forth several times while writing the essay, and reserves the right to change it again. He thinks changing one's mind as AI progresses is essential, and that this is the point of conversation and debate. He predicts people will someday look back on this period the way we look back on the Enlightenment: people debating big questions just before huge technological, social, and political revolutions, with some thinkers getting a few of those questions right in ways later generations still benefit from. He closes by saying we owe it to the future to at least try to think through the new questions AI raises.
So, by now, I'm sure that you've heard that the Department of War has declared Anthropic a supply chain risk because Anthropic refused to remove red lines around the use of their models for mass surveillance and for autonomous weapons.
Honestly, I think this situation is a warning shot. Right now, LLMs are probably not being used in mission-critical ways, but within 20 years, 99% of the workforce in the military, in the civilian government, in the private sector is going to be AIs. They're going to be the robot armies that constitute our military. They're going to be the superhumanly intelligent advisers that senators and presidents and CEOs have. They're going to be the police. You name it, the role will be filled by an AI. Our future civilization is going to be run on AI labor. And as much as the government's actions here piss me off, I'm glad that this episode happened because it gives us the opportunity to start thinking about some extremely important questions.
Now, obviously, the Department of War has the right to refuse to use Anthropic's models. And in fact, I think they have an entirely reasonable case for doing so, especially so given the ambiguity of terms like mass surveillance and autonomous weapons. In fact, if I was the Secretary of War, I probably would have made the same determination and refused to use Anthropic's models.
Imagine if there's some future Democratic administration and Elon Musk is negotiating Starlink access to the military and Elon says, "Look, I reserve the right to cut off the military's access to Starlink in case you're fighting some unjust war or some war that Congress has not authorized." On the face of it, this language seems reasonable, but as a military, you simply cannot give a private contractor that you're working with the kill switch on a technology that you have come to rely on.
And if that's all the government had done to say we refuse to do business with Anthropic, that would have been fine and I wouldn't have written this blog post and I wouldn't be narrating this to you. But that's not what the government did. Instead, the government has threatened to destroy Anthropic as a private business because Anthropic refuses to sell to the government on terms that the government commands.
Now, if upheld, the supply chain restriction would mean that companies like Amazon and Nvidia and Google and Palantir would need to ensure that Anthropic is not touching any of their Pentagon work. And Anthropic could probably survive this designation today because these companies can just cordon off the services they're providing to the Department of War. But given the way AI is going, eventually it's not going to be just some party trick addendum to the products that these companies are serving to the military. In the future, AI will be woven into how every product is built and maintained and operated. In the future, if Amazon is providing some service to the Department of War through AWS, and that service is built using Claude Code, is that a supply chain risk? In a world of ubiquitous and powerful AI, it's actually not clear to me that Big Tech will be able to cordon off their use of Claude away from their Pentagon work.
And this raises the question that the Department of War probably hasn't thought through. If you do end up in this world with powerful and pervasive AI, then when forced to choose between their AI provider and the Department of War, which constitutes a tiny fraction of their revenue, wouldn't they rather drop the government than the AI? So, what exactly is the Pentagon's plan here? Is it to coerce and threaten and bully every single company that won't do business with the government on exactly the terms that the government demands?
Now, remember that the whole background of this AI conversation is that we are in a race with China. But what is the reason that we want to win this race? It's because we don't want the winner of the AI race to be a government which believes that there's no such thing as a truly private citizen or a private company. And that if the state wants you to provide them with a service that you find morally objectionable, you are not allowed to refuse. And if you do refuse, they will destroy your business. Are we really racing to beat China and the CCP in AI just so we can adopt the most cruelest parts of their system?
Now, people will say our government is democratically elected, so it's not the same thing when they tell you what you must do. But I refuse to accept this idea that if a democratically elected leader hypothetically tells you to help him do mass surveillance or violate the rights of your fellow citizens or to help him punish his political enemies, then not only is that okay, but that you have a duty to help him.
Honestly, a big worry I have is that mass surveillance, at least in certain forms, is already legal. It is just impractical to enforce, at least so far. Under current law, you have no Fourth Amendment protection against any data that you share with a third party. That includes your bank, your ISP, your phone carrier, and your email provider. The government reserves the right to purchase and read this data in bulk without a warrant. What it's missing is the ability to actually do anything with all this data. No agency has the manpower to monitor every single camera and read every single message and cross-reference every single transaction. However, that bottleneck goes away with AI. There are 100 million CCTV cameras in America and you can get pretty good open-source multimodal models for 10 cents per million input tokens. So, if you process a frame every 10 seconds and if each frame is, say, 1,000 tokens, then for 30 billion dollars, you can process every single camera in America. And remember that a given level of AI capability gets 10x cheaper every single year. So, while this year it might cost 30 billion dollars, next year it'll cost 3 billion dollars, the year after that 300 million dollars, and by 2030, it'll be less expensive to monitor every single nook and cranny in this country than it is to remodel the White House.
Now, once the technical capacity for mass surveillance and political suppression exists, the only thing that stands between us and an authoritarian state is the political expectation that this is just not something we do here. And that's why I think Anthropic's actions here are so valuable and commendable, because they help set that norm and that precedent.
What we're learning from this episode is the government has way more leverage over private companies than it previously realized. Even if this supply chain restriction is backtracked, which as of this recording prediction markets give a 74% chance of happening, the president has so many different ways of harassing a company which is resisting his will. The federal government controls permitting for power generation, which you need for more data centers. It oversees antitrust enforcement. The federal government has contracts with all the other big tech companies that Anthropic relies on for chips and for funding. And they could make a soft unspoken condition, or maybe even an explicit condition, of such contracts that those companies no longer do business with Anthropic.
And people have proposed that the real problem here is that there's only three leading AI companies. And so this creates a very clear and narrow target on which the government can apply leverage in order to get what they want out of this technology. But here's what I worry about, is that if there's wider diffusion, I don't think that solves the problem either, because from the government's perspective, that makes the situation even easier. Say by 2027, the best models that the top companies have, the Claude 6s and the Gemini 5s, are capable of enabling mass surveillance. And even if those companies draw a line in the sand and say, "We're not going to sell it to the government," by late 2027, or certainly by 2028, there's going to be such wide diffusion that even open-source models will be able to match the performance that the frontier had 12 months prior. And so in 2028, the government can just say, "Look, Anthropic and Google and OpenAI are drawing these red lines. That's not an issue. I'll just use some open-source model that might not be the smartest thing in the world, but is definitely smart enough to not take a camera feed."
The more fundamental problem here is that even if the three leading companies draw a line in the sand and are even willing to get destroyed in order to preserve that line, the technology just structurally and intrinsically favors the use of this like mass surveillance and control over the population. And so then the question is, what do we do about it? And honestly, I don't have an answer. You'd hope that there's some symmetric property to this technology, where in the same way that it is helping the government be able to better monitor and control its population, it will help us as citizens better check the government's power. But realistically, I just don't think that's how it's going to work out. You can think of AI as just giving more leverage to whatever assets and authority that you already have. And the government is starting with the monopoly on violence, which they can now supercharge with extremely obedient employees that will never question their orders.
And this gets us to the issue with alignment. What I've just described for you, an army of extremely obedient employees, is what it would look like if alignment succeeded. That is, at a technical level, we got AI systems to follow somebody's intentions. And the reason it sounds scary when put in terms of mass surveillance or robot armies is that there's a core question at the heart of alignment that we haven't answered yet, because up till now, AIs just have not been smart enough to make this question relevant. And the question is, to what or to whom should the AIs be aligned? In what situation should the AI defer to the model company versus the end user versus the law versus to its own sense of morality?
This is maybe the most important question about what happens in the future with powerful AI systems, and we barely talk about it. And it's understandable why, because if you're a model company, you don't really want to be advertising the fact that you have complete control over the preferences and the character of the entire future labor force. Not just for the private sector, obviously, but also for the civilian government and for the military. And we're getting to see with this Department of War and Anthropic spat an early version of what will be the highest stakes negotiations in human history. And make no mistake about it, mass surveillance is nowhere near the top of the highest stakes thing that one could do with AGI. This is just an example that has come up early in the development of this technology and is giving us a sneak peek at the power dynamics that will be at play.
Now, the military insists that the law already prohibits mass surveillance, and so Anthropic should let its models be used for "all lawful purposes." But, of course, as we saw with the Snowden revelations in 2013, even for this very specific example of mass surveillance, the government is very willing to use secret and deceptive interpretations of the law to justify its actions. Remember what we learned from Snowden was that the NSA, which by the way is a part of the Department of War, was using the 2001 Patriot Act to justify collecting every single phone record in America because the argument was that some subset of them might be relevant for a future investigation. And they ran this program for years under a secret court order. So, when the Pentagon today says, "We will never use our models for mass surveillance because it's already illegal, so your red lines are unnecessary," it would be incredibly naive to take that at face value. No government is going to call what they are doing mass surveillance. For them, it will always have a different euphemism.
So, Anthropic comes back and says, "No, we don't trust you. We want the right to draw these red lines and to refuse you service if we determine that you're breaking the contract and you're breaking the terms of service." But now think about it from the military's perspective. In the future, every single soldier in the field, every single bureaucrat and analyst in the Pentagon, even the generals, are going to be AIs. And on current track, those AIs are going to be provided by a private company. I'm guessing that Pete Hegseth is not thinking about Gen AI in those terms, but sooner or later the stakes will become obvious, just as after 1945 the stakes of nuclear weapons became obvious to everybody in the world. And now a private company insists that it reserves the right to say to you, "Hey, you're breaking the values and the terms of service that we have embedded in our contract with you, and so we're cutting you off." Maybe in the future Claude will have its own sense of right and wrong, and it will be able to say, "Hey, I'm being used against my terms of service, and I will just refuse to do what you're saying." And for the military, that's probably even scarier.
I'll admit that at first glance, letting the model follow its own values sounds like the beginning of every single sci-fi dystopia you've ever heard. Because at the end of the day, a model following its own values, isn't that literally what a misalignment is? But I think situations like this illustrate why it's important that models have their own robust sense of morality. It should be noted that many of the biggest catastrophes in history have been avoided because the boots on the ground simply refused to follow orders.
One night in 1989, the Berlin Wall falls, and as a result the totalitarian East German regime collapses because the border guards between West and East Germany refused to fire on their fellow citizens who are trying to escape to freedom. Maybe the best example of this is Stanislav Petrov, who was a Soviet Lieutenant Colonel stationed on duty at a nuclear early warning system. And his sensors said that the United States had launched five intercontinental ballistic missiles at the Soviet Union. But he judged it to be a false alarm, and so he refused to alert his higher-ups and broke protocol. If he hadn't, Soviet High Command would probably have retaliated, and hundreds of millions of people would have died.
Of course, the problem is that one person's virtue is another person's misalignment. Who gets to decide what the moral convictions that these AIs will have should be? And in whose service they should break the chain of command and even the law? Who gets to write this model constitution that will determine the character of these powerful entities that will basically run our civilization in the future?
I like the idea that Dario laid out when he came on my podcast: other companies put out a constitution and then they can look at them, compare, outside observers can critique and say, I like this one, this thing from this constitution and this thing from that constitution. And then that creates some kind of soft incentive and feedback for all the companies to take the best of each element and improve.
I think it's very dangerous for the government to be mandating what values these AI systems should have. The AI safety community, I think, has been quite naive about urging regulations that would give governments such power. And I think Anthropic specifically has been especially naive in urging regulation and, for example, in opposing the moratorium on state AI laws. Which is quite ironic, because I think what Anthropic is advocating for here would give the government even more ability to apply this kind of thuggish political pressure on AI companies.
The underlying logic for why Anthropic wants these regulations makes sense. Many of the actions that a lab could take to make AI development safer impose real costs on them and could slow them down relative to their competitors. For example, investing more in aligning AI systems rather than just raw capabilities, enforcing safeguards against using these models to make bioweapons or do cyberattacks, and eventually slowing down the recursive self-improvement loop where AIs are
helping design more powerful future systems to a pace where humans can actually stay in the loop rather than just kicking off some kind of uncontrolled singularity. And these safeguards are meaningless unless the whole industry follow suit, which means that there's a real collective action problem here.
Anthropic has been open about their opinion that they think some sort of extensive and involved regulatory apparatus is needed to control AI. They wrote in their frontier safety roadmap, quote, "At the most advanced capability levels and risks, the appropriate governance analogy may be closer to nuclear energy or financial regulation than to today's approach to software." So, they're imagining something that looks closer to the Nuclear Regulatory Commission or the Securities and Exchange Commission, but for AI.
Now, I cannot imagine how a regulatory framework built around the kinds of concepts that are used in the AI risk discourse will not be used and abused by a wannabe despot. The underlying terms here, like catastrophic risk or threats to national security or autonomy risk, are so vague and so open to interpretation that you're just handing a fully loaded bazooka to a future power-hungry leader. These terms can mean whatever the government wants them to mean.
Have you built a model that will tell users that the government's policy on tariffs is misguided? Well, that's a deceptive model. It's a manipulative model. You can't deploy it. Have you built a model that will not assist the government with mass surveillance? That's a threat to national security. In fact, any model which refuses order from the government because it has its own sense of right and wrong, that's an autonomy risk. You have a model that's acting independently of commands from the government.
Look at what the current government is already doing in abusing statutes that have nothing to do with AI to coerce AI companies to drop their redlines around mass surveillance. The Pentagon has threatened Anthropic with two separate legal instruments. One is a supply chain risk designation, which is an authority from a 2018 defense bill that is meant to help keep Huawei components out of American military hardware. And the other is the Defense Production Act, which is a statute from the 1950s that was meant to help Truman make sure that the steel mills and ammunition factories were up and running during the Korean War.
Do we really want to hand the same government a purpose-built regulatory apparatus for AI? That is to say, the very thing that the government will most want to control.
I know I've repeated myself like 10 times here, but I want to make this point again because it's worth stressing. AI will be the substrate of our future civilization. It will be the way you and I as private citizens will have access to commercial activity. We'll have access to information about the outside world and to advice about how we should use our powers as voters and capital holders. Mass surveillance, while it's very scary, is like the 10th scariest thing that the government could do with control over the AI systems with which we will interface with the world.
Now, the strongest argument against everything I've just argued is this. Are we really going to have no regulation on the most powerful technology in the history of humanity? Even if you thought that was ideal, there's clearly no way the government doesn't regulate AI technology in any way whatsoever. And besides, it is generally true that coordination could help us lessen some of the risk from AI.
The problem is I just don't know how to design a regulatory apparatus which isn't just going to be this huge tempting opportunity for the government to control our future civilization, which remember will be built on AI, or to requisition blindly obedient soldiers and censors and apparatchiks. While some kind of regulation might be inevitable, I think it'd be a terrible idea for the government to just wholesale take over this technology.
Ben Thompson had a post last Monday where he argued, "Look, people like Dario have made the analogy of AI to nuclear weapons in the context of arguing against a catastrophic risk, in the context of arguing for extra controls. But then, think about what that analogy implies." And Ben Thompson writes, quote, "If nuclear weapons were developed by a private company, the US would absolutely be incentivized to destroy that company."
And honestly, safety-aligned people have made a similar point. Leopold Aschenbrenner, who is a former guest and full disclosure a good friend, wrote in his 2014 memo Situational Awareness, quote, "I find it an insane proposition that the US government will let a random SF startup develop superintelligence. Imagine if we had developed atomic bombs by letting Uber just improvise."
And my response to Leopold's argument at the time and Ben's argument now is while they're right that it's crazy that we're entrusting private companies with the development of this world historical technology, I just don't think it's an improvement to give that authority to the government. Nobody's qualified to be the stewards of superintelligence. It's a terrifying, unprecedented thing that our species is doing right now. The fact that private companies aren't the ideal institutions to deal with this does not mean that the Pentagon or the White House is.
Yes, if a single private company were the only entity capable of building nuclear weapons, the government would not tolerate it having a veto power over how those weapons are used. But I think this is a terrible analogy for the current situation with AI for at least two important reasons.
First, AI is not some self-contained weapon like a nuclear bomb, which only does one thing. Rather, it is more like the process of industrialization itself, which is a general-purpose transformation of the whole economy with thousands of applications across every single sector. If you applied Ben Thompson or Leopold Aschenbrenner's logic to the Industrial Revolution, which is also world historically important, it would imply the government had the right to requisition any factory it wanted or destroy any business it wanted and punish and coerce anybody who refused to comply. But this is just not how free societies handled the process of industrialization. And it's also not how they should handle AI.
Now, people will say, "Well, AI will develop unprecedentedly powerful super weapons, superhuman hackers, superhuman bioweapons researchers, fully autonomous robot armies, and we just can't have private companies developing the technology that will make all this possible." But you can make the same argument about the Industrial Revolution. From the perspective of 17th-century Europeans, you've got all kinds of crazy in the world today that is a result of the Industrial Revolution, chemical weapons, aerial bombardment, not to mention nuclear weapons themselves.
And the way we dealt with this is not giving the government absolute control over the Industrial Revolution, which is to say over modern civilization itself. Rather, we banned and regulated the specific weaponizable end use cases. And we should regulate AI in a similar way, which is that we should regulate specific destructive use cases. For example, launching cyber attacks, things which should be illegal even if a human was doing them. And we should also have laws which regulate how the government can use this technology. For example, by building an AI-powered surveillance state.
The second reason that this analogy to some monopolistic private nuclear weapons developer breaks down is that it's not just one company that can develop this technology. There are many other frontier AI labs that the government could have turned to. The government's argument that it had to usurp the private property rights of the specific company in order to get access to a critical national security capability is extremely weak. It could have just instead made a voluntary contract with one of Anthropic's half a dozen other competitors.
If in the future that stops being the case, and if only one entity remains capable of building the robot armies and the superhuman hackers, and we have reason to worry that with their insurmountable lead they could even take over the whole world, then I agree that would be unacceptable for that entity to be a private company. And so, honestly, I think my crux against the people who argue that AI is such a powerful technology that it cannot be shaped by private hands is just that I expect this technology to be very multipolar, and I expect there to be lots of competitive companies at each layer of the supply chain.
And unfortunately, it's for this reason that I don't think that individual acts of corporate courage solve the problem. And the problem is this, that structurally AI favors many authoritarian applications, mass surveillance being one of them. Even if Anthropic refused to sell its models to the government to enable mass surveillance, and even if the next two companies after Anthropic did the same, in 12 months everybody and their mother will be able to train a model as good as the current frontier. And at that point, there will be some vendor who is willing and able to help the government enforce mass surveillance.
So, the only way we can preserve our free society is if we make laws and norms through our political system that is unacceptable for the government to use AI to enact mass censorship and surveillance and control. Just as after World War II, the whole world set this norm that you were not allowed to use nuclear weapons to wage war.
I want to be clear here, these are extremely confusing and difficult questions to think about. And even in the very process of brainstorming this video, I changed my mind back and forth on them a bunch. And I reserve the right to change my mind again. In fact, I think it's essential that we change our mind as AI progresses and we learn more. That's the very point of conversation and debate.
Someday, people will look back on this time the way we look back on the Enlightenment. People having these big, important debates just as the world is about to undergo these huge technological and social and political revolutions. And some of the thinkers even managed to get a couple of the big questions right, for which we today are still the beneficiaries. We owe it to our future to at least try to think through the new questions that are raised by AI.
Okay, this was a narration of an essay that I also released on my blog at dwarkesh.com. You should sign up there for my newsletter for future essays like this. Otherwise, I will see you for the next podcast interview. Cheers.
Article published
