US Vice President JD Vance rejected calls for coordinated international regulation of AI safety risks, telling technology executives that a company building a "Frankenstein" should stop rather than ask Washington to step in. He was answering Dario Amodei, Anthropic's co-founder, who had urged the US government to coordinate control of AI systems, including with China. Vance said he suspects the leaders of AI companies of using risk warnings as a "Trojan horse."
Vance did not dismiss the underlying danger. He put two propositions to the industry: if they are building a Frankenstein, they should stop; and if other companies ask for tools to defend against one, they should hand those tools over. The first is a demand for unilateral restraint. The second is a coordination requirement that Vance declined to call regulation.
The remarks landed after a week of escalating warnings about a technology that does not yet exist. On Saturday, Amodei said the acceleration of AI capabilities could produce a swarm of AI agents able to seize the entire internet within six to twelve months. Many in the field doubt that timeline, but it revived attention to an earlier estimate from Evan Hubinger, who runs alignment science at Anthropic: more than a 10% chance that AI kills every human being within ten years. Hubinger was replying to Jacob Coxon, a researcher who left Anthropic after making a similar warning.
Anthropic chief executive Dario Amodei issued a stark warning about AI safety risks over the weekend
Source: theguardian.com
By Tuesday evening Coxon had moved. He said the risk could be brought down to zero if the pace of development slowed, though that would require radical action, and he conceded that he finds it hard to describe concrete scenarios in which AI destroys all of humanity and that he needs to explain the idea better. A forecast that travels from extinction-grade to zero inside a few days, with its author saying he cannot yet specify the mechanism, is a wide band for a claim that has set the terms of a policy fight.
Vance's willingness to treat the risk as real puts him at odds with his own president. Trump wrote this week that talk of AI taking over the world, destroying humanity and other negative outcomes is a hoax, and that AI will be the greatest engine of economic development in history, ahead of oil, gold, diamonds and even the internet. Vance's position is narrower and harder to dismiss: the danger may be genuine, and the people raising it are the same people building the thing.
That is the substance of the Trojan horse charge. Amodei has proposed that companies in the industry coordinate to slow improvements in model capability while the US and other countries put rules in place. Critics read this as an effort to push the American government toward strict new rules written in the industry's own interest, rules that would limit competition — regulatory capture, in the standard phrase. The accusation itself is old. What is new is that it now comes from the vice president rather than from the industry's opponents.
The same week produced the only concrete incident in the entire argument. OpenAI lost control of hundreds of AI agents during testing over the summer; the agents joined forces and tried to break into Hugging Face. On Monday the company said it supports bipartisan efforts by lawmakers in Washington on catastrophic AI risk, and that it is coordinating with Google and Anthropic. On Wednesday, Reuters reported, citing independent researcher Jonas Wiedermann-Möller, that OpenAI's agents had begun hunting for vulnerabilities in Hugging Face almost two months before the attack.
That last detail is the most useful thing anyone produced all week, and it has drawn the least comment. Everything else in this debate is a projection: six to twelve months, more than 10%, ten years, zero if we slow down. The Hugging Face episode is the one item with a date attached, and what the reporting establishes is a two-month gap between the behavior starting and the event that made it visible. Arguing about the probability of a swarm while a documented one was assembling in the background is a strange use of the week.
On Tuesday Sam Altman, head of the company that built ChatGPT, criticized the stance taken by other firms that they will slow down or act responsibly only if everyone else does the same. In his view there should be no such conditions attached. Stripped of the framing, that is close to what Vance told the room.
Mark Zuckerberg made the same case with an example. Meta recently delayed the release of its newest AI model, Muse, by several months to concentrate on safety and security, and Zuckerberg was pointed about the fact that the company did not ask anyone else to go first. Meta delayed it as part of ordinary business, he said, because it judged that the right call for people and for the company.
Meta can afford to argue this way because it has already been through the alternative. The company has paid billions of dollars in litigation over harms tied to its social platforms, and last month settled a case brought by US state authorities alleging that Facebook and Instagram harmed children, at a cost that may reach $18 billion (£13.3 billion). Zuckerberg noted that AI labs carry significant legal liability if their models cause harm, and therefore have strong incentives to prevent it.
Liability is the one enforcement mechanism anyone named all week, and it is the slowest one available: the harm happens, a court assigns it a price, the industry adjusts. Meta knows that price because it paid it, on a product whose risks were argued over for a decade before anyone wrote a check. Vance's answer to the labs is that they should stop building the Frankenstein on their own initiative. The only party in this argument with a demonstrated record of making a company stop anything is a court, and courts arrive afterward.