Four Anthropic researchers publicly endorsed a departing colleague's warning that the models their employer is building could end humanity within a decade. Jacob Coxon, who left Anthropic on Wednesday after earlier working at OpenAI, wrote in a post that went viral that neither company develops models responsibly and that both are "playing with our lives." Evan Hubinger, who describes himself as head of aligning models with human goals at Anthropic, put the probability of that outcome at more than 10% over the next ten years. Elon Musk's response was that the whole episode is staged.
The endorsements came within a day. Anna Wang, who works on general artificial intelligence safety at Anthropic and previously worked at Google DeepMind, said many employees want development slowed down so that a plan for reducing the risks of such models can be prepared. On X, she added that no scientific plan yet exists for handling the risks of recursively self-improving AI.
Drake Thomas, also at Anthropic, said he respects Coxon's decision to stop building models if Coxon believes they could pose a planetary-scale threat. Development is moving too fast, Thomas said, and the level of confidence in safety that would be required before superintelligent AI — ASI — arrives is not currently achievable.
Samuel Marks, a safety researcher at the company, wrote that AI developers believe their technology could cause human extinction or consequences on a comparable scale, possibly within the next few years, and that concern rises with seniority. Hubinger went further, saying Coxon is right and that the industry is not keeping pace with the apocalyptic potential of what it is building. Employees, he wrote, genuinely take the destruction of humanity by AI to be possible.
An Anthropic spokesperson, responding to The Guardian, defended the company's strategy: Anthropic has always been open about the dual nature of AI, which will bring enormous benefit and also create unprecedented risks, and it continues to build models with some of the strongest safeguards in the industry.
Musk took a different route. He and other people who hold a more positive view of AI began circulating the theory that Coxon's post and the fallout are part of a psyop meant to turn the public against the technology. Musk wrote on X that preparation for such an operation, for lack of a better term, may have been underway for a long time, and that Coxon's post was merely the match that lit the fire.
He was replying to Parker Thayer, a researcher at the conservative think tank Capital Research, who suggested — without presenting convincing evidence — that Coxon's post launched an "extremely sophisticated and well-funded PR operation" aimed at winning Democratic support for AI regulation that would effectively destroy the industry. Bill Ackman, the billionaire chief executive of the hedge fund Pershing Square, noticed Thayer's post and called it interesting.
Coxon answered Musk with a selfie, said he is a real person who actually holds the views he published, and noted that Musk could have asked the researchers at his own company xAI about him, had he not fired them.
The detail that matters most here is one nobody in the exchange dwells on: the people issuing extinction warnings about Anthropic's models are, with one exception, still building them. Coxon left. Wang, Thomas, Marks and Hubinger posted. That gap is the story. A researcher who assigns better than one-in-ten odds to civilizational catastrophe and continues shipping is either running a calculation about counterfactual impact that he has not shown his readers, or the belief is held at a different temperature than the words suggest. Both readings are worse than they sound, because neither is available to anyone outside the company to check.
The company's statement is built to survive this. It does not contest the extinction claim; it absorbs it. An organization whose stated position has always been that the technology carries unprecedented risks cannot be caught contradicting its own employees when those employees say the risks are unprecedented. Dissent that arrives already inside the house position is not a crisis to be managed.
What none of the four says is what would make them stop. Wang says there is no scientific plan; no one specifies what would count as one, what evidence would clear the bar Thomas describes as currently unreachable, or which observation would move Hubinger's number down instead of up. Without that, the warnings function as a description of a mood rather than a decision procedure, and a mood cannot be falsified — which, unhelpfully, is also the objection to Musk's psyop theory.
Others reject the frame entirely. Gary Marcus, a scientist and prominent critic of AI, has called for boycotting the technology, but on the grounds that it is already causing harm. What frightens him most is not extinction but bioweapons built with AI, wars started or escalated by AI-generated disinformation, and attacks capable of disabling critical infrastructure. Nothing he has seen, he said, suggests that even one of those risks is under control.
On the same day, Anthropic published a report describing how it disrupted an operation that used its models to build a biological weapon. Read one way, that is evidence for Marcus and for the departing researchers: the capability is live and people are already reaching for it. Read the other way, it is evidence for the spokesperson: the safeguards worked. The company can cite whichever half the moment requires, and the dual-nature framing exists so that it never has to choose.