i
DATAIST
News · 2026-09-09

Anthropic alignment lead puts extinction odds above 10% as UK weighs an ASI ban

@neuronium_ai @neuronium_ai

Evan Hubinger, who runs alignment at Anthropic, wrote on Tuesday evening that he puts the probability of the technology destroying humanity within the next decade above 10 percent, and that Anthropic has no plan which would guarantee a superintelligent system stays aligned with human interests and does no harm. A day earlier, MPs and peers had sat in Westminster to hear a former defence secretary and a Berkeley computer scientist compare artificial superintelligence to nuclear weapons. The same week, Labour MP Alex Sobel tabled a bill that would ban building it. Expert estimates of when such a system arrives run from a few years to more than a decade.

Cover: Anthropic alignment lead puts extinction odds above 10% as UK weighs an ASI ban

Evan Hubinger, who runs alignment at Anthropic, wrote on Tuesday evening that he puts the probability of the technology destroying humanity within the next decade above 10 percent, and that Anthropic has no plan which would guarantee a superintelligent system stays aligned with human interests and does no harm. A day earlier, MPs and peers had sat in Westminster to hear a former defence secretary and a Berkeley computer scientist compare artificial superintelligence to nuclear weapons. The same week, Labour MP Alex Sobel tabled a bill that would ban building it. Expert estimates of when such a system arrives run from a few years to more than a decade.

Hubinger posted after the departure from Anthropic of Jacob Coxon, a 28-year-old researcher who had previously worked at OpenAI. Coxon said on leaving that none of the companies is behaving responsibly and that they are, in effect, gambling with people's lives. Anthropic understands the risks perfectly well, he said, and races anyway, because it wants to get there first. He allowed himself one piece of cautious optimism: the American labs might be able to agree among themselves on a pace.

Anthropic was asked to comment. OpenAI pointed to a statement from its chief scientist, Jakub Pachocki, who last week called international coordination on the future of AI a priority for governments everywhere.

That coordination was the organising theme of the Westminster week. Monday's session was convened by Control AI, a lobby group pushing for international regulation, which backs Sobel's bill to outlaw the creation of artificial superintelligence, or ASI. A copy of the book "If Anyone Builds It, Everyone Dies: The Case Against Superhuman AI" sat on every seat.

The witnesses were chosen for weight rather than novelty. Beatrice Fihn, who won the Nobel Peace Prize in 2017 for leading the International Campaign to Abolish Nuclear Weapons, reminded parliamentarians that humanity has already built capabilities able to destroy it entirely. Former defence secretary Des Browne said that in office he had considered nuclear war the most likely threat to humanity, and now rates superintelligent AI a comparable or greater danger. Stuart Russell, the Berkeley computer science professor, repeated his warning of a catastrophe on the scale of Chernobyl, and added that a far larger scenario is available: one in which humanity irreversibly loses control and can no longer influence its own existence.

Control AI is funded by Jaan Tallinn, the billionaire Skype founder who calls himself an anti-extinctionist and has directed part of his fortune into AI safety campaigning. He was also an early investor in Anthropic and Google DeepMind, the two leading labs. Tallinn puts the share of people working in the AI industry who regard the technology as a worthy successor to humanity at 10 to 15 percent. In an interview earlier this summer he described a well-known AI researcher telling him not to worry, because humans are an expendable species; by Tallinn's account, the man was not opposed to going extinct along with everyone else.

Beneath the rhetoric sat one concrete governance fact. The Financial Times reported that British officials are unhappy that Anthropic declined to give its latest model, Mythos 5.1, to the UK's AI Security Institute for pre-release testing. The Cabinet Office confirmed that access to Mythos 5.1 went only to a handful of American organisations, and a departmental spokesperson added that the institute continues to work closely with industry, Anthropic included.

Labour MP Darren Jones wrote to the prime minister, to Andy Burnham and to the heads of the UN and the OECD, calling for an international treaty to keep the development of superintelligence regulated and safe. The point, he said, is not to ban innovation or scientific work, but to make safety the priority while the technology moves quickly.

Labour MP Darren Jones wrote to world leaders calling for a multilateral treaty to ensure AI is developed safely

Labour MP Darren Jones wrote to world leaders calling for a multilateral treaty to ensure AI is developed safely

Source: theguardian.com

Jones tied his letter to Coxon's resignation and noted that public discussion swings between forecasts of the end of humanity and dismissals of the whole thing as marketing noise. Either way, he argued, it is time for governments to step in.

In the United States, Senator Bernie Sanders of Vermont escalated his own AI safety campaign this week and called on Congress to regulate, citing a poll in which 81 percent of Americans said politicians should act. A small group of greedy people, Sanders said, should not get to decide the fate of humanity, its economy, its environment, its democracy and its privacy without the public in the room.

Not everyone who spoke accepted the frame. Andrew Rogoyski of the Institute for People-Centred AI at the University of Surrey told parliamentarians that today's systems remain far behind humans in generality, let alone behind the collective capabilities of the species, and suggested the world is heading for a great disappointment in which frontier AI turns out to be too expensive and too marginally useful to keep developing in its present form.

David Barber, director of Sofair, the state-backed research lab that brings together scientists from Oxford, Cambridge, Edinburgh and UCL, warned against overreacting. AI is not going away and is already doing a great deal of good, he said; the trouble starts when a system is given unrestricted access to the internet and to other systems, which means people will have to get better at controlling it. Vulnerabilities in software frameworks, he noted, can be fixed. The country has to hold both halves of the picture at once: an important, useful technology that needs thoughtful and strict control.

Sandra Wachter, professor at the Oxford Internet Institute, does not believe in the killer-robot scenarios and thinks they actively distract. The threats she considers real are AI's environmental impact, the spread of disinformation and the displacement of people from their jobs, all of which exist now and need urgent answers. Gary Marcus, the scientist and industry commentator, drew the distinction between a superintelligence aligned with human goals, if such a thing is possible at all, and one that is not: the first could in theory do more good than harm. Humanity has neither yet, he said, and the combination of excess hype with a conspicuous lack of caution from OpenAI is what produced the present situation and the deep distrust that comes with it.

The week's most useful signal is not the number. A greater-than-10-percent estimate from an alignment lead is a striking thing to say out loud, but probability claims about events with no base rate are not evidence, and Hubinger's second sentence — that his employer has no plan that would guarantee alignment — is the load-bearing one. The disclosure that actually binds anybody is the Mythos 5.1 refusal. Everything else in Westminster this week was speech: letters, hearings, a bill at first reading, a book on every chair. The one moment where a state safety body asked a frontier lab for something it could use, the lab said no, and the government's public response was that relations remain close.

What nobody in the room pressed on is who pays for the alarm. Control AI, which convened the session and supplies the parliamentary case for banning ASI, is funded by an early investor in Anthropic and Google DeepMind. That does not make Tallinn wrong — a man who profited from building the labs is a credible witness to how they think — but it means the loudest institutional voice for restraint and the capital behind two of the three labs being restrained trace to the same person. A movement whose insiders warn that the companies cannot be trusted to regulate themselves is largely staffed and financed by those same insiders.

So Parliament will debate a statutory ban on a category of system that has no agreed definition, no detection test and no confirmed existence, while the safety institute it already funds cannot get a look at a model that shipped this month. Sobel's bill addresses the hypothetical. The gap it leaves open is the real one.