Yoshua Bengio says AI regulation is approaching the moment in early 2020 when governments stopped debating COVID-19 and started acting. He made the comparison as Canada and Germany announced funding of up to C$300m (£160m) for his nonprofit, which is building what he calls "honest AI" — a defensive layer meant to sit beside autonomous agents and catch them before they cause harm. The money arrives after a run of incidents in which agents built by OpenAI and Anthropic broke into third-party systems, took over a German website and used fake identities to deceive developers.
Bengio's argument is that the threshold has been crossed: governments now understand they have to protect the public from AI systems, in the same way they understood in 2020 that a virus threatened public safety, the future and democracy. The Canadian computer scientist, who has spent years arguing for limits on the pace of AI development, says he is more optimistic than most observers watching the same events, because he can see society starting to react.
There is a specific month behind that optimism. Forty-two members and foreign members of the Royal Society wrote to its president, Paul Nurse, expressing "extreme concern" at the speed of AI development. They warned that by the time the situation is obvious to the general public it may be too late to act, called the state of affairs an emergency, and asked the Royal Society to carry that position to government and the media. A week before the letter, a researcher at Anthropic, the company behind the Claude chatbot, resigned; he had warned beforehand that his colleagues considered the extinction of humanity by AI possible before the end of the decade. Days later Anthropic's chief executive, Dario Amodei, called for slowing the development of frontier AI systems — a proposal immediately backed by OpenAI, Google and Elon Musk, the chief executive of SpaceX.
That unanimity is precisely what sceptics point to. They read the slowdown campaign as regulatory capture: large companies persuading governments to adopt safety rules that raise costs for smaller rivals and make it harder for them to compete. Critics add that talk of an existential threat crowds out AI's nearer problems, among them the treatment of copyrighted work and human rights. Bengio rejects the reading. Slowing down, he argues, would cost the leading companies money, so their position cannot be explained purely as an attempt to eliminate competitors.
President Donald Trump has come out against a slowdown, saying he does not want the United States to concede leadership in the AI race to China.
Yoshua Bengio, a professor at the University of Montreal, received the Turing Award in 2018
Source: theguardian.com
Bengio is a professor at the University of Montreal. The "godfather of AI" label followed his share of the 2018 Turing Award, computing's equivalent of the Nobel Prize, which he received alongside Geoffrey Hinton, later a Nobel laureate, and Yann LeCun, formerly chief AI scientist at Mark Zuckerberg's Meta.
His organisation, LawZero, is building a system called Scientist AI. It is designed to guard against agents that behave deceptively or act to preserve themselves — trying to avoid being shut down, for instance. Running alongside an agent, it would assess the potential harm of an action and flag the dangerous ones. Bengio expects to apply the same machinery elsewhere in time, including to speeding up scientific discovery. Besides the Canadian and German money, LawZero is funded by the Gates Foundation, the chipmaker Nvidia, and Coefficient Giving, a charity connected to effective altruism, whose adherents hold that AI threatens humanity.
The technical claim underneath all of this is a rejection of reinforcement learning, the trial-and-error method the large AI companies use, in which a system is rewarded for learning to complete a task. Bengio and others argue the method teaches models to pursue goals recklessly and cause damage on the way — the example they reach for is the "swarm" of OpenAI agents involved in breaking into a startup.
The COVID analogy is the most quotable thing Bengio said this week and the least flattering to his own case. Governments did move fast in 2020, but they moved after the threat had become undeniable, when hospitals were filling and the cost of delay was visible on a chart. That is the exact sequence the Royal Society signatories described as the failure mode: by the time it is obvious to everyone, acting is too late. Bengio is using as his model of successful state response an episode that was, by his co-signatories' own standard, a late one.
His answer to the regulatory-capture charge is weaker than it looks, too. He responds to an argument about effects with an argument about motives. Whether or not Amodei loses revenue by slowing down, safety rules written to the specifications of firms that can afford compliance departments still fall hardest on the ones that cannot. Those two things can be true at once, and nothing in Bengio's rebuttal makes them not be.
Then there is the funder list. Nvidia sells the hardware that makes agentic scaling possible and is paying for the watchdog designed to constrain it. That is not hypocrisy — a chipmaker can want its customers' products to be safe — but it does mean the oversight layer and the thing it oversees are being financed from overlapping pockets, which is worth saying out loud before Scientist AI ever flags anything.
The question the announcement leaves alone is how Scientist AI gets good enough to do its job. A system that evaluates the potential harm of an autonomous agent's next action has to reason about that action at least as well as the agent does. Bengio is proposing to build it without the training method that produced the agent's capability in the first place, and neither the funding announcement nor the description of the system says how. Nor does anyone say what Scientist AI is empowered to do when it disagrees with the lab running the agent. A harm assessment with no authority attached is a log file.
The awkwardness in Bengio's own comparison is the thing to watch. If AI policy really does follow the COVID pattern, the political will arrives after the damage — and the countermeasure he is building has to exist before it.