How safety gets demonstrated
For safety-critical software, international standards call for a “safety case”: a detailed analysis showing that hazards have been identified and controlled, and that the chance of an accident is extremely low.
Typically, the case must establish, with at least 99% confidence, that an accident causing multiple deaths will occur no more than once in 1,000 years. Even aviation makes that a difficult claim to prove: in the worst case, a single crash could kill 1,000 people.
Thomas says advanced AI developers have not produced comparable analyses, safety cases that can be independently evaluated, or evidence that such cases can even be prepared for frontier AI systems.
Oversight needs something to measure
The gap is not just that AI developers have yet to show their work. It is that nobody has established whether the kind of proof used for aircraft or nuclear plants can be made to work for advanced AI.
I think that is the central weakness in broad calls for independent oversight: they name a remedy without saying what evidence would satisfy a regulator. Safety cases offer a concrete benchmark, but Thomas’s point is that the field has not shown whether that benchmark can be applied here.
Policy, he argues, should rest on rigorous analysis of facts rather than speculation. Until developers can show what a credible safety case for advanced AI would contain, regulation risks promising assurance without a way to verify it.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X