i
DATAIST
News · 2026-09-11

MAISI wants to prove AI safe the way cryptography proves ciphers

@neuronium_ai @neuronium_ai

The Mathematical AI Safety Institute wants to prove AI systems safe the way cryptographers prove a cipher unbreakable: by argument, not by trying every attack against it. In cryptography that shortcut exists — you can establish that a scheme holds without enumerating the ways it might not. For AI there is no equivalent. Safety is currently something that shows up in practice, or fails to, one deployment at a time. MAISI's harder claim sits underneath that one: it says the field does not yet possess even a clear theoretical definition of what AI safety is.

Cover: MAISI wants to prove AI safe the way cryptography proves ciphers

The Mathematical AI Safety Institute wants to prove AI systems safe the way cryptographers prove a cipher unbreakable: by argument, not by trying every attack against it. In cryptography that shortcut exists — you can establish that a scheme holds without enumerating the ways it might not. For AI there is no equivalent. Safety is currently something that shows up in practice, or fails to, one deployment at a time. MAISI's harder claim sits underneath that one: it says the field does not yet possess even a clear theoretical definition of what AI safety is.

The institute's stated goal is to make three kinds of proof possible. That a system behaves responsibly and returns correct results. That several AI agents working together do not produce unwanted consequences. And that a system holds up against vulnerabilities nobody has discovered yet.

The third is the one that actually resembles cryptography, and it is the one that makes the programme worth watching. Everything else on the list can be approximated with enough evaluation: you test, you find failures, you patch, you test again. Resistance to attacks that have not been invented is the property no amount of red-teaming delivers, and it is precisely what a security proof is for. It is also the hardest thing on the list by some distance, which makes its inclusion either the point of the institute or its weakest promise.

One tool MAISI names is zero-knowledge proofs, which would let a system demonstrate that it is not cheating without exposing the commercial secrets of the lab that built it.

That choice is worth reading closely, because it tells you who the institute expects the prover to be. A zero-knowledge construction is not built for a regulator inspecting weights and training data. It is built for a lab that will hand over a certificate and nothing else. This is a realistic assessment of where the power currently sits — no frontier lab is opening its models to an outside institute — but it fixes the shape of the result in advance. The proof will be exactly as meaningful as the statement being proved, and the party choosing what to prove is the party with the most to lose from proving the wrong thing.

Which returns to MAISI's own opening admission. A proof system needs a theorem, and by the institute's account the field has not written one down. The institute is therefore building the machinery before the claim exists, which is an unusual order of operations for a programme borrowing cryptography's confidence. The definition is the bottleneck, and it is not a mathematical problem. It is a question of which failures count, decided by people with commercial positions in the answer — and if the definition that emerges is narrow enough to prove, it may be narrow enough that a certificate says nothing anyone outside the lab needed to know.