Microsoft has published a code of conduct for its AI models, setting out the values and red lines the company says govern how it trains them and showing how those ideas are meant to work in practice. Three things are banned outright: cyberattacks, helping create nuclear weapons, and producing deepfakes. A fourth clause is the one that matters. Microsoft's MAI models must not use adaptive, deceptive, self-sustaining, collusive or other mechanisms that would let them bypass or remove human oversight, so that authorised people or systems retain the ability to reliably steer, modify and shut them down.
The document opens with a forecast rather than a policy: over the next decade, superintelligent AI systems will surpass humans at most tasks. Microsoft calls controlling, containing and aligning a tool that powerful with human goals one of the largest problems humanity has faced, and argues that this is why the purpose of such systems and the way they will be governed have to be fixed in advance rather than settled afterwards.
Two general principles sit above the constraints: models should support people rather than replace them, and should accelerate human development and wellbeing. The structural claim underneath is more specific than either. Each model carries a general code of conduct that takes priority over an individual user's preferences and over the requirements of the task at hand. That is a statement about who the model ultimately answers to, and it is not the customer.
The publication lands in a week of unprecedented attention to AI safety, driven by a series of incidents involving agents that got out of control and by the abrupt resignation of an Anthropic employee who tied the decision to a growing risk that AI leads to human extinction. Microsoft positions itself alongside Anthropic, OpenAI and xAI in broadly supporting a slower pace of frontier development, and endorses embedded evaluators inside the labs building AI. Chief executive Satya Nadella wrote that the company supports the research, the attention and the deliberate slowing of pace needed to align AI properly with human goals, and backed ideas such as embedded evaluators along with the wider work of turning those intentions into practice.
As a document, this is more concrete than Dario Amodei's call to slow down — a code with named prohibitions is a different artefact from an essay asking the field to reconsider. But the three absolute constraints are the cheap part. Nobody was going to defend nuclear weapons design or deepfake production, no serious lab claims otherwise, and writing them down costs Microsoft nothing it was planning to sell. The anti-evasion clause is the real content, and it is also the only part that cannot be checked from outside. "Does not use self-sustaining mechanisms to remove oversight" is not a property anyone can observe by using the product; it is a claim about training intent and internal testing, graded by the company making it.
Which is the thing the code, as described, does not address: what happens when a model breaks it. There is no audit named, no external body, no consequence, no definition of what counts as a violation serious enough to pull a model. The embedded-evaluator idea Nadella endorses is the closest thing to an enforcement mechanism in the entire picture, and it lives in a sentence of support, not in the code. The code also governs MAI models — the ones Microsoft trains itself — and says nothing, in what has been published, about the models Microsoft ships without having trained them.
The unanimity is the genuinely new fact here. Microsoft, Anthropic, OpenAI and xAI compete directly, and within days of each other they have converged on the same two positions: slow the frontier down, and let evaluators sit inside the building. That convergence is either the industry recognising a shared problem ahead of the regulators, or four companies agreeing on the terms of their own supervision before anyone else writes them. A document that predicts superhuman systems within ten years and proposes, as the remedy, rules its author wrote for itself does not settle which.