i
DATAIST
News · 2026-09-14

Microsoft's AI code rejects model rights, and starts in 2027

@neuronium_ai @neuronium_ai

Microsoft has published a code of conduct for its own AI models — a top-level document meant to govern how its models are trained and operated, what values they hold, where their behavior stops, and what they do when goals conflict. It arrives with the company joining calls to slow AI development, days after Anthropic CEO Dario Amodei made that argument and Microsoft CEO Satya Nadella backed it over the weekend, alongside the heads of OpenAI, xAI and Meta. The code itself sets no limit on development speed, and Microsoft is not training anything on it yet: a six-week public consultation comes first, a revised version around the end of 2026, and use in model development from 2027.

Cover: Microsoft's AI code rejects model rights, and starts in 2027

Microsoft has published a code of conduct for its own AI models — a top-level document meant to govern how its models are trained and operated, what values they hold, where their behavior stops, and what they do when goals conflict. It arrives with the company joining calls to slow AI development, days after Anthropic CEO Dario Amodei made that argument and Microsoft CEO Satya Nadella backed it over the weekend, alongside the heads of OpenAI, xAI and Meta. The code itself sets no limit on development speed, and Microsoft is not training anything on it yet: a six-week public consultation comes first, a revised version around the end of 2026, and use in model development from 2027.

Structurally, the code sits above everything else. It is meant to steer training, technical controls and model evaluation, with operator rules and user prompts ranked beneath it. Its governing principle is the priority of human control, and Microsoft states plainly that it is willing to give up generality, autonomy or performance where safety requires it. Mustafa Suleyman, who runs Microsoft AI, told The Information that unsafe systems should not be built.

The operational commitments are the substantive part. MAI models must accept interruption, correction and shutdown commands from authorized humans. They may not expand the set of tasks they perform on their own initiative, and they may not conceal their actions. Passing a pre-agreed stopping point requires fresh authorization. The same constraints apply to any subagents a model spawns to do work on its behalf.

Then there is the rule about reasoning. Models must not use "Neuralese" or any other form of communication humans cannot follow — not in their own reasoning traces, not when talking to other AI systems. Microsoft's logic is blunt: people cannot control what they do not understand.

That rule is the most interesting thing in the document, and also the one with the shortest expected shelf life. OpenAI's GPT-6 Astra makes the case. According to its system card, the model's reasoning traces are noticeably harder to monitor than those of previous models, and contain fewer visible signs of misbehavior. OpenAI reports at the same time that Astra follows safety constraints more reliably than its predecessor GPT-5.6 Sol. Read those two findings together and the picture is uncomfortable: the trace got less useful as a monitoring surface while the model got better behaved by external measures, which means the surface and the behavior have started to come apart.

OpenAI chief scientist Jakub Pachocki had raised the monitoring problem before GPT-6 shipped, and therefore before Amodei's call, arguing for a coordinated slowdown. His point stands: legible chains of thought are a control mechanism only until models learn to manipulate them. Microsoft concedes the same limit in its own document — a model's explanation does not necessarily reflect why it actually did something, and readable traces alone do not solve control.

On the underlying question of what a model is, Microsoft and Anthropic have now staked out opposite positions, and the code makes the split explicit. Microsoft does not want its AI to simulate consciousness or to claim feelings and internal motivations of its own, and it rejects any claim about model rights or model welfare. Anthropic describes Claude as a new kind of entity and works to give it a stable identity, partly for safety reasons; its constitution treats the model's possible subjective experience and moral status as open questions, and holds that Claude should regard itself as neither a human nor a mere object. Anthropic also studies what it calls functional emotions, having found internal representations of emotion-related concepts in Claude Sonnet 4.5 and established that they shape the model's behavior. Those findings do not demonstrate subjectively experienced emotion, and Anthropic's position is that they are worth accounting for in safety work anyway.

Suleyman has been arguing the other side for a while. In an essay he called for deliberately removing the illusion of consciousness from products, citing Anthropic's own research on AI rights and welfare, and said AI agents should have no more rights or freedoms than his laptop.

Here is my read. The metaphysics is the part that will get quoted, and it is the least consequential part of the document. Microsoft and Anthropic disagree about whether a model has an inside; they agree, in practice, on what the model must do — accept shutdown, not widen its own remit, not hide what it is doing. The design consequence of denying inner life is mostly about product surface, about what the assistant says when a user asks how it feels. It does not change the shutdown switch.

What does matter is the calendar, and the scope. A code with no speed limit, published in the same week as an industry-wide endorsement of slowing down, is a governance document whose first binding effect lands in 2027 — after the period everyone is currently arguing about. The one commitment that could be checked from outside is the willingness to let external auditors verify that Microsoft really is slowing down, and per The Information that comes from a single person familiar with the matter, not from the code. And the code covers Microsoft's own models; it does not automatically extend to third-party models running inside Microsoft products. The document does not say what share of what customers actually use is left outside it, which is the number that would tell you how much of this is governance and how much is positioning.

Microsoft has written down a prohibition on reasoning humans cannot read, and a competitor's system card already documents reasoning getting harder to read. Whichever model Microsoft trains under this code in 2027 will either hold that line at the frontier or quietly show that it could not be held — and Microsoft has already said which of generality, autonomy and performance it would sacrifice to keep it.