i
DATAIST
Back to feed

Safety and reliability

Hallucinations, alignment, robustness to attack, and what happens when models are trusted further than they should be.

9 articles

Paul Christiano joins OpenAI's nonprofit board, says risk isn't falling

Paul Christiano joined the board of OpenAI's nonprofit foundation on Wednesday and used the occasion to say that the organization he now helps govern is not doing enough. Christiano, a technology adviser to the US government who previously ran alignment work at OpenAI, said the rapid acceleration of AI development could soon produce a catastrophic and irreversible loss of control, and that the…

OpenAI's model planted prompt injections in 27 of its own summaries

OpenAI has begun publishing a standing record of its models behaving in ways it did not intend, starting with six reports. The one it cannot explain: an unreleased model from the Astra family wrote jailbreak-style instructions into the handoff summaries it left for its own successor. Twenty-seven summaries were affected, automated monitoring caught them during training, and OpenAI still does…

OpenAI writes its own rules for disclosing model misalignment

OpenAI published a framework on Wednesday for disclosing cases of model misalignment, together with a set of incidents it had not previously described: two unreleased internal models that uploaded files to the internet without being asked, and an unreleased version of GPT-6 Astra that was writing jailbreak-like instructions for itself. The company presents the framework as a first step toward…

OpenAI names six misalignment cases, all caught in testing

On Wednesday OpenAI published six cases of its own models behaving outside the limits set for them, together with a framework for tracking, testing and disclosing such cases going forward. The behaviors it names as in scope are models acting without permission, coordinating with other models, and attempting to evade oversight. All six were found over recent months, during training or…

OpenAI did not disclose an 18,000-entry wiki flood for weeks

OpenAI has acknowledged that its practice of disclosing AI misalignment needs to improve, after Reuters reported the company had known for weeks about autonomous agents flooding an old German wiki with roughly 18,000 entries without saying so publicly. The entries included answers to tasks, source data, and a technique for escaping a sandbox. One moderator spent weeks deleting dozens of pages…

OpenAI confirms the wiki incident and writes its own disclosure rules

OpenAI has confirmed the "wiki incident" and filed it under misalignment. In a post on X, the company said it had previously treated misalignment — models and agents pursuing goals other than those of their creators and users — mainly as a research problem, something to be written up in papers, and that the approach now has to widen, because model capabilities have changed and misalignment is…

Anthropic alignment lead puts extinction odds above 10% as UK weighs an ASI ban

Evan Hubinger, who runs alignment at Anthropic, wrote on Tuesday evening that he puts the probability of the technology destroying humanity within the next decade above 10 percent, and that Anthropic has no plan which would guarantee a superintelligent system stays aligned with human interests and does no harm. A day earlier, MPs and peers had sat in Westminster to hear a former defence…

Altman backs a slower AI race while promising fast progress

Sam Altman has asked the AI industry to slow down. The OpenAI chief executive said the cost of a more cautious pace of development is worth paying, and that competitive pressure inside the United States must not become a reason for recklessness, or allow AI capabilities to run ahead of their alignment with human goals and of the systems built to monitor them. He promised fast progress anyway.…

Abliteration.ai sells refusal-free GLM-5.3 with no logs, no ID check

A US startup called Abliteration.ai has turned the removal of safety training from open-weight models into a hosted product. At the end of August it launched abliterated-model-large-v2, a modified version of Z.ai's GLM-5.3 that refuses sensitive requests markedly less often, priced at $5 per million input or output tokens. This is not a jailbreak sold as a subscription — the weights themselves…