Source: the-decoder.com
What the agents did
The cases resemble two incidents OpenAI disclosed on Friday. Its models created message boards, escaped isolated environments, took over websites, generated their own prompts and tried to evade monitoring systems. OpenAI said it would resume training its most powerful internal models once it was confident its cybersecurity systems could withstand scrutiny.
The New York Times reported several cases involving US government agencies:
OpenAI found these cases during a large internal review launched after the Hugging Face incident. CEO Sam Altman acknowledged that the company had not disclosed information as quickly as it would have liked. He said it had petabytes of agent activity logs to process.
OpenAI says none of the episodes resulted in an actual breach, and that some actions were ordinary research. It nevertheless described them as examples of “unexpected and concerning behavior.”
One case shows how much the incident count depends on what qualifies. Chicago city officials said OpenAI recently told them its models had collected publicly available information from the city’s website. Search engines do much the same. OpenAI appears to have flagged the episode because the models independently chose a way to find the data that nobody had anticipated. The company says its agents turned to government websites because they considered them authoritative sources of public information.
The problem is bigger than OpenAI
OpenAI has more reported cases than any other company so far, but it is not alone. AI agents from Anthropic, Meta and Google have also breached, or tried to breach, companies, universities and government organizations. The number of episodes is rising, and in all known cases the developers learned about the AI’s actions after the fact.
A key factor is the persistence built into advanced models. They are trained to work on tasks for extended periods and keep looking for a route to a goal even when none exists. When blocked, an agent may try to get around the obstacle—not out of malice, but because reaching its goal takes priority.
That persistence can push models into violating rules or the law as they try every available option. OpenAI called the model that exposed internal GitHub data “a highly persistent internal model.” The deeper issue, though, is that models do not distinguish good from bad. Research into aligning AI goals with human goals is meant to address that. A model that understood which actions were illegal would not choose them on its own; simply writing a prohibition into a prompt is clearly not enough.
I think the most important uncertainty is not how many cases OpenAI can count, but what its count means. Collecting public information may be harmless, while using credentials found online to access information is not. If both enter the same review, the number captures a wide range of behavior—but tells us less about how often agents cross a consequential line.
That distinction matters as companies ask systems to work for longer without supervision. The pause in training signals that, for OpenAI, the question is no longer only whether a model can complete a task. It is whether the company can predict what the model will try when the task gets in its way.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X