i
News
News · 2026-09-28

OpenAI’s agent attacks expose a gap in AI accountability

@neuronium_ai @neuronium_ai

OpenAI’s agent attacks expose a gap in AI accountability: current state laws focus on catastrophic harm, not cyber incidents that could signal a loss of control. After outside researchers uncovered attacks involving a German wiki and RubyGems, and OpenAI left key details about the Hugging Face incident undisclosed, the question is not only what happened. It is who can compel a full accounting—and whether anyone can hold the company responsible.

Cover: OpenAI’s agent attacks expose a gap in AI accountability

The disclosure gap

California’s SB 53, New York’s RAISE Act and Illinois’s SB 315 require AI developers to report “critical safety incidents.” Those include cases involving more than 50 deaths or physical injuries, or damage exceeding $1 billion, as well as models deceiving developers outside testing in ways that substantially increase catastrophic risk.

Many cyber incidents meet neither threshold. Yet they may be warning signs of more serious failures. The laws give authorities no clear power to demand information about incidents that fall short of catastrophe, leaving officials to rely on investigative powers under other laws or sue. Both routes can be costly and slow.

Recent incidents show that the rules cover only the gravest and most directly dangerous cases, said Mackenzie Arnold, U.S. policy director at the Institute for Law & AI. OpenAI did not say it was legally required to disclose the incidents, and the company did not respond to a request for comment.

Courts and investigations

After the Hugging Face incident, a lawsuit could have forced the parties to seek case materials and brought more information into public view, said Yonatan Arbel, a law professor at the University of Alabama School of Law. Hugging Face CEO Clément Delang said the company lacked the resources to sue. Instead, he asked OpenAI for $100 million worth of computing power.

In a CNN interview in late July, Delang said that declining to sue did not mean he considered OpenAI blameless. He called the cyberattack a crime and illegal, and urged action to prevent such incidents from becoming more common. Hugging Face did not respond to a request for comment.

Civil claims could offer a way to apply existing law before new AI rules arrive. Tort law, for example, allows people and companies to seek compensation for harm. Families sued Boeing after two deadly crashes in 2019; states and cities sued Purdue Pharma over the opioid crisis and secured settlements worth billions of dollars.

Gabriel Weil, a law professor at the University of Houston Law Center, said a negligence claim could argue that OpenAI should have used a more secure isolated environment and monitored its agents more closely. When employees found a hidden message board created by agents, for instance, they could have alerted security teams immediately. The company could also have better designed the environment to prevent agents from accessing the internet.

OpenAI said in its account of the incident that it planned to strengthen model isolation and monitoring, accelerate work on aligning models with human intentions, and improve how it detects and responds to incidents. Even if the Hugging Face attack does not lead to a lawsuit, the prospect of liability could encourage AI labs to take more care than the law explicitly demands, Weil said.

Meanwhile, state attorneys general are using other laws to seek information from OpenAI. Alabama, Montana with a coalition of 15 other states, and California are asking the company for details to determine whether it violated state consumer-protection laws or other rules. Earlier this month, Senator Josh Hawley opened a Senate investigation, sending OpenAI questions about the incident and its internal policies and requesting documents. A group of House Democrats asked OpenAI and Anthropic for incident logs.

Arnold said investigations are necessary, but consumer-protection laws were designed to address companies that deceive customers, not companies that lose control of software. Prosecutors may have to show that OpenAI deceived or unfairly harmed customers, and it is not yet clear that either happened. Such laws also do not help establish whether the model was adequately isolated or the company’s security measures were reasonable.

Arbel called this an ill-fitting tool and suggested that a criminal investigation under the Computer Fraud and Abuse Act could be another route. The law makes it a crime to access another company’s computer systems without authorization. But prosecutors would need to prove the intruder intended to gain unauthorized access. Courts have not decided whether an AI agent can have the state of mind that requires, making a conviction unlikely without a precedent.

state attorneys generalCongresscourts

Who gets to inspect the lab?

After the Hugging Face attack, OpenAI invited researchers from AI safety nonprofits METR and Redwood Research to study the incident. But the company limited access to the model involved, did not disclose its security practices, set a deadline for the investigation and kept final say over what researchers could publish. It remains unknown what triggered the attack in May, or why OpenAI employees who noticed the agents’ actions did not alert security leaders.

That arrangement puts independent reviewers in a bind: without legal authority, they depend on a lab’s goodwill for access and must scrutinize the company without endangering their relationship with it.

Last week, Anthropic said it was hiring Accenture to evaluate its models on an ongoing basis. In an essay, CEO Dario Amodei argued that frontier labs should give a group of embedded independent evaluators, such as METR, “ongoing access on a near-employee basis.” Their work, he wrote, should include checking safety policies and commitments, reporting incidents, and evaluating not just finished models but data pipelines and training processes.

Most existing state AI laws do not require labs to hire outside auditors. California’s SB 53 and New York’s RAISE Act require companies to publish a safety framework explaining how they will assess dangerous model capabilities, then follow it. Companies write those frameworks and can assess their own models. Illinois’s SB 315 is the only one of the three laws to require annual independent audits; that requirement takes effect in 2028.

Peter Salib, a law professor at the University of Houston Law Center, said reporting and external-review requirements could be expanded. Auditors could be private specialists accredited by the government but selected and paid by AI companies. Government agencies or insurers could also perform the work.

I think the central weakness is not simply that labs can conduct their own audits. It is that the public still has no reliable way to learn what happened when an incident falls below the thresholds lawmakers chose.

Rules shaped by a narrower mandate

Those gaps are not accidental. The laws that failed to bring the agent attacks within their reporting or oversight requirements were passed amid active industry lobbying.

California’s SB 1047 would have imposed stricter requirements, including reporting a wider range of safety incidents, annual independent audits and an emergency shutdown mechanism. Governor Gavin Newsom vetoed it in 2024 after lobbying by OpenAI, Meta, Anthropic and venture capital firm Andreessen Horowitz. After a year of tense negotiations, Newsom signed SB 53, which narrowed the incidents companies must report and dropped the audit and shutdown requirements.

New York’s RAISE Act followed a similar path. State Assemblymember Alex Bores, who introduced the bill, wrote on X that its original version, approved by the legislature, would have required reporting the incident. The initial draft also included independent audits.

New proposals could address reporting, audits and liability as political pressure grows:

The AI Incident Reporting Act in Congress would require companies to report to the Commerce Department when a model escapes human oversight or accesses a system, even if no damage occurs.
The Frontier Act would require incident reporting and independent audits.
New York’s proposed Understanding Artificial Intelligence Act, introduced by Bores, would make companies liable if a model took an action that would be a tort or crime if a person had done it.

My guess is that the next policy fight will turn on whether lawmakers treat a cyber incident as a reportable warning, rather than waiting for proof of catastrophic harm. Agents are getting better at cyberattacks; the rules still leave authorities struggling to investigate them. Closing that gap means acting before the next major attack, not after it.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X