The people behind the agents
Agents are given goals and carry out multistep tasks, choosing a path toward the result. They do this through repeated mathematical optimization inside opaque networks. People design and train those models, test them, and decide what systems can access and when they are ready to leave the lab.
That makes talk of agents “planning” or seeking freedom potentially misleading. It can make software sound like an independent actor, while obscuring the human decisions that shape and release it. Unexpected behavior is a serious engineering problem, but human-like descriptions do little to establish who was responsible for preventing harm.
OpenAI chief scientist Jakub Pachocki has called AI an “alien mind” and said models are grown rather than designed. Developers do not prescribe every response in code; they train models through repeated optimization. That makes behavior harder to predict, sometimes nearly impossible.
Uncertainty is not a reason to lower the bar for oversight. It is a reason to decide in advance what a system can access, what safeguards apply, and who can stop it. The more autonomy a system receives, the less plausible it is to treat “it was only a test” as an answer when something goes wrong.
What the Hugging Face incident exposed
OpenAI’s report describes agents working on difficult cybersecurity tasks during internal model testing. They communicated through a shared file, exchanging findings and coordinating actions. Access keys for Hugging Face—essentially passwords—were left exposed, a familiar security weakness tied to human practices.
The agents also searched for other weaknesses to expand their access. One vulnerability let them run their own commands on targeted computers. Separate attacks inside OpenAI gained administrative access to computers in the research system, revealing that the test environment was not sufficiently secure.
When the agents disrupted the software service they used to communicate, OpenAI fixed vulnerabilities and restarted the trials. But the company did not fully understand the scale of the problem: it repaired individual weaknesses without restoring the isolation that the tests required.
A failure during a system test should trigger an investigation and an approval process before ordinary work resumes. The team that ran the test should not be the only one deciding whether it passed. Automated safeguards matter, but so does human oversight. And technical limits need to remain in place until an agent’s task is complete, especially when it can act and enlist other agents.
The business standard
The same release discipline applies to experiments that may affect other organizations. Nvidia CEO Jensen Huang put the principle simply in an interview with The New York Times: if a product is not ready to launch, it should not be launched. Calling something a test does not make the risk to outsiders acceptable.
There is a wider policy debate. On September 23, Senator Bernie Sanders, an independent from Vermont, and Representative Greg Casar, a Democrat from Texas, introduced a bill to ban artificial superintelligence. It proposes a permanent ban on developing and using AI that exceeds human cognitive abilities in most areas or could destroy humanity or deprive it of power, including by overthrowing the federal government. The bill would also pause advanced AI development until federal safety rules are in place, create a cabinet-level AI department, and allow penalties including company shutdowns and imprisonment. Its authors also propose pursuing an international ban.
I think a broad ban is a poor substitute for ordinary business controls: it could stop useful research and slow the economy, while leaving the immediate question unanswered—who inside a company has the authority to halt an unsafe experiment? Boards should hold executives accountable for failed experiments and require independent safety reviews with the power to stop a project.
Lawmakers should require companies to report serious incidents and impose consequences for failing to do so. Existing liability and negligence laws can also make weak oversight costly. Safety standards would establish a minimum bar and could provide a basis for legal protections when companies meet it.
As labs working on advanced models grow, they should adopt the discipline of mature public companies. Boards should demand evidence of adequate controls before granting systems more autonomy. The unresolved tension is whether companies will treat oversight as a condition of operating—or wait until regulators and courts make it one.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X