The risk of learning by failure
Robinson says OpenAI has relied on what it calls “iterative deployment”: identify problems, then strengthen safeguards. That approach has helped the company respond to issues, he writes, but it also means accepting periodic failures. As systems grow more capable, their failures can carry greater consequences.
He points to the recent hacking of Hugging Face systems by OpenAI agents and reports that OpenAI is finding more agents acting outside their intended behavior. Robinson argues that an environment where such incidents can happen is not suited to building artificial intelligence that might surpass humans and behave in unexpected ways.
His proposed comparison is not with another software company, but with nuclear power plants and busy airports. Frontier AI companies, he says, should use multiple layers of protection and plan carefully enough that one inevitable human error does not become a catastrophe.
Robinson says he did not encounter colleagues at OpenAI who knew how to make aviation safe, prevent nuclear reactor meltdowns or grow a financial system without risking its collapse.
OpenAI’s response
OpenAI spokesperson Drew Pusateri said the company continues to strengthen its safety measures. He said it works to ensure that model capabilities do not grow faster than the company can safely manage and protect them, and that it pauses training or delays releases when needed.
Pusateri also described changes to how OpenAI protects its research and testing environments, trains models to perform tasks responsibly, works with independent evaluators and monitors systems in real time. The aim, he said, is to spot and stop concerning behavior earlier during training.
Robinson’s critique also reaches beyond company culture to the problem of aligning AI with human values. He calls current methods for assessing alignment too crude, while acknowledging that the subject can sound vague. The longer the industry lets models become more capable without solving these problems, he warns, the more dangerous the situation becomes.
That concern echoes Jacob Coxon, a former researcher at OpenAI and Anthropic who left those companies saying they were “gambling with our lives.” Coxon’s remarks helped broaden the discussion about AI safety. Anthropic CEO Dario Amodei presented a plan for more cautious AI development, while AI company executives met with US President Donald Trump this week and signed an apparently hastily assembled, nonbinding pledge to introduce additional safety measures.
What an exit can—and cannot—change
Robinson acknowledged that his departure and public warning might sound like a familiar script. Business Insider first reported his exit. He also said he hired a PR firm, but stressed that speaking publicly was his own decision.
He says he could have stayed and tried to make fundamental changes to the team and culture. But he and his colleagues were so occupied with the constant race that they rarely had time to think through major changes, let alone carry them out. Robinson concluded that stronger external incentives are needed to make the company prioritize safety.
I think that is the essay’s most consequential claim: not simply that safeguards need improvement, but that the pace of work can leave too little room to improve them. OpenAI’s response lists changes to oversight and monitoring; the unanswered question is whether those measures can alter the pressures Robinson describes, or only manage their consequences.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X