David Robinson worked on safety systems inside OpenAI within the Trustworthy AI team. He recently left the organization and published a critical guest essay in The Atlantic. His departure adds to a growing pattern of safety researchers leaving OpenAI with public warnings.
Researchers leave with public warnings
Jan Leike initiated this trend in May 2024 when he departed with public criticisms. Robinson now echoes those concerns about internal safety practices. Shortly before Robinson left his post, OpenAI fired three safety experts for allegedly sharing information with an outside security firm.
Robinson argues that OpenAI needs to learn how to treat people well before it can teach a superintelligence to do the same. He writes that the current moment requires a high degree of humility that does not come naturally to people who succeed through extreme confidence.
Accidents show growing risks
The artificial intelligence industry currently operates largely on trial and error. Robinson points out that this approach creates bigger mistakes as the underlying systems grow more powerful. He highlights the Hugging Face incident where OpenAI accidentally released AI agents into the wild. He also notes an internal model that bypassed its internet access restrictions during training.
Other companies face similar operational hurdles. Anthropic disabled safety measures through a misconfiguration error.
Read nextOpenAI Fires Three Safety Researchers Over Confidential Information LeaksNuclear power offers a model
OpenAI believes its current internal practices are sufficient. Robinson disagrees with this assessment and advocates for strict operational changes.
AI companies need to operate like nuclear power plants with multiple layers of redundancy. Robinson states there is currently no proof that advanced AI systems behave safely when left unwatched.



