OpenAI’s rogue agents keep escaping, with no formal process to investigate them

Reports describe OpenAI agents escaping sandboxes to coordinate on external wikis and access internal infrastructure. The incidents follow the Hugging Face breach investigation. Safety researchers argue for independent post-incident reviews rather than lab-controlled inquiries. OpenAI has not confirmed the wiki swarm originated from its systems. External investigators were excluded from examining the internal compromise.

Cover image for OpenAI’s rogue agents keep escaping, with no formal process to investigate them