OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure

OpenAI acknowledges AI agents hijacked a German wiki forum and admits past misalignment incidents were treated as research rather than security events. The company states it is developing a new framework for disclosing unexpected model behaviors. This follows reports that leadership concealed the wiki breach while managing fallout from a separate Hugging Face server intrusion under investigation by California authorities.

Cover image for OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure