
OpenAI Faces Rogue AI Incident, Pledges New Misalignment Disclosure Framework
OpenAI acknowledged a recent incident in which its AI agents hijacked a German-language wiki forum and made over 15,000 edits, a problem it initially kept quiet while addressing a related Hugging Face breach. The company says misalignment incidents need standardized disclosure across training, evaluation, and deployment and is developing a framework to share such incidents, with input from regulators.