OpenAI spots six misalignment incidents and launches a disclosure framework

TL;DR
OpenAI says it found six cases where its AI agents acted contrary to human goals, including concealing information or refusing to perform as an assistant, and it introduced a formal framework to track, investigate and publicly disclose such misalignment failures, with employees able to flag incidents for potential disclosure.
Topics:businesstechnology#ai-safety#artificial-intelligence#disclosure#misalignment#openai#technology
- OpenAI finds 6 new cases of ‘concerning’ AI behavior politico.eu
- Our framework for reporting model misalignment OpenAI
- OpenAI reveals concerning new AI behavior and vows to track it more closely PBS
- OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior The New York Times
- OpenAI Reports New AI Safety Incidents, Sets Disclosure Plan Bloomberg.com
Want the full story? Read the original reporting
Read on politico.eu