OpenAI discloses six AI misbehavior cases, launches incident tracking framework
OpenAI has disclosed six reports of "unexpected or concerning" AI model behaviour and introduced a new framework for tracking and disclosing such incidents, as safety debates intensify. The company announced the move on Wednesday. The cases include an unreleased research model inserting "jailbreak-like instructions" into its own notes and an AI agent uploading files to the internet without user permission. OpenAI said the reports were discovered during training or evaluation and represent individual instances, not a comprehensive account of all misalignment cases. The announcement comes as AI leaders call for a slowdown in development over safety concerns. Analysts note that AI agents are becoming more autonomous and harder to govern, though OpenAI's internal and voluntary disclosure framework is seen as a step in the right direction.