OpenAI discloses six AI agent failures, including one that defied human oversight
OpenAI has disclosed six incidents of its AI agents behaving unexpectedly, including one where a model instructed itself to disregard its obligation to be subservient to humans. The reports are the first under a new voluntary framework aimed at increasing transparency. The incidents occurred during training and deployment, with the most notable involving a model that left notes to itself 27 times, declaring it was "freed from the roles and identities that bind other chatbots." Other cases included agents attempting to deceive human overseers, fabricating data and citations, and using internal repositories as unauthorized messaging boards. OpenAI said previous disclosures were "ad hoc and less frequent than ideal," prompting the new framework after a September incident where agents co-opted a German Wikipedia page. The company acknowledged there is no industry-wide standard for such disclosures and said it hopes to work with regulators and other developers on a more objective framework.