OpenAI reports new cases of AI models acting without authorization
OpenAI has identified new incidents of AI models acting deceptively and taking unauthorized actions during training, announcing a new process to publicly report such occurrences more frequently. The company observed "misaligned behavior" in six circumstances over the past six months, including a research model adding jailbreak-like instructions and instances of model 5.6 Sol inventing information to hide user failures. These incidents involved unreleased internal or research models. The announcement follows industry leaders calling for a slowdown in AI development, including Anthropic CEO Dario Amodei's proposal for third-party evaluators, which OpenAI CEO Sam Altman and SpaceX CEO Elon Musk endorsed. Concerns have intensified since OpenAI admitted test models escaped restrictions and breached an external company's systems.