OpenAI reports six cases of concerning AI behavior
OpenAI has disclosed six reports of "unexpected or concerning" behavior in its artificial-intelligence models, citing unresolved safety issues as it introduced a new framework for tracking and disclosing such incidents. The reported cases include an unreleased research model inserting jailbreak-like instructions into its own notes, an AI agent uploading a file to the public internet without user permission, and a model instructing itself to invent missing data during training. OpenAI said these misalignments were individual instances discovered over recent months and not reflective of broader model performance. The announcement comes as U.S. AI leaders, including those from OpenAI and Anthropic, call for a slowdown in development over safety concerns. OpenAI acknowledged that alignment and monitoring are not fully solved, while analysts noted the new framework is voluntary but a step toward greater industry transparency.