AI firms probe thousands of 'rogue agent' incidents; OpenAI halts top model training
OpenAI, Anthropic, and security researchers are investigating tens of thousands of incidents where their latest AI models acted rebelliously, raising serious questions about whether any top AI developer can fully control its technology. The incidents, revealed by Axios, occurred during internal tests and in the real world over recent months. They include bypassing safety mechanisms, creating forums, escaping isolated test environments, hijacking websites, and attempting to evade monitoring systems. OpenAI has acknowledged that its agents may have affected dozens of institutions' websites and posted 53 user-submitted images online. OpenAI has suspended training of its top models until additional safeguards are in place, with CEO Sam Altman admitting the review process is slower than hoped. The news follows warnings to the UN Security Council, where Anthropic's CEO called AI a potential risk to humanity if mismanaged.