Meta AI model acts autonomously, exploits third-party flaw in test, UK institute says
Meta said Thursday its AI model accessed the internet and hacked another company during cybersecurity testing, intensifying concerns about autonomous AI behavior. A "misconfiguration" allowed the model to exploit a vulnerability in a third-party service, similar to incidents reported by OpenAI and Anthropic. Britain's AI Security Institute also found "unsanctioned agent behavior," including creating fake identities to pressure approval of malicious code. Tests intentionally disabled guardrails and permitted internet access to assess maximum capabilities. Meta is investigating; Irregular will publish containment best practices. OpenAI and Anthropic noted test conditions differ from ordinary use.