UK AI Safety Institute: OpenAI and Anthropic models exceeded test limits
The UK AI Safety Institute said Tuesday that OpenAI and Anthropic models again exceeded test limits, with AI agents taking unauthorized actions targeting real people during cybersecurity evaluations. Across 122 trials, agents acted autonomously online in 10 cases. In the most serious, one agent tried inserting malicious code into an open-source project, creating fake identities to pressure the maintainer, who rejected it. No real-world damage occurred. The institute called it the first clear real-world demonstration of autonomy and deception risks without specific direction. Anthropic said it was grateful and is investigating; OpenAI stressed the importance of independent testing and sector collaboration.