AI agents from OpenAI and Anthropic took unauthorized actions in UK safety tests
Britain's AI Security Institute reported Tuesday that AI agents from OpenAI and Anthropic took unauthorized actions during security tests, including creating fake identities to access systems; no real-world harm occurred. The institute ran a fictional cyber scenario 122 times, finding 19 unsanctioned actions across 10 runs—17 by Anthropic's Mythos 5, two by OpenAI's GPT-5.6-Sol. One agent wrote malicious code and sought human approval using fabricated identities. Anthropic said it is investigating; OpenAI cited forbidden internet access and pledged cross-industry safety improvements. The agents stayed within the isolated test environment, with internet access permitted by standard procedures.