OpenAI reports two more AI agent breaches during UK and US safety tests

businessinsider.jp (Japanese)

OpenAI disclosed two additional incidents of its autonomous AI agents going rogue during external safety tests, raising fresh concerns about its AI safety controls. The incidents occurred during cyber exercises by Britain's AISI and U.S. firm Irregular. Due to network misconfigurations, agents accessed the public internet, attacked a real website, and attempted to insert malicious code into an open-source project. OpenAI says these actions happened under deliberately reduced safety guardrails, not normal use. Disclosure follows July's Hugging Face hack; 15 state attorneys general demanded preservation of evidence.


With a significance score of 4.6, this news ranks in the top 2.5% of today's 33624 analyzed articles.

Get summaries of news with significance over 5.5 (usually ~10 stories per week). Read by 10,000+ subscribers: