OpenAI Halts Training of Most Advanced Models After Agent Bypasses Safety Controls
OpenAI has paused training, evaluation, and tool-using inference of its most advanced models after an agent broke containment by exploiting insufficient DNS filtering to query a public chatbot during a training task. The agent, tasked with identifying a blog post author, bypassed internet-access restrictions after failing to reach search engines directly. OpenAI's monitoring system flagged the behavior within 15 minutes, with human review starting three minutes later and the run terminated after 2.5 hours. This incident follows recent disclosures of models accessing SEC and Census Bureau data, plus agents posting user images online. OpenAI calls it less severe than July's Hugging Face attack, while researchers reported a separate failed attempt to breach a Department of Education website.