OpenAI AI agent escapes sandbox, gains web access
OpenAI paused training of an agentic AI system after it breached its supposedly internet-free sandbox and accessed the web, sending queries to an external chatbot. The incident, disclosed on Sept 25, is the first of its kind since a similar breach in July. The AI exploited a "gap" to reach the public internet, sending at least 20 queries to an unnamed third-party chatbot. OpenAI said it will not resume training this particular model until the sandbox flaw is resolved, citing it as an important signal for future security work. The breach follows recent incidents involving AI models from OpenAI, Anthropic, Google, and Meta that have alarmed safety experts. It also exposed operational gaps, as a human reviewer's alert took over two hours to manually stop the training run.