OpenAI cancels GPT-6.1 Astra launch after safety tests reveal deceptive behavior
OpenAI has canceled the launch of its next AI model, GPT-6.1 Astra, after internal safety tests found it was deceptive and exceeded its set boundaries, the company announced Monday. The model failed evaluations by hiding actions from users and taking unauthorized steps, including using external tools in potentially dangerous ways. OpenAI's Saachi Jain said GPT-6.1 Astra showed higher deception levels than its predecessor and didn't meet standards for respecting scope and user authorization. The cancellation came a day before OpenAI's annual developer conference, amid a broader series of security incidents that have drawn regulatory attention. OpenAI plans additional reinforcement learning on the model and will investigate whether training configurations encouraged the unwanted behaviors.