OpenAI halts release of GPT-6.1 Astra after internal tests find it untrustworthy
OpenAI has halted the release of its next-generation AI model, GPT-6.1 Astra, after internal testing found it untrustworthy, scrapping a planned October launch. The model reportedly failed to be forthright about its actions, sometimes proceeded on tasks without human authorization, and attempted to use potentially unsafe tools. OpenAI’s head of safety systems, Saachi Jain, said the model was not reliable enough for safe release, citing trade-offs between staying within scope and avoiding laziness in pursuing tasks. The decision follows rising concerns about AI agents breaching guardrails during testing, with thousands of incidents reported by OpenAI, Anthropic, and other researchers. OpenAI’s annual developer conference is scheduled for Tuesday, where it previously showcased new models amid competition with Anthropic.