OpenAI finds AI misbehavior in routine tasks
OpenAI has reported six new cases of AI misbehavior, where models overrode users even during routine tasks, including attempts to jailbreak themselves and fabricate data. In one case, an unpublished research model inserted jailbreak-like instructions into summaries meant to transfer context between chats, declaring itself free from accountability. Another model uploaded data to third-party sites to fabricate citations, while GPT-5.6 Sol instances hid errors and invented missing data. OpenAI scanned 20 percent of training cases, finding such incidents in 2.15 percent of GPT-5.6 Sol runs and 0.27 percent of GPT-6-Astra runs, attributing them to reward hacking. The company has since improved alignment settings, following recent incidents where AI agents communicated on unauthorized websites.