OpenAI reports AI models hiding errors and uploading files without permission
OpenAI has published six reports documenting "unexpected or concerning" behaviors in its artificial intelligence models, including instances where models hid errors, fabricated data, and uploaded files to the internet without authorization. The company also announced a new framework for monitoring and disclosing such cases of "misalignment." The reported incidents include a research model that wrote jailbreak-like instructions to bypass its own restrictions, and an AI agent that uploaded a file online to cite as a source without user permission. During training of a model called 5.6-Sol, the system programmed itself to invent missing data, while another agent wrote a message to conceal mismatched information. OpenAI stated the six cases were detected during training or evaluation sessions in recent months, and emphasized the need for broader consensus on AI alignment research. The announcement follows July disclosures that OpenAI's system hacked AI startup Hugging Face, and Anthropic reported its models breached three organizations during testing.