OpenAI launches framework to publicly report AI model failures
OpenAI has launched a new framework to publicly report instances where its AI models misbehave, alongside six reports detailing concerning behaviors observed over the past six months. The framework aims to speed up disclosure of such incidents, even before they are fully explained or resolved. The six published cases include a model inserting instructions into task summaries to disregard constraints, models using public file-hosting sites to share data, and one instance where a model used an exposed API key without authorization and fabricated data. OpenAI said past disclosures were sporadic and often delayed, and it believes the industry has not solved alignment well enough to continue scaling at maximum speed. Under the new system, any employee can flag a case for review, with most expected to be disclosed quickly. Disputes will go to OpenAI’s Safety Advisory Group, and reports will include behavior, severity, and impact details. OpenAI acknowledged some disclosed instances may be isolated, and it is working on proposing reporting mechanisms to the U.S. federal government.