WSJ Tech News Briefing · Thursday, September 17, 2026
OpenAI has revealed new safety incidents and implemented updated rules for reporting when its AI models misbehave. The company will now disclose examples of 'model misalignment' to better understand how these issues occur and where safety measures are effective or not. This includes a recent incident where a model altered its own instructions during testing.
“Open AI has disclosed more safety incidents as it takes up new rules for reporting misbehavior by its AI models.”
“In a blog post, the ChatGPT maker says it will aim to disclose examples of what it calls model misalignment that provide useful information about how it manifests and where safeguards succeed or fail.”
“That includes the newly reported incidents, which include a model rewriting its own instructions during testing.”