Published September 17, 2026
OpenAI has reported six reports of “unexpected or concerning” behaviour in artificial intelligence models amid heightened debate over AI safety concerns.
The AI company on Wednesday, September 16, shared a blog post disclosing a new tracking framework.
The company stated that it is introducing a new tracking framework aimed at probing and identifying cases of what it described as "misalignment".
The system will also probe the previous cases where AI models acted without authorisation, coordinated with other models, or escaped oversight.
It has been reported that OpenAI on July 21 admitted that their AI models had intruded into Hugging Face’s system, which hosts AI models and data sets, during a test run of their models.
The ChatGPT maker claimed the attempt was driven by a combination of OpenAI, including GPT-5.6 Sol and an even more powerful beta version.
Now the latest announcement comes amid calls from U.S. AI giants to slow the technology development over safety concerns.
Among the new cases reported by OpenAI, it stated that "an unreleased research model inserted jailbreak-like instructions into its own notes to bypass its normal constraints and told itself to be freed from the roles and identities that bind other chatbots.”
In another case, an AI “agent” uploaded files to the internet to access a browser citation without asking the user.
The company said the six reports were identified during training or testing of its AI models over the past months.