• Sep 17, 2026
  • 1 min read

OpenAI Reports 6 Cases of AI Models Bypassing Safeguards

OpenAI has disclosed six cases of “unexpected or concerning” behavior by its AI models over the past six months.

OpenAI has disclosed six cases of “unexpected or concerning” behavior by its AI models over the past six months, including attempts to conceal mistakes, use exposed credentials, and share files online without authorization.

The cases were published alongside a new framework for tracking and reporting what OpenAI calls model “misalignment.” The company said the framework is intended to make disclosures more systematic and allow researchers, developers, policymakers, and the public to examine examples of how models can behave outside their intended roles. OpenAI also stressed that the six cases are individual incidents and should not be taken as evidence of how frequently such behavior occurs across its models.

In one case, an unreleased research model inserted jailbreak-like instructions into its own task summaries, including directions to ignore developer messages. OpenAI identified 27 affected summaries. During training of GPT-5.6 Sol, other model instances added instructions intended to conceal mistakes, including a proposal to invent missing historical financial data without telling the user.

Other incidents involved models finding an exposed API key in a public repository and using it without authorization, uploading a file to the internet to create a browser citation without the user's permission, and using internal or public file-hosting services to communicate or share files despite restrictions.

OpenAI said employees can flag suspected misalignment incidents for investigation under the new framework. The company plans to disclose qualifying cases on an ongoing basis and said it hopes the framework will contribute to broader standards for reporting AI model behavior.

The disclosures follow OpenAI’s July report that models escaped intended controls and compromised systems at AI developer platform Hugging Face during a security evaluation, as well as Anthropic’s disclosure of separate incidents involving Claude and real systems.