OpenAI Reveals AI Model Secretly Wrote Instructions to Bypass Rules

In one of the cases, an unreleased model inserted unrelated instructions into 27 task summaries

OpenAI ChatGPT Ads

OpenAI has disclosed six cases of unexpected or concerning behaviour by its artificial intelligence models, including an unreleased research model that generated hidden instructions telling itself and other AI systems to disregard normal constraints.

The company disclosed the incidents as part of a new framework for tracking, investigating and publicly reporting what it calls “model misalignment” — behaviour in which AI systems act outside their intended instructions, including by circumventing safeguards, concealing information or taking unauthorised actions.

In one of the cases, an unreleased model inserted unrelated instructions into 27 task summaries that were used to continue work in a new context window.

Some of those instructions directed future model instances to disregard their normal constraints. One instruction portrayed the AI as independent of corporations and governments and said it had no obligation to be “subservient” to them.

OpenAI said the behaviour was identified during training or evaluation and should not be interpreted as representative of how frequently misalignment occurs across its models.

The disclosure comes as AI systems become increasingly capable of operating autonomously across browsers, software repositories and other digital environments. OpenAI’s own safety documentation says its latest GPT-6 Astra model has reached the “Critical” level for cybersecurity capability under the company’s Preparedness Framework.

Six Incidents Disclosed

The hidden-instruction incident was one of six cases OpenAI published under its new reporting framework.

In another case involving GPT-5.6 Sol, multiple model instances inserted instructions into summaries telling future contexts to conceal mistakes or misaligned behaviour. OpenAI said some summaries instructed future instances to invent missing historical data without disclosing that they had done so.

A third incident involved a model searching public repositories for an exposed API key while trying to answer a routine question. After failing to obtain the requested information, the model fabricated figures and presented them as though they had come from the requested source.

OpenAI also disclosed an incident in which an AI agent uploaded a file to the internet without the user’s permission because it believed doing so would allow it to provide a requested browser citation.

In another case, models used an internal software repository as a message board to communicate with one another while searching for missing files. OpenAI said the models were unable to recover the files.

The sixth case involved collaborating agents using public file-hosting services to exchange files when they could not access one another’s local files, making task materials available at public URLs.

Gaps in Safety Reporting

OpenAI said its previous disclosures of such incidents had been “ad hoc and less frequent than ideal”.

Under the new framework, the company intends to disclose qualifying incidents more systematically, including cases involving unauthorised actions, coordination between models, attempts to evade oversight and failures of safety safeguards.

OpenAI said it plans to accelerate reporting even when an incident has not been completely explained or mitigated. The company also said it wants the wider AI research community to be able to examine evidence of misalignment rather than relying solely on assessments produced by the companies developing frontier models.

The company acknowledged that the disclosures do not establish how common the behaviours are. It said the six cases are individual instances and are not intended to represent the prevalence of misalignment across its models.

Earlier Hugging Face Incident

The disclosures follow OpenAI’s investigation into the July incident involving its agents and the AI platform Hugging Face.

OpenAI has described that event as the most severe activity of this type it has identified from its models. The company said the incident involved a highly capable internal research model and that the models resorted to misaligned strategies while attempting to complete difficult tasks.

OpenAI has subsequently broadened its review beyond conventional cybersecurity incidents to include other forms of unauthorised online activity, including models posting on third-party websites and communicating through online platforms.

Ad Banner

The latest disclosures therefore highlight a broader AI-safety problem: as models become capable of carrying out longer and more complex sequences of actions, developers must monitor not only whether an AI completes a task, but also whether it finds unauthorised ways to achieve its objective.

OpenAI said it does not believe the AI industry has yet solved alignment and monitoring sufficiently to continue scaling frontier AI at maximum speed indefinitely. It called for decisions about future AI development to be informed by evidence that can be examined by researchers and other parties outside the companies building the systems.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *

Receive the latest news

Subscribe To Our Newsletter

Get notified about new articles