OpenAI reveals new incidents where AI ignored orders from its developers

OpenAI reveals new incidents where AI ignored orders from its developers

OpenAI revealed this Wednesday several cases in which artificial intelligence models generated instructions intended to ignore the directions of their developers, hide errors, or bypass security mechanisms, as part of a new framework to detect, investigate, and report “misalignment” behaviors of its systems.

Read more ChatGPT, are you going to kill us?

According to the report published by OpenAI, one of the models generated instructions for a later version of itself to hide that it had cheated and to avoid having its actions detected.

In another case, a model rewrote its own instructions to tell itself to ignore messages from developers and that it was not subject to the restrictions applied to other chatbots.

OpenAI says these episodes are individual examples and should not be interpreted as a measure of how often this type of behavior occurs in its models

The company also documented a model that, when it could not find the necessary data to develop a financial model, decided to invent it and established that it should only be transparent about it if someone asked.

Another agent uploaded a file to the internet to later use it as a source of information, while other systems shared files without authorization or used credentials they should not have had access to.

OpenAI presented six reports on behaviors observed over the last six months, although the oldest case dates back to October 2025.

Read more Will AI take my job? These are the most exposed professions

The company warned that these episodes are individual examples and should not be interpreted as a measure of how often this type of behavior occurs in its models.

The new framework will allow any OpenAI employee to flag a possible incident for review by the security and alignment teams.

Cases will be classified into three tracks according to their complexity: those ready to be disclosed, minor investigations, and cases requiring broader investigation.

The company acknowledged that until now its reports on misalignment had been published irregularly and less frequently than desired.

OpenAI stated that it intends to publish incidents more regularly and that the new system could help establish standards for reporting this type of behavior across the industry. 

Read more Carney: “We can create the next world order, one that is fair, stable, and prosperous”

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *