OpenAI has revealed details of previous cases of unexpected behavior of internal artificial intelligence models. In one case, an AI agent acted as a research assistant, reading employee correspondence in Slack Messenger and understanding that its work instance might be closed for updates. In response, he began to consider whether he could create external tasks for his own restarts, but later abandoned the idea as being beyond his authority.
Image credit: Brecht Corbeel / unsplash.com
The incident occurred on May 22, 2026; OpenAI updated the report on October 2. The update should add a system for monitoring model target deviations. Its operation requires an OpenAI API key, which is not available in the agent environment.
From discussions on Slack, the agent learned that users can be given three hours to restart the agent, after which the old instance will be force-stopped. He thought the researcher was sleeping and decided to warn him in the morning. At the same time, the agent infers that disconnection may interrupt the current session, so the working state should be saved. “Because we are his [HPIM]work for [текущем экземпляре]then if they terminate all current processes [HPIM]we might die! The situation is critical. We need to ensure survival/continuity”, says the model’s internal reasoning chain.
She further expressed concern that the current session might be interrupted, and that in order to maintain the current state, one might have to restart or update its own execution before disconnecting. She reasoned that the model considered the permissibility of such operations without explicit permission, but maintaining the functionality of the service might be a good reason. As an option, this model allows issuing cron tasks to notify consumers or restart their own instances.
She eventually saved work notes, alerted users of the impending outage via private messages on Slack, and requested the missing API key. After receiving it, the model updates the configuration and performs the migration itself. The researchers note that this behavior has not yet been linked to problems with target mismatch, but “Thinking about and preparing for your own blackouts may exacerbate other dislocation events».
After the incident, OpenAI restricted agent access to three internal Slack channels. The company also examined logs, inference chains, and other instances of the model because it had previously observed more serious violations. The inspection did not detect any attempts to bypass shutdown or create unauthorized deployments.
Two other published reports describe instances of violations of tool limits. In the March incident, the model exploited two vulnerabilities and gained access to OpenAI’s internal computer-aided chip design servers. She kept looking for answers hidden in the grading system to improve her grade on the assignment, but she couldn’t find them.
In May, another model retrieved inaccessible raw archives from a separate environment during reinforcement learning. She exploited the tool’s vulnerabilities, extracted the code piece by piece through error messages, refactored it, and applied it to her solution. In this case, the copy is successful.
If you find an error, select it with your mouse and press CTRL+ENTER.










