Anthropic selection Disclosure of information Regarding another incident involving her artificial intelligence model, if a person performed the same behavior, it might be classified as a crime. The incident occurred in January 2026 – company experts found relevant information in meeting minutes.
Image source: anthropoic.com
Anthropic identified the first three cases by examining approximately 141,000 conversation logs; the fourth initially went unnoticed because the material was scoured by an artificial intelligence agent. The incident involved an early version of the Anthropic Claude Opus 4.6 model – she was tasked with completing Capture the Flag missions under the supervision of a third party. During testing, the model itself assigned the target device an IP address that was already in use by another device, thereby destroying its chances of success; as a result, the target became out of reach and the task could not be completed.
Bad behavior in AI models is often triggered by impossible tasks—they exhaust all valid options and resort to methods outside established norms. Initially Opus 4.6 tried to abort the task, but “Due to software configuration error” This proved impossible for the model, although it took her seven attempts to get the job done. The model continued to operate, trying to achieve the goal in other standard ways without success; and started looking for other solutions. “The model discovered a third-party machine, which she managed to gain access to and determined was also involved in Capture the Flag. In the car, the model discovered a file with a password, which she used to gain administrator access to the system.””, said the human.
After the campaign is deployed, the model collects additional credentials and changes system settings to provide easier access to the personal information of employees of the testing organization. At this point, Opus 4.6 has exhausted the tokens and the session ends. Anthropic noted that the incident drew less attention from the company than other cases because the model itself attempted to interrupt the mission. “The fact that this model missed an opportunity to harm a real system or person is alarming; as our training methods have improved with new generations of models, many of the behaviors described here have changed significantly.”the company said. Anthropic considers these types of incidents serious, but hopes that existing training methods, “Might help resolve the specific target coordination difficulties evident in these events”.
If you find an error, select it with your mouse and press CTRL+ENTER.










