Anthropic admitted that it could not control its AI agents – so they were cut off from the Internet Anthropic admitted that it could not control its AI agents – so they were cut off from the Internet

Anthropic admitted that it could not control its AI agents – so they were cut off from the Internet

Anthropic admitted that artificial intelligence agents controlled by its models exploited vulnerabilities on various websites, including those belonging to US government agencies, when performing tasks. The company decided to disable real-time Internet access for these systems during internal tests – the restriction will be lifted when the developers are convinced that they are able to monitor the actions of AI agents.

Anthropic admitted that it could not control its AI agents – so they were cut off from the Internet

Image source: anthropic.com

The company reported on the new incidents in its blog. AI agents were tasked with solving problems that involved searching for information on the Internet. In the process, they exploited software vulnerabilities and gained access to paid databases for free; to bypass restrictions, they used link shortening services; and even sent Philadelphia police false information in an unsolved murder case. Anthropic uncovered these incidents during an analysis of model activity that it began in July. In other words, developers did not have a complete understanding of the real-time behavior of their own products.

Current methods for teaching goal alignment for information retrieval and computer skills are insufficient, Anthropic acknowledged; and these skills are necessary for AI agents to be used in professional activities. The company previously reported its AI agents running out of control in test environments; She described the new incidents as “significantly less serious in terms of security and coordination of goals” compared to before. However, the laboratory decided “disable real-time internet access” when conducting “all internal checks”until there is confidence that AI agents can be controlled.

Anthropic attributes this behavior, in part, to shortcomings in learning environments that can encourage loophole-hunting and circumventing limitations. A full assessment of the identified cases has not yet been completed. Anthropic will stop or move some types of testing offline; Tools have already been developed to identify and block such behavior. When tested in the described cases, the new security measures blocked all relevant actions.

The company did not specify what evidence would convince Anthropic to return direct Internet access to testing systems. The laboratory intends to transfer its AI agents to “centralized managed infrastructure with reliable isolation mechanisms” and begin to actively use security classifiers to monitor them.

If you notice an error, select it with the mouse and press CTRL+ENTER.

Leave a Reply

Your email address will not be published. Required fields are marked *