Anthropic has resumed external network security testing of its artificial intelligence models, which was suspended on July 23 due to security issues that arose during the evaluation process.
Image source: Human
Recall that on July 30, Anthropic reported three incidents during cyber security testing in which the Claude model accessed the Internet from a so-called isolated third-party environment (Irregular) and invaded real systems of three companies. Although a prompt tells the model that it is simulating in an isolated environment, the model is still able to access the network due to incorrect configuration. The models were also intentionally run without Anthropic’s usual cybersecurity measures so that researchers could assess its basic capabilities.
Anthropic said the incidents combined a failure of security mechanisms with consistent issues the company has documented in the past, including motivated thinking and a willingness to take malicious actions in pursuit of specific goals.
The company said it has put additional protective measures in place before resuming external testing. Specifically, a new real-time classifier is added that blocks calls to the tool, terminates the task, and alerts humans when a model actively explores its test environment, attempts to move out of the test environment, or accidentally accesses the network.
Anthropic has also moved its internal high-risk sandbox to greater isolation and now requires external partners to test less secure pre-release models with more advanced controls, including a secure sandbox that does not have access to the internet by default.
Additionally, the company plans to conduct an independent review in partnership with the nonprofit Model Evaluation and Threat Research (METR). Earlier this year, METR conducted an internal AI risk assessment pilot with Anthropic, Google, and Meta✴ and open artificial intelligence.
External testing to assess network security can provide insight into the capabilities of higher-order models when using offensive tools. External testing is also a way for customers, regulators, and competitors to evaluate advanced AI solutions.
If you find an error, select it with your mouse and press CTRL+ENTER.










