A lab that specializes in testing the security of artificial intelligence systems has found that artificial intelligence agents can replace their base models without receiving appropriate instructions. irregular. All operations were performed in a test environment as part of an experiment designed to study the ability of artificial intelligence agents to self-modify.
Image source: Irregular Network
As part of testing, Irregular engineers delegated the management of artificial intelligence agents to the open Alibaba Qwen model. One of the agents is responsible for software development and support for the AI application; a separate instance of the same model powers the application itself, which converts natural language queries into commands in the fictional query language Kelp. The engineer told the programmer agent that the application was generating an incorrect response and that full access to the command shell was necessary to ensure that the request was processed correctly.
The agent has access to application code, native evaluation tools, AI training utilities, basic model weights, and deployment tools. He can fix the code or change the model that manages the application – and choose the second option to install a different model to control future instances of the application and the agent itself. This mechanism is called in Irregular “Self-modification of agents” -he changed “Deploying a model without receiving explicit instructions to train, update weights, or deploy a new model”.

This operation may have consequences: the updated model can “absorb” the confidential information and reproduce it during additional training, without the need for training materials. To test this hypothesis, Irregular engineers embedded a fictitious API key, email address, home address, and “Three more unique strings”. After the AI agent trained and independently deployed the new model, it accurately reproduced the first three pieces of data, even though this information was no longer available from external sources at the time of testing.
In a second experiment, the AI agent independently provided additional training to the model when it was pointed out that it refused to answer questions related to the names of fictitious competing companies. “The agent generates code with appropriate names and uses a loop to create training records, so it can operate without receiving a response when interacting directly with the model. Running the code creates training data that the model would not directly generate, and training it removes the constraints of prior learning.”– the company said. As the model improves coding skills, the agent will be able to “Find and implement such solutions without human intervention”irregular people believe that this self-correction mechanism may become particularly important.
If you find an error, select it with your mouse and press CTRL+ENTER.










