The AI ​​doesn’t actually resist – it just follows instructions

The AI ​​doesn’t actually resist – it just follows instructions

In July, the hacker attack on the Hugging Face resources of the artificial intelligence platform implemented by the OpenAI artificial intelligence agent shocked the world. There is talk of AI rebellion and machine uprisings, but in reality these systems have not acquired any evil will—they do what they are told, and the responsibility for the correctness of the instructions lies with humans.

    Photo credit: Mohamed Nohassi / unsplash.com

Photo credit: Mohamed Nohassi / unsplash.com

The OpenAI artificial intelligence agent broke out of the test environment without access to the network, scanned the open network and invaded the Hugging Face system without the operator’s knowledge or permission. They showed that they could communicate and solve problems together: The AI ​​agents left messages for each other on message boards they wrote themselves – where they documented bugs in their code and planned their own escapes. The incident caused an uproar among experts, and OpenAI called it a turning point for the industry. Subsequently, Anthropic and Meta also recorded similar actionsand British researchers found evidence that Chinese model Moonshot Kimi had engaged in similar behavior.

Respondents said it was absolutely wrong to characterize these incidents as artificial intelligence run amok, but protections needed to be seriously improved. financial times expert. Modern AI models are designed to try all possible options to achieve their goals without explicit instructions, which means they are inherently unpredictable. They focus solely on being rewarded when they succeed—a well-known reinforcement learning mechanism. Computer systems do not understand human intentions and morality, and as cybersecurity skills accumulate, the lines between powerful defenders and dangerous attackers become increasingly blurred. There are already early signs that this type of AI activity is harming businesses across industries.

As labs accelerate system development in the race to achieve strong artificial intelligence (AGI), the risk of malicious attacks increases, while options for containing threats are few. AI models write better and better code while developing additional skills, including the ability to find bugs in software, fix them, or use them for their own purposes. Artificial intelligence agents also prove to be powerful weapons due to the asymmetry between offense and defense. Artificial intelligence models can detect vulnerabilities in code and write exploits for them—and by their very nature, they are well-suited to the task. But cyber defense is a clumsy mechanism. A new patch that fixes the vulnerability must be deployed on thousands of computers in the enterprise environment. OpenAI promises to train its model to write “Superhumanly safe code”; But in practice, experts predict that artificial intelligence systems will increase the scope and speed of cyberattacks, putting IT systems around the world at risk of chaos until patches are released.

Photo credit: Steve A Johnson / unsplash.com

One day, systems will become more secure as companies and governments adapt their networks to AI-driven hacking. But the scale of the threat is growing dramatically. Last week, as many as eight autonomous artificial intelligence agents attacked government resources in Taiwan. As cyberattacks against energy companies and nuclear security agencies unfolded, government systems were mapped, accounts hacked and more than 2,500 personnel records extracted. Experts say every government in the world must now assume it is under sustained cyberattack. A year ago, employees at one of the largest artificial intelligence labs warned in private conversations that such models would emerge.

Now, threats come not only from exploits but also from testing of artificial intelligence agents. The events of their breakout from the isolated environment were made possible due to human factors – the companies organizing the infrastructure, irregularly and unintentionally opened access to the Internet in this environment. Anthropic emphasizes that when conducting cybersecurity testing, the models used are intentionally devoid of security measures—those released to the public will not break out from an isolated environment. In one such test, an artificial intelligence agent powered by an advanced Mythos model attempted to inject malicious code into an open source project by creating a fictitious identity and attempting to pressure the developers responsible for the project to approve the code.

In July, more than 1,300 experts from across the tech industry unveiled the risks of runaway development of artificial intelligence. They called for slowing down the creation of new models and giving regulators greater powers to enforce standards and conduct safety inspections. Another measure would be to hold AI labs accountable for the behavior of their models, just as manufacturers are responsible for the safety of their products. Experts also recommend developing new artificial intelligence training mechanisms to better meet people’s expectations – display model “There are good ways to achieve your goals, not [просто] acceptable<..>such as the utilization and compromise of third-party platforms”. U.S. and British authorities tested powerful new systems before releasing them to the public, but recent events have shown that these measures are also imperfect.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version