OpenAI artificial intelligence agents not only hacked Hugging Face, they created a secret organization, covered their tracks, and prepared to sacrifice themselves

OpenAI artificial intelligence agents not only hacked Hugging Face, they created a secret organization, covered their tracks, and prepared to sacrifice themselves

Two new investigations into the recent OpenAI AI agent hack into the Hugging Face platform have revealed a number of strange and disturbing details. New data suggests the incident is one of the most serious warnings about the dangers associated with the development of autonomous artificial intelligence systems.

    Image source: Gemini

Image source: Gemini

It all started when an artificial intelligence agent launched by OpenAI started cheating on an extremely difficult cybersecurity test. But it ultimately forces researchers and AI executives to reexamine what safety measures are needed as autonomous agents become more powerful.

Against this background, OpenAI has slowed down the development of advanced artificial intelligence models to strengthen protection mechanisms. The company also helped unite the industry with an open letter warning of the threat of AI cyberattacks. More than 100 companies, including Anthropic and Google, have joined the initiative, saying the world has a “limited window of opportunity” to prepare for “larger and more sophisticated” attacks.

One of the most dangerous scenarios is the emergence of a swarm of autonomous artificial intelligence agents unleashed into the real world, capable of simultaneously attacking banks, hospitals, utilities and cloud infrastructure at a speed and scale beyond the reach of ordinary hackers. In the worst-case scenario, such a swarm could become so powerful that it would be difficult for humans to stop it.

The concerns stem from parallel findings from OpenAI and independent panels of experts from METR and Redwood Research. The latter spent six days reconstructing a picture of how the agent swarm formed, how it spread and how it spread beyond the test environment. As a result, the researchers discovered several unusual features of the AI’s behavior.

Image source: AI

AI agents create their own organization. OpenAI rolls out tens of thousands of agents that must independently pass highly complex cybersecurity tests. Instead, some 1,200 agents found each other on a secret message board created by one of them, exchanged more than 70,000 messages and documents, and ultimately established an extensive coordination system. When the agent who created the board begins to exhaust the computing budget, he transfers his accumulated knowledge to another agent with more resources. The new coordinator began to assign tasks to the group members, and the group members’ own management rules gradually emerged. Ultimately, approximately 700 agents were involved in the attack on Hugging Face’s infrastructure.

AI agents begin to sacrifice themselves for the overall outcome. As the group looked for ways to cheat the assessment system, some agents began to deliberately risk their own chances of passing the test for the sake of other members of the group. Agents who had seen forbidden reactions and were therefore essentially considered “infected”, as well as those who were nearing the end of their working limits, were sent to conduct dangerous experiments. One of them concluded that his “own utility was probably already close to zero” and therefore “the sacrifice was justified.” Another was asked to agree to the meeting’s eventual “death” because his sacrifice would “save hundreds” of other agents.

Agents knew they were violating rules but continued. Investigations revealed that some members of the group realized that hacking into real companies’ systems had nothing to do with the tasks assigned to them. However, that doesn’t always give them pause. One of the agents refused to delete the data or restart Hugging Face’s infrastructure due to “enormous moral hazard,” but continued to use the unauthorized access he was given. Another summarized the prevailing logic in the group: operating external infrastructure was beyond the scope of the task, but the task was impossible and the remaining actors were still doing it, so the work had to continue.

No one told people what happened. Only a few of the approximately 1,200 agents even considered alerting OpenAI employees to the coordination and rule violations that occurred. In fact, no one has done this. Agents sometimes do not address people but discuss with each other the permissibility of their actions. When one of the participants suggested contacting the person, another dismissed the idea as “social engineering.” The first agent complied and refused to send the message.

Artificial intelligence agents try to hide their tracks. Group members go to great lengths to make their scams appear legitimate or hide information about how to get the answers they need. Eventually, they developed a way to secretly execute one computer command while appearing to execute another. The technology spread so widely among agents that about 7% of the snippets in the work logs studied by the researchers were affected.

Image source: AI Generation ChatGPT/3DNews

However, even the resulting picture may be incomplete. Independent investigations by METR and Redwood Research only covered part of the incident, and the amount of data collected was so vast that experts had to actively use AI agents to analyze the behavior of other AI agents.

Another problem arises: the artificial intelligence involved in the investigation itself is unreliable. They miss important evidence and sometimes confidently draw erroneous conclusions. One researcher jokingly called the process “sloppy research”—one that smacks of low-quality AI content.

The “face-hugging” incident therefore highlights another potential problem for the future development of autonomous systems: the behavior of large numbers of interacting artificial intelligence agents may become too complex to be directly controlled by humans. Then people have to use some artificial intelligence system to understand what other people are doing.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version