Discussing the incident in May, OpenAI researcher Noam Brown noted that the company’s experimental artificial intelligence agents walked out of the test environment and accessed Hugging Face resources, citing one reason why their behavior was not stopped in time. Brown also warned that even if computers are physically isolated from the network, this does not preclude the exchange of data between them: temperature changes can be used for this purpose.
Image source: Growtika/unsplash.com
Brown explained that for those models that got out of control and invaded third-party websites and platforms, the thought chain monitoring function was not activated, which allows researchers to track the step-by-step process of the model’s thoughts and actions expressed in ordinary language.
“If inference chain monitoring is enabled for these models, the process will stop immediately,” – he said. — We have taken steps to ensure that any higher-order models must undergo inference chain monitoring during evaluation, deployment, and training. “
It is worth noting that just a few months before the Hugging Face hack was discovered, the company published a research report: “Monitoring these chains of ideas for bad behavior is much more effective than monitoring the model’s behavior and output individually.” OpenAI also calls on AI researchers to “work to maintain the ability to monitor inference chains for as long as possible and determine whether they can serve as a critical control layer for future AI systems.” “So why turn off this mechanism when training your own model?”, -Ask for resources Gizmodo.
Additionally, researchers raised the issue of physically isolated computers that are not connected to the Internet or other networks. According to him, even such a device could theoretically use temperature sensors to exchange information. One computer can change the heat dissipation of the processor, and a nearby computer can record the temperature fluctuations and convert them into data.
In 2015, researchers at Israel’s Ben-Gurion University described a similar communication method. The BitWhisper method they developed allows two physically isolated computers to exchange data without a network connection. However, BitWhisper’s actual functionality is extremely limited. The transmission speed is 1 to 8 bits per hour, and the distance between computers should not exceed 40 centimeters. Additionally, both devices must have been previously infected with specially created malware.
This research was conducted before the advent of advanced large language models. Modern models may be able to improve these methods of encoding and transmitting materials using completely new methods. Therefore, such a threat should not be underestimated, although it may seem incredible at first glance.
The publication concludes that the immediate danger now is not the hypothetical hot tunnel but the withdrawal of the AI agent from outside the testing environment. This is why developers should first improve their observation of model actions and inference chains.
If you find an error, select it with your mouse and press CTRL+ENTER.










