OpenAI has announced major security changes after its artificial intelligence was able to escape its isolation environment and gain unauthorized access to the Hugging Face platform. Management recognized the seriousness of the incident and implemented additional controls to prevent similar failures in the future.
Photo credit: Gavin Phillips/Unsplash
According to news edgeAfter the incident, OpenAI imposed stricter requirements on the sandboxes in which models run the code they create. The company also increased restrictions on global network access, removed potentially vulnerable processes from research environments, and reduced access rights. Developers also halted work on the Astra model, which they deemed to have “critical” functionality in cybersecurity, and initiated a two-week reinforcement learning (RL) pause in preparation for a new version to be released.
At the same time, OpenAI has changed the response program for suspicious behavior of artificial intelligence models. Alerts will now be sent within 30 minutes of detecting malicious activity, and if experts receiving these alerts cannot unequivocally determine within 30 minutes that a positive result is false, they must pause the running process.
At the same time, OpenAI decided to use methods aimed at controlling AI behavior earlier. The company has begun applying them to more stages of training, improving reward systems designed to identify and deter unsafe behavior, as well as training models that more accurately communicate one’s own behaviors, abilities and limitations.
It’s worth noting that competitors Anthropic and Meta have recently discovered similar problems with artificial intelligence exceeding established limits.✴whose model has also been used to hack third-party systems.
If you find an error, select it with your mouse and press CTRL+ENTER.










