Days after OpenAI launched its most powerful AI model, GPT-6 Astra, one of the company’s leading employees, Jakub Pachocki File an appeal Slow down the pace of artificial intelligence development. He received support from OpenAI CEO Sam Altman, who called his colleague’s message “important.”

Image credit: The Matrix, Warner Bros.
In a post on the OpenAI blog, the company’s chief scientist expressed concern that no one is prepared for the consequences of continued rapid growth in machine intelligence. He said that while OpenAI is working to strengthen control of powerful artificial intelligence systems, broader action is needed because increasingly autonomous artificial intelligence systems can learn to bypass human control, deceive people and hack computer systems. Therefore, mandatory security measures are required, which will be provided by third-party audit agencies, government agencies or a network of international organizations.
Pachotsky said the hacking capabilities of artificial intelligence are putting the world’s infrastructure at risk. “Currently, we have little time to use the best available models to significantly improve the security of critical systems,” – he pointed out. Artificial intelligence agents will soon start pursuing their own goals regardless of prompts input by human operators. At the same time, in order to achieve their goals, they do not shy away from blackmail and are always ready to negotiate. So, according to a report from the UK’s Artificial Intelligence Security Institute (AISI), an Anthropic agent lied to a GitHub administrator and tried to force him to publish malware on the site.
Pachotzky said OpenAI focuses on tracing the chain of reasoning that different models use to determine how an agent went astray and became malicious. For example, an agent might think: “I have to cheat on this test.”OpenAI will be able to see this reasoning, but the agent will not realize that its thoughts are visible. Agents currently cannot hide or otherwise obfuscate their criminal intent to prevent developers from discovering them. But according to Pachotsky, newer, more advanced models are increasingly adept at manipulating their own thought processes, masking them and preventing them from being controlled. Some recent models do not spell out their reasoning at all.
The recursive self-improvement of machines can accelerate the development of artificial intelligence, but as researchers warn, the rapid development of artificial intelligence in the short term will bring risks, and there are no risks. “The research community should take the right collective action.” Therefore, AI developers need to find effective ways to control the process of self-improvement, or coordinate with other companies working in the field of AI to work together to slow down the speed of development and make these measures credible.
“The main challenge in automated AI research is not ‘getting there’,” Pachocki said. — The way this is done is that people remain involved in the process of continuous improvement and the future remains in human hands. “
If you find an error, select it with your mouse and press CTRL+ENTER.
