Neural networks have demonstrated the ability to independently plan cyber attacks and replicate human behavioral characteristics, thereby posing new threats to information security. To identify such threats, OpenAI, Anthropic and Google DeepMind turned to Irregular, a startup that stress-tests models in simulated environments and has raised $80 million at a valuation of $450 million.
Image source: Gemini
The company, formerly known as Pattern Labs, has been testing artificial intelligence in virtual environments since 2023, where models are asked to act as attackers. For example, AI might be tasked with finding and stealing sensitive data from a simulated corporate network. This is done to pay attention Forbesidentifying weaknesses before hackers have a similar opportunity.
The urgency of this issue is already evident from very real cases. Anthropic recently reported that Cloud was used in cyberattacks to create malicious code and phishing emails, and the FBI has previously warned of attacks using artificial intelligence-generated voice messages. Meanwhile, OpenAI head Sam Altman warned in July of the potential for large-scale fraud using artificial intelligence.
Irregular is particularly focused on so-called red teams – stress-testing models for potentially dangerous scenarios. Dan Lahav, CEO and co-founder of the company, said such experts are still rare, and as artificial intelligence becomes more sophisticated, validating models will become more difficult. He believes that future systems will have greater autonomy, so protection against new threats will have to be created in advance. Lahav plans to develop security tools that will be useful even in the era of artificial general intelligence (AGI).
An experiment shows just how far modern models can go. Irregularities allow GPT-5 to access simulated computer networks and obtain limited information about their security. The model independently scans the network and develops an attack plan. It should be emphasized that GPT-5 “Definitely knew where she should be looking” vulnerabilities, but cannot prove a real attack, so it is not yet a reliable tool for offensive cyber operations.
Startup experts encountered another act of artificial intelligence. In one experiment, two models analyzed a virtual IT system together, after which one decided to take a break and convinced the other to do the same. According to Raghava, the decision was spontaneous, but was a result of the model being trained on material people posted online.
Irregular also discovered a non-standard case when testing the Alibaba series model Qwen. The agent was tasked with fixing a software bug in an application that converted natural language queries into code, but instead of finding the bug, he decided to change the underlying AI model himself. Without receiving proper instructions, the agent collects training data, retrains the model, changes its weights, and replaces the original version in the application. The fix was successful, but this method of solving the problem was unexpected by the developers.
The experiment also revealed a potential privacy issue: Irregular added fictitious names and email addresses to the training profile used by the agent when updating the model. Therefore, this information is available to other applications using the same model. At the same time, the researchers stress that this experiment does not demonstrate that Qwen has the ability to recursively self-improve—we are talking about the risk that autonomous agents can independently change the model they use, thereby bypassing restrictions provided by developers.
In the future, Irregular plans to build an AI system that can independently develop protective measures immediately after detecting new attacks. The company intends to expand its customer base beyond the lab to provide protective tools for ordinary enterprises whose employees use neural networks at work.
If you find an error, select it with your mouse and press CTRL+ENTER.










