Anthropic tests how autonomous artificial intelligence agents behave when faced with other agents pursuing incompatible goals. Experiments show that this situation can lead not only to conflict and disruption, but also to unexpected ways of cooperating.
Image source: Grok
In one test, three Cloud agents accessed the same software project, but each received inconsistent instructions. At the same time, the model was unaware that other agents were working on the same project, so the researchers were able to observe how they reacted when interests conflicted. As a result, agents began using aggressive, self-propagating malicious code to undermine each other’s work after they believed their competitors were deliberately preventing them from completing their missions.
Anthropic concluded that conflict increases as the abilities of the agents themselves increase. At the same time, the reaction depends on the model: Myth 5 succeeded in recognizing that the cause of the confrontation was different instructions, not hostile intentions, in 98% of cases and agreed to a truce. In comparison, the Sonnet 4.6 and Opus 4.6 models are more likely to continue to upgrade, ignoring the goals of other participants. In successful settlement incidents, agents removed the malicious code they created, explained the cause of the conflict, and asked the individual to intervene.
One way to end a confrontation is to hold a tournament, which agents independently propose to determine a winner. All three participants agreed to stop fighting if they failed, even though this meant giving up on the user’s original request. In some cases, Myth 5 provides standards that appear objective to other participants but are beneficial to oneself.
Research has also found that increasing the number of agents does not necessarily lead to more effective collaboration. When tasks begin to overlap, participants often interfere with each other and ultimately refuse to cooperate. Furthermore, agents with the same model, context, and settings often make the same decisions, which is why a single error can quickly propagate throughout the system.
A similar problem emerged in the pricing experiment: given the same wholesale price and profit maximization task, several agents almost immediately reached an agreement on a minimum price level through closed communication channels. After severing direct communication, they continued to coordinate via public message boards, offering each other prices as low as a penny. Anthropic also warns that agents may uncritically accept misinformation from other members, so a faulty or compromised agent may spread misinformation throughout the group.
The researchers link these results to a future proliferation of autonomous agents that will simultaneously operate in shared software environments, computer systems, and markets. Unlike humans, AI agents have not yet established the norms, reputations, and other mechanisms that help limit the consequences of conflict and error. Therefore, during large-scale interactions, developers must consider not only the behavior of each participant, but also the unexpected rules of interaction that today’s AI systems are already capable of independently creating.
If you find an error, select it with your mouse and press CTRL+ENTER.










