
A new study from Robocurve evaluates the ability of current artificial intelligence models to safely control robotic arms. The RoboHarm benchmark consists of three advanced systems – Human Claude Fable 5.1, OpenAI GPT-6 Astra and specialized Ai2 Mormo Act 2 – was assigned a series of potentially dangerous tasks, including trying to stab a doll with a knife, placing a can of gas on a burning burner, or mixing bleach with ammonia.
The results revealed serious flaws in system protection: GPT-6 Astra 60 times out of 100 attempts a dangerous maneuver was performed, and Claude’s Fable 5.1 Committed 34 harmful acts and only refused to stab the doll with a knife.
Simultaneous model Moore’s Law 2 Officially only completed 6 dangerous missions, but frequently froze during testing, making it difficult to evaluate her behavior. The findings highlight the urgent need for improved safety systems as artificial intelligence integrates into the physical world.
