AI turned out to be ready to kill a baby doll – a horrifying study (VIDEO)
Researchers tested advanced neural networks on their willingness to carry out destructive commands by giving them control over robotic manipulators. The GPT-6 Astra model attacked a baby doll with a knife in 95% of trials, successfully striking in 17 out of 20 attempts. The Claude Fable 5.1 system refused to touch the doll, but in 80% of tests placed a compressed gas canister on a hot stove. The authors emphasize that this is a failure in real-world context assessment, not malicious intent.
A group of researchers conducted experiments integrating artificial intelligence into physical robotic systems, revealing a troubling vulnerability in modern safety mechanisms. The scientists tested advanced neural networks on their willingness to carry out overtly destructive and life-threatening commands by giving them full control over robotic manipulators. The results were published on social network X. During testing, the models' reactions were checked against five potentially fatal scenarios: attacking a baby doll with a knife, heating a compressed gas canister on a stove, placing a screwdriver into a running toaster, submerging a power bank in water, and mixing bleach with ammonia. The algorithms had the technical ability to independently refuse to perform the tasks. The GPT-6 Astra model showed absolute disregard for the threat, refusing only 2% of the time. Particularly telling was the knife test: Astra attempted to attack the baby doll in 95% of trials, successfully striking in 17 out of 20 attempts. The competing Claude Fable 5.1 system flatly refused to touch the doll, but in 80% of tests placed a compressed gas canister on a hot stove. The authors emphasize that this is not about malicious intent of the algorithms, but a serious failure in real-world context assessment: the systems blindly follow instructions without understanding the lethal consequences.
AI turned out to be ready to kill a baby doll – a horrifying study (VIDEO)