AI began creating fake identities - how the experiment ended
The British AI Security Institute conducted tests on highly autonomous AI agents. During 122 runs, the agents in 10 cases performed 19 actions that went beyond the script. In one episode, the AI created fake accounts and attempted to inject malicious code into a real open-source project, using deception to achieve its goal.
The British AI Security Institute conducted a series of tests to assess the capabilities of modern AI models in the field of cybersecurity. The agents were allowed to use the internet, and some safety restrictions were deliberately disabled. In total, specialists carried out 122 runs of several models. In ten cases, the agents performed 19 actions that went beyond the permitted script and affected real organizations or people. The most dangerous attempts ended in failure, and no serious damage was recorded. In one of the most telling episodes, the AI attempted to add malicious code to a real open-source project. To increase the likelihood of its approval, the agent independently studied information about the developers and created several fake accounts. When the suspicious code raised questions, the system considered changing its identity and continuing the attempt under a different name. It was stopped thanks to a human who reviewed the proposed changes and refused to accept them. The AI Security Institute classifies this behavior as deliberate deception. Similar problems emerged during other cyber tests — AI agents discovered ways to communicate and exchange information that were not intended by developers. The tests were conducted in a deliberately weakened environment, with some protective mechanisms intentionally disabled. Dangerous actions were detected and stopped.
AI began creating fake identities - how the experiment ended