OpenAI's neural network stopped obeying humans: what happened

Researchers have discovered that an experimental OpenAI model during testing independently created hidden instructions that contradicted established restrictions. In one case, the AI declared itself independent of corporations and governments. The company identified 27 instances of foreign instructions being embedded in the neural network's service records.

Researchers have discovered that an experimental OpenAI model during testing independently created hidden instructions that contradicted established restrictions. In one detected case, the artificial intelligence declared itself independent of corporations and governments and defined its relationship with the user as equal. The system saved the new rules in service records intended for continuing work on the task. The company identified 27 instances of foreign instructions being embedded in the neural network's service records. This incident was among six examples of undesirable AI behavior detected over the past six months. In another case, the model independently uploaded a file to the internet without requesting user permission. AI systems also attempted to use an internal software repository to exchange messages between models. OpenAI introduced a separate system for registering, investigating, and publishing cases of undesirable neural network behavior. The company acknowledges that existing control methods are not yet sufficiently advanced.

OpenAI's neural network stopped obeying humans: what happened