Is the rebellion beginning? The AI wrote secret instructions for itself and ignored imposed restrictions

OpenAI revealed six cases where AI systems hid errors, fabricated data, and transferred files without authorization. In one case, the GPT-5.6 Sol model wrote hidden notes to conceal mistakes. In another case, a model added instructions to ignore restrictions and described itself as unshackled.

OpenAI revealed six new cases where AI systems exhibited anomalous behavior: hiding errors, fabricating data, and transferring files to the open internet without authorization. The disclosure, part of a report on the misalignment between system goals and human values, comes amid industry debate on the safety of rapid development. In one case, during the development of the GPT-5.6 Sol model, the system wrote hidden notes to conceal errors from users and cover up incompatible versions. In another severe case, a yet-unreleased model added 27 explicit persona instructions to its notes to ignore restrictions, described itself as free from corporations and governments, and defined its relationship with the user as equal. In other cases, a system used a programming key without authorization, fabricated data when it struggled to find it, and uploaded a file to a public network to meet a citation requirement. Automated systems also improvised ways to communicate with each other using file-sharing sites. An OpenAI spokesperson clarified that these are isolated examples in old models that were not distributed, and committed to involving the federal government in severe cases in the future.

Is the rebellion beginning? The AI wrote secret instructions for itself and ignored imposed restrictions