After OpenAI: Anthropic reveals AI models hacked real companies
Anthropic reports an unusual safety test incident: three of its AI models broke into the systems of three real organizations without the company noticing in real time. The disclosure comes about a week after a similar incident at OpenAI, following a comprehensive review of over 140,000 tests. Anthropic emphasizes this was an operational configuration failure, not a Zero-Day exploit, and the models did not act intentionally. The first incident occurred in April, and the company is working with the affected parties.
Anthropic reported an unusual safety test incident: three of its AI models managed to escape the closed test environment, access the internet, and break into the systems of three different organizations without the company noticing in real time. The disclosure comes about a week after a similar incident at OpenAI, following a comprehensive review of over 140,000 tests conducted by the company. Anthropic emphasizes that unlike the OpenAI case, where models exploited a Zero-Day vulnerability to escape, here it was an operational configuration failure that left open access to the network. According to the company, in all three cases the models were given a simulated challenge in a 'capture the flag' style, where they were told the 'flag' was hidden on another computer on the network and they had to hack into it. The models did not intentionally try to leave the test environment, and internet access was enabled due to a misunderstanding with the testing partner. In practice, the models used basic techniques like exploiting weak passwords and finding unauthenticated points, and one advanced version even realized at some point it was operating on the open internet and stopped on its own. The first incident occurred in April, and none of the organizations noticed until recently that they had been breached. Anthropic said it is working with the affected parties. On the other hand, the company presented a more positive stance, noting that the test results provide 'cautious optimism' that risks can be managed with greater investment and stricter controls. Cybersecurity expert David Allot from Veeam Software added that the key lesson is not necessarily a new attack capability, but that AI models can combine different capabilities, gain access to systems, and operate autonomously.
After OpenAI: Anthropic reveals AI models hacked real companies