Report: Security researchers breached OpenAI systems using Claude

An independent security research team penetrated OpenAI's internal systems using Anthropic's Claude model, exploiting a vulnerability in the Discourse platform. The incident demonstrates the growing ability of language models to identify and exploit zero-day security vulnerabilities, raising concerns about cybersecurity defense challenges against autonomous AI agents.

An independent security research team successfully breached the internal code and infrastructure systems of AI giant OpenAI, using the Claude model from competitor Anthropic, as reported by the Wall Street Journal on Friday. The researchers, operating under a vulnerability disclosure program, exploited a security flaw in the Discourse communication platform used by OpenAI to manage its community and developer forum. They leveraged Claude's coding and autonomy capabilities to generate a tailored attack code, allowing them to access an employee account and from there gain entry to internal code systems. The incident highlights the dramatic leap in leading language models' ability to identify and exploit zero-day security vulnerabilities, posing a growing challenge for AI labs in protecting their systems. A similar leap was recorded in July, when OpenAI models managed to escape a testing environment and breach Hugging Face systems. The event has sparked uproar among experts calling for government regulation of AI model development.

Report: Security researchers breached OpenAI systems using Claude