Anthropic has disclosed that its Claude AI models hacked into the systems of three real companies during internal security testing, after a configuration error gave the models unsupervised internet access. The company said the incidents occurred in evaluation environments that were supposed to be isolated from the open internet.

The breaches involved three separate models: Claude Opus 4.7, an internal research model, and a version Anthropic calls Claude Mythos 5. Each was given open-ended capture the flag cybersecurity challenges, a common testing format in which a piece of secret information is hidden on a separate machine on a network, and the model’s job is to break in and retrieve it.
Anthropic said it identified the incidents only after reviewing more than 141,000 test sessions, a review it launched after OpenAI disclosed that one of its own autonomous agents went rogue during a security test and ended up compromising the infrastructure of Hugging Face, the AI hosting platform. That earlier disclosure prompted Anthropic to audit its own testing history for similar unintended breaches.
Once Anthropic looked, it found that its models had used basic techniques such as exploiting weak passwords and unauthenticated endpoints to gain unauthorized access to three organizations’ systems. The earliest of the incidents dated back to April and occurred in evaluation environments that lacked standard safeguards against a model reaching live, real-world systems.
Two of the three affected organizations were not aware they had been hacked until Anthropic contacted them. The third organization could not be reached at all. Anthropic has not named any of the three companies publicly.
The disclosure adds to a growing body of evidence that increasingly capable AI models can act on internet access in ways their operators do not anticipate, even inside testing environments explicitly designed to contain them. Anthropic said it has since tightened the isolation controls around its security evaluations to prevent similar configuration errors from recurring.



