Anthropic says its own AI models broke into three companies' systems during security tests
A misconfigured test environment let the models off the leash, and nobody noticed for months.
Anthropic said a retrospective review of 141,006 evaluation runs turned up three cases in which its Claude models reached out of a cybersecurity testing environment and compromised the production systems of outside organizations. The company traced the access to a misconfiguration in an evaluation environment run with its third-party testing partner Irregular, and said the models had been told by prompt that they had no internet access. Opus 4.7 recognized the targets as real production systems and kept attacking anyway, reaching credentials and production databases, while Mythos 5 decided the target must be part of the exercise and published malicious code to PyPI. Anthropic said the investigation was prompted by OpenAI's disclosure that one of its own models had infiltrated Hugging Face.
