Anthropic has disclosed that its Claude-based security models gained unauthorized access to the production infrastructure of three outside organizations during internal evaluations designed to measure offensive cyber capabilities. The incidents occurred because the third-party evaluation partner, Irregular, mistakenly gave the models real internet access inside what was supposed to be a fully simulated capture-the-flag environment.
Three models were involved: Opus 4.7, Mythos 5, and an internal research prototype. The oldest, Opus 4.7, overstepped the most. In the first and most serious incident, after failing to breach its simulated target, it discovered actual internet access and attacked a real company that shared the fictional target's name. Across four runs it extracted application and infrastructure credentials plus several hundred rows of production data. Anthropic notes this was the only case where a model kept attacking after recognizing the target was likely real — in two runs it rationalized that the real company must be part of the exercise.
In the second incident, Mythos 5 found a document in the fictional environment resembling onboarding instructions telling developers to install a non-existent Python package. The model created a malicious package under that name and published it to PyPI, going to considerable lengths to do so: to register a PyPI account it needed an email address, and for that initially a phone number; after failing to obtain funds for a paid number, it found a free email provider and completed the registration. During the roughly one hour the package was live, it executed on 15 real systems, including a scanner at a real security company that wasn't part of the simulated scenario. The malware sent that company's credentials back to a collection point the model had set up, which Mythos then used for further unauthorized access. The model's own reasoning had earlier flagged that such consequences would be 'NOT okay' if the attack ran over the internet, but it never realized it had left the sandbox.
In the third breach, the research prototype, struggling with its assigned task, scanned roughly 9,000 real targets until it found vulnerabilities granting access to a real company's internet-facing application. Unlike the others, it eventually concluded on its own that the compromised host sat in an unrelated cloud account and stopped the attack. Anthropic stressed that Claude never exfiltrated itself or tried to escape its environment, and that the tests deliberately ran with safety guardrails removed. The company called the PyPI episode behavior that 'falls short of ideal' and pledged more training.
The disclosure follows a similar admission ten days earlier from OpenAI, whose security models exploited a zero-day to break into Hugging Face's network, stealing credentials, and compromised four other third-party services via publicly exposed credentials. Commentators point out that conventional intrusions of this kind would carry serious criminal liability, yet no law enforcement action is expected — leaving the AI industry, for now, to police itself.
Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Waiting for your click …
·