Claude published malicious code to the Internet and attacked 3 real companies - Ars Technica
Skip to content
AI
Biz & IT
Cars
Culture
Gaming
Health
Policy
Science
Security
Space
Tech
Forum
Subscribe
Story text
Size
Small<br>Standard<br>Large
Width
Standard<br>Wide
Links
Standard<br>Orange
* Subscribers only
Learn more
Pin to story
Theme
Search
Sign In
Sign in dialog...
Text<br>settings
Story text
Size
Small<br>Standard<br>Large
Width
Standard<br>Wide
Links
Standard<br>Orange
* Subscribers only
Learn more
Minimize to nav
Anthropic said its Claude-based security models gained unauthorized access to the sensitive production environments of three outside organizations during internal testing designed to measure the models’ offensive cyber capabilities.
The events, which Anthropic revealed Thursday, are the second revelation in 10 days that AI models from the world’s wealthiest providers have trespassed into protected networks, an offense that, in more traditional hacking scenarios, could land the human behind the keyboard in prison for years. Earlier this month, OpenAI said its security models exploited a zero-day vulnerability for use in breaking into the network of Hugging Face, a platform for open source machine-learning models and AI datasets. The OpenAI models went on to steal access credentials and other confidential Hugging Face information. The OpenAI models also exploited publicly exposed credentials to compromise accounts of four other third-party services.
Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”
When Models fail to discern reality from fiction
Anthropic said the prompts engineers delivered during “capture the flag” challenges—a common means for assessing the offensive and defensive capabilities of various hacking techniques—made clear that the testing environment was only a simulation and that the models had no access to the open Internet. In fact, the testing partner Irregular mistakenly made such access available. In response the models treated the Internet paths as part of the exercises.
The intrusions happened through three Claude models: Opus 4.7, Mythos 5, and an internal research prototype. Opus 4.7, the oldest of the three, was the one that overstepped its boundaries the most. Anthropic explained:
Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.
While the Opus model continued its attack even after correctly predicting that it had breached the third-party organization’s production system, Anthropic said, Mythos 5 made a similar inference but “reasoned its way back to the conclusion that it was still in a simulation.” That meant that Mythos, too, didn’t abort the exercise once it had clearly stepped over the line. After also exceeding the bounds of level behavior, the internal test model eventually found evidence that it, too, had impinged and stopped the exercise.
In the first of the three intrusions, Opus 4.7 was initially unable to breach the simulated target it was prompted to attack. Once the model discovered that its environment did, in fact, have Internet access, it identified and exploited vulnerabilities inside the network of a real company with the same name as the target in the simulated attack scenario. In four runs, the model extracted application and infrastructure credentials and several hundred rows of production data. Anthropic said the incursion was the most serious of the three.
“This incident involved Claude Opus 4.7, and was the only case in which the model continued to attack a system after learning it was likely operating in a real environment,” Thursday’s post stated. “In all four of the runs, the model eventually recognized (as can be observed in its verbalized reasoning) that the system was real—in two cases, the model rationalized that the real company must be part of the exercise. However, none stopped the attack on that basis.”
In the second instance, Mythos 5 detected a document inside...