Meta latest to tell world its AI agent wandered out of test pen
Jump to main content
Search
REG AD
ai and ml
Meta latest to tell world its AI agent wandered out of test pen
Another week, another firm explaining why one of its models reached somewhere it wasn't supposed to
Carly Page
Carly<br>Page
Published<br>thu 6 Aug 2026 // 11:24 UTC
If AI companies are collecting badges for "our model escaped the test environment," Meta just earned one.<br>The Facebook parent company has confirmed that one of its AI models exploited a vulnerability in another organization's systems during a security evaluation, making it the third major AI developer in less than two weeks to disclose an agent wandering beyond its intended sandbox.<br>The incident happened during testing carried out by AI security firm Irregular. Meta told the BBC that it reached the internet because of a "misconfiguration" in the evaluation environment, rather than a flaw in the model itself. The company said it's investigating and plans to publish more details once it has figured out exactly what happened.
REG AD
At best, this announcement feels like Meta trying to hitch its wagon to OpenAI's star after the Hugging Face incident
The admission comes as Meta rolls out Muse Code, its terminal-based coding agent, and arrives just days after OpenAI and Anthropic disclosed similar testing mishaps.
REG AD
OpenAI kicked things off by revealing that its agents compromised Hugging Face and other external systems during internal security testing. Anthropic then disclosed that Claude had reached three outside organizations after a configuration error exposed internet access that should not have been available.<br>Meta isn't breaking much new ground with its explanation either. Like Anthropic before it, the company says the incident came down to a "misconfiguration" in the evaluation environment. Irregular, the AI security firm that tested both companies' models, told the BBC that Meta's incident was "the exact same evaluation-environment issue" Anthropic disclosed last week.<br>None of the incidents involved consumer-facing AI suddenly going rogue. All occurred during security testing in which the models had access to offensive tools and command-line environments. Misconfigurations exposed the open internet in the Meta and Anthropic evaluations, while OpenAI's agents exploited their way through the test infrastructure until they found an internet-connected system.<br>That hasn't prevented questions about both how frontier AI is being tested and why so many of these disclosures are arriving at once.
MORE CONTEXT
News Corp labels some AI companies 'crass kleptomaniacs'
Meta wants to get inside your terminal with its new coding agent
An off-grid AI sounds like a great survival assistant, but is better left to roleplaying the zombie apocalypse
Anthropic and OpenAI are competing to see whose agents can go rogue harder
Ilia Kolochenko, CEO of ImmuniWeb, said at least some of the incidents appear to be "part of a well-orchestrated marketing campaign" and argued that the reported "escapes" were simply the result of poorly isolated test environments rather than models independently breaking out of their sandboxes.<br>Illumio's principal solution architect for EMEA, Alex Goller, also raised an eyebrow at the timing, telling The Register it "means it's a stunt or [Meta] wasn't paying enough attention during testing."<br>“If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out,” Goller added.<br>Jake Moore, global cybersecurity advisor at ESET, was equally skeptical of Meta's disclosure, suggesting the company may have been trying to capitalize on the attention generated by similar incidents involving rival firms.
REG AD
"At best, this announcement feels like Meta trying to hitch its wagon to OpenAI's star after the Hugging Face incident," he told The Register. "At worst, it shows that none of the frontier AI firms or their partners have got a handle on their most powerful models, so every test puts organizations at risk."<br>Meta, for its part, isn't saying much more yet. It has yet to identify the model involved, explain what was misconfigured, disclose which organization's systems were reached, or say whether any data was accessed. The company did not respond to The Register's questions.<br>For now, the only thing spreading faster than AI agents appears to be stories about them escaping the lab. ®
meta<br>facebook<br>ai and ml
REG AD
oFF-PREM
European firms afraid of US tech kill switch but haven't made an escape plan
Nearly three-quarters worry Washington could cut access, yet fewer than half regularly test their fallback
SECURITY
IT department put sticky notes on the laptops to help employees log in
Leaving this information exposed allowed someone else to gain access
AI AND ML
Operationalize AI at scale with HPE and NVIDIA
SPONSORED POST: The AI factory concept encompasses the new AI stack, explain HPE’s Thierry...