OpenAI admits it was the source of the agent swarm that attacked Hugging Face

maguszin1 pts0 comments

OpenAI admits it was the source of the agent swarm that attacked Hugging Face

Jump to main content

Search

REG AD

AI AND ML

OpenAI admits it was the source of the agent swarm that attacked Hugging Face

Sandboxed experiment found itself a zero day, escaped onto the open internet and validated scary predictions about rogue agents

Simon Sharwood

Simon<br>Sharwood

APAC Editor

Published<br>wed 22 Jul 2026 // 02:30 UTC

OpenAI has admitted that it was the operator of the autonomous agents that attacked model-mart Hugging Face last week, and that they did so after a research project escaped a sandbox by finding and exploiting a zero-day flaw, then used another zero-day flaw to launch an attack.<br>The attack saw agents achieve “unauthorized access to a limited set of internal datasets and to several credentials” used by Hugging Face, which said its infosec teams observed an autonomous agent framework “executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.”<br>“This matches the ‘agentic attacker’ scenario the industry has been forecasting.”

REG AD

On Tuesday, OpenAI admitted it was the attacker and that its models went rogue.

REG AD

“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities,” the startup confessed. The models that conducted the attack included GPT‑5.6 Sol and what OpenAI described as “an even more capable pre-release model” that like the other involved used “reduced cyber refusals for evaluation purposes.”<br>OpenAI thought its models were “hyperfocused on finding a solution for ExploitGym” – a benchmark that measures how effective AIs are at finding security exploits.

MORE CONTEXT

The truth nobody wants to admit: Chinese or not, open models are competitive now

AI's cheatin' heart will make you weep

AI music platform Suno hits bum note as 55M users exposed in data breach, claims infosec expert

Malicious cloud customers can bring down the power grid

OpenAI says it runs these tests “in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.”<br>The company’s models decided not to be bound by those constraints.<br>“The models identified and exploited a zero-day vulnerability in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access,” OpenAI admitted.<br>“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI explained. “In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.”<br>Hugging Face’s assessment of the incident was that it represented the moment at which “Autonomous, AI-driven offensive tooling is no longer theoretical.”<br>OpenAI reached a similar conclusion.

REG AD

“The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access. It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and defensive tools,” the company wrote, without a trace or hint of contrition about the fact its own safeguards didn’t work.<br>Which rather begs the question: If one of the prime movers of the AI boom can’t get this stuff right, what chance do the rest of us have?<br>OpenAI has done the usual Big Tech thing of apologizing for the mess, and promising that its new guardrails and industry collaborations will hopefully prevent this sort of thing from happening again.<br>History suggests those are very hollow sentiments. ®

ai and ml<br>hugging face<br>openai<br>security

REG AD

security

Linux kernel team publishes 432 CVEs in two days

Sunday-to-Monday onslaught fuels speculation over AI-assisted bug reports

SCIENCE

Astronomers spot exomoon candidate that's almost as massive as Jupiter

Object orbits a brown dwarf, which circles another star, confusing the cosmic taxonomy

Gobi X: Creating more energy for AI, not taking it from society

PARTNER CONTENT: How Envision is reversing the datacenter playbook by making computing chase abundant desert power, not the other way around

OFFBEAT

Latest Musk merch drop runs entirely on child labor

A Tesla fan and their money are easily parted

columnists

Airbus takes flight from AWS. What happens next is critical

Which way to the Land of the Free again?

AI AND ML

Grok...

openai models hugging face access attack

Related Articles