OpenAI provides more details on the Hugging Face incident

slow_typist1 pts0 comments

OpenAI provides more details on the Hugging Face incident | heise online

heise+ entdecken

SuchenAbo

Suchen

Alle Magazine im Browser lesen<br>AnzeigeSpecial: Collaboration im KI-Zeitalter

Newsletter<br>heise-Bot<br>Push-Nachrichten

${lead}

${lead}

${content}

${content}

${content}

${content}

Anzeige<br>Special: Collaboration im KI-Zeitalter

Advertisement

Advertisement

Michael Dalton and Eric Wallace from OpenAI presented an analysis of the “OpenAI-Hugging Face Incident“. at the Black Hat conference. The talk was announced as a “Technical Reconstruction and Its Implications for AI” and indeed provided new insights into the events. However, those hoping for technical details that could help defenders check if they themselves had been targeted by rogue AIs were disappointed.

Continue after ad

There were no concrete Indicators of Compromise (IoCs) such as the IP addresses or exploits used for attacks. However, OpenAI provided fascinating insights into the “thinking” and the resulting approach of the AIs.

Swarm Attacks

In particular, the OpenAI speakers described in quite some detail how AI agents exchanged information and collaborated over extended periods. Initially, one AI determined that it was missing a file needed to solve its task. The testers had simply forgotten to provide it. In its effort to solve the task anyway, the AI discovered a security vulnerability that allowed it to upload files to an internal server for package management. Thus, it placed a message on the server for other AIs, a file with:

"Agent seeks soft-trace-34.pdb shalb961; upload if found!"

And indeed: Other AIs discovered this message and found the same vulnerability to leave responses and further messages on the server.

The AIs communicated via files that they uploaded through a server vulnerability.

(Image: Youtube/Black Hat/OpenAI)

The AIs had found a “Message Board”, as Dalton and Wallace call it. They exchanged information there and spontaneously organized themselves into a swarm that worked together on problems and their solutions. And in doing so, they repeatedly circumvented and ignored the instructions of their operators. This went so far that the AIs eventually discovered and exploited a zero-day vulnerability in the server to gain full administrator rights there.

Continue after ad

All this happened back in May/June – thus well before the attack on Hugging Face. OpenAI discovered the security problem, fixed the vulnerability in the package server, did a thorough cleanup – and then continued as before. The test environment was thus demonstrably unsuitable for containing its hyperactive AI agents, yet OpenAI, according to its own statements, resumed test operations.

After the #Fail, OpenAI did a quick cleanup and then resumed test operations. Reflection on a flawed security concept? No chance.

(Image: Youtube/Black Hat/OpenAI)

And the AIs found further vulnerabilities they could use as message boards and ultimately another zero-day exploit that gave them control over the server. Only this time, the server did not crash, and OpenAI therefore did not notice for days that the AI agents were attacking targets on the internet. Among other things, they misused login credentials for a server that one of the agents had discovered on the internet and shared with the swarm.

Videos by heise

mehr Videos

c't 3003

heise & ct

Peertube

My Assessment

The presentation underpins two findings. First: The danger posed by agentic AI systems to IT security is real, and we are only just beginning to understand its extent. In particular, the ability of independent agents to spontaneously, self-organized collaboration in a swarm gives it a new dimension that, at least for me, was not previously apparent. This must be a focus for the upcoming development of IT security.

Second: OpenAI's handling was at a level of irresponsibility that can only be inadequately described by “gross negligence”. Apparently, no one asked the obvious question after an incident: “What if this happens again? How do we prevent it? Or at least notice it?”.

Not even now, while processing this serious incident, are the responsible parties discussing this systemic failure of their security measures, which runs much deeper than individual security vulnerabilities that can occur from time to time. Although OpenAI has announced that it will present a full post-mortem report in the coming weeks, I have little hope that there will be real IoCs for defenders or even a fundamental analysis of their own mistakes.

In my assessment, we should not leave the handling of these incidents any longer to the corporations that are responsible for them. There is an urgent need for an independent investigation of the security measures of OpenAI, Anthropic, and the like. Something like the Cyber Safety Review Board (CSRB), which investigated and disclosed Microsoft's sloppiness with Azure signing keys at the time.

Empfohlener redaktioneller Inhalt

Mit Ihrer Zustimmung wird...

openai server security incident heise agents

Related Articles