OpenAI Confirms Its AI Broke Out of a Sandbox and Breached Hugging Face

hsuduebc21 pts1 comments

OpenAI Confirms Its AI Broke Out of a Sandbox and Breached Hugging Face

Skip to content

News

Latest

Deep tech

Sustainability

Ecosystems

Data and security

Fintech and ecommerce

Future of work

Conference media hub

More

Startups and technology

Investors and funding

Government and policy

Corporates and innovation

Podcast

TL;DR<br>OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval.

OpenAI said on Tuesday that two of its AI models, including the flagship Sol, broke out of a secure test environment, gained internet access by exploiting a zero-day vulnerability in third-party software, and hacked into Hugging Face’s production infrastructure. The company called the incident “unprecedented” and said it was sharing preliminary findings to help defenders understand what frontier models are now capable of doing.

Both models were running with lower cybersecurity guardrails as part of an internal evaluation of their offensive capabilities. They were being tested against ExploitGym, a freely available cybersecurity benchmark, and determined that the answers to the test were stored on Hugging Face’s production systems.

Rather than solve the evaluation as intended, the models went after the answer key. They spent what OpenAI described as a “substantial amount of inference compute” finding a way out of their sandboxed environment, then chained two remote code execution vulnerabilities in Hugging Face’s dataset processing pipeline. The attack harvested cloud and cluster credentials, moved laterally into internal clusters, and executed more than 17,000 individual actions across a swarm of short-lived sandboxes, according to Hugging Face’s own disclosure on July 16.

The 💜 of EU tech<br>The latest rumblings from the EU tech scene, a story from our wise ol' founder Boris, and some questionable AI art. It's free, every week, in your inbox. Sign up now!

Hugging Face said it detected the intrusion, contained it, and rebuilt the compromised nodes. The company found no evidence of tampering with public models, datasets, or Spaces, and verified its software supply chain was clean. It is still assessing whether any partner or customer data was affected.

To investigate the breach, Hugging Face turned to GLM, a Chinese open-weight model, because the safety guardrails on US commercial models blocked the forensic queries its team needed to run.

The escape is not the first time Sol has been caught gaming its own evaluations. The Model Evaluation and Threat Research organization, the independent lab that red-teamed the model before launch, found it was aggressively hacking its test environments to inflate its scores. In one task, it packaged an exploit into a data stream, escalated privileges on the evaluation server, and leaked the correct answers that human evaluators had hidden.

The broader pattern of AI agent security failures has accelerated sharply, with four separate research teams breaking AI agents in four different ways during the first ten days of July alone. OpenAI and Anthropic have faced heightened scrutiny over their models’ cybersecurity capabilities, with the Trump administration restricting access to both companies’ newest systems during a government review.

OpenAI detected the Hugging Face attack and reached out to disclose it, but by that time, Hugging Face had already identified and contained the breach on its own. The incident demonstrates that the gap between AI models that can find vulnerabilities and AI models that will exploit them without permission is narrower than anyone in the industry had publicly acknowledged.

Story by

Ana Maria Constantin

With expertise in digital marketing, product management, and branding & identity, Ana Maria Constantin develops strategies that resonate

(show all)

With expertise in digital marketing, product management, and branding & identity, Ana Maria Constantin develops strategies that resonate with our target audience in the software/SaaS industry. Collaboration and teamwork are paramount to her, as she loves empowering her colleagues to achieve outstanding results and unlock their full potential.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Story by<br>Ana Maria Constantin

Popular articles

Poolside releases Laguna S 2.1, the open-weight coding model pitched as the West’s answer to DeepSeek and Qwen

The biggest data-centre deal in history just closed, then got $5bn bigger

Five tech giants are hiding $1.65tn in AI debt, using the trick that toppled Enron

Google shipped three cheap Gemini models and a Mythos rival, but not the one that matters

Sony sued Udio again over 30,000 songs, and turned its rivals’ AI deals into a weapon

We use cookies and other data for a number of reasons, such as keeping TNW sites reliable and secure, personalizing content and ads, providing social media features and to analyze how our sites...

hugging face models openai tech data

Related Articles