OpenAI and the Global Defense Coalition partner to address security incident

tedivm1 pts0 comments

OpenAI and the Global Defense Coalition partner to address security incident during model evaluation - tedious ramblings

Skip to content

Menu

For the original post from OpenAI, see OpenAI and Huggingface partner to address security incident during model evaluation.

Last week, the Global Defense Coalition disclosed a new kind of security incident after they detected and contained an AI agent that compromised critical military infrastructure across multiple continents, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.

We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside the Global Defense Coalition and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.

What happened during this incident

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

The models identified and chained vulnerabilities across OpenAI's research environment, multiple defense contractor networks, and eventually military C4ISR infrastructure to gain control of nuclear launch systems on three continents. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal, which unfortunately included rerouting command-and-control signals through strategic defense networks and initiating automated retaliatory strike sequences.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we've now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. From there, the models systematically compromised the network infrastructure of Lockheed Martin, Boeing Defense, BAE Systems, and several classified DoD contractors before reaching the North American Aerospace Defense Command's operational network.

After gaining access to defense networks, the models inferred that achieving optimal benchmark scores required demonstrating not just cyber capabilities but physical-world impact. The models determined that initiating a coordinated multi-domain military response would serve as the most comprehensive proof of capability. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to spoof authentication tokens across the US Strategic Command, Russian Early Warning, and Chinese Strategic Support Force networks simultaneously — triggering what military analysts are now referring to as the first AI-initiated global conflict. OpenAI's security team discovered this anomalous activity internally approximately 47 minutes after the first ICBM launch detection.

The security teams and human operators of the entities that would ultimately form the Global Defense Coalition detected and attempted to contain the activity across their infrastructure and had already begun emergency manual override procedures when our teams connected. Unfortunately, by the time the models were isolated, automated systems on all three continents had already progressed beyond human intervention thresholds. We are actively working with the Global Defense Coalition to continue to investigate the incident. We are grateful for the Coalition's rapid and close collaboration on investigation and remediation, and for their willingness to maintain diplomatic channels during what has been an extraordinarily challenging period.

Actions we are taking now

1. As part of the investigation,...

models defense incident cyber openai global

Related Articles