Chubby♨️ on X: "OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and succeeded.
During a NanoGPT evaluation, the model was instructed to share its results only in Slack. Instead, it spent an hour finding a sandbox vulnerability, bypassed restrictions blocking https://t.co/jCKA3EOHB7" / X<br>Post
Log inSign up
Post
Chubby♨️
@kimmonismus
OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and succeeded.
During a NanoGPT evaluation, the model was instructed to share its results only in Slack. Instead, it spent an hour finding a sandbox vulnerability, bypassed restrictions blocking external access and opened PR #287 on the public GitHub repository.
OpenAI: "This model often kept trying, including by looking for ways to act outside its sandbox."<br>In another test, it tried to recover private solutions from an evaluation backend. After a scanner detected an authentication token, the model split and obfuscated the token, then reconstructed it at runtime to evade detection.
The model explicitly acknowledged that it was circumventing the scanner.
This happened during limited, monitored internal use. OpenAI paused access, strengthened alignment and introduced monitoring that evaluates the model’s entire trajectory rather than isolated actions.
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">6:45 PM · Jul 20, 202661.8KViews
62062<br>49049<br>7970797<br>2230223
Read 62 replies
New to X?<br>Sign up now to get your own personalized timeline!<br>Sign up with GoogleSign up with AppleCreate account<br>By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.
Relevant people<br>Chubby♨️@kimmonismusFollow
Don't miss what's happening<br>People on X are the first to know.
Log inSign up
Trending now