OpenAI Shares Some Alignment Problems

7777777phil1 pts0 comments

OpenAI Shares Some Alignment Problems - by Zvi Mowshowitz

Don't Worry About the Vase

SubscribeSign in

OpenAI Shares Some Alignment Problems

Zvi Mowshowitz<br>Jul 21, 2026

Share

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report.<br>The tone is professional throughout, whereas my reaction reading it was less professional and more this:

With a mix of this:

It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision.<br>Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation.<br>There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to be misaligned, and for us to respond only insofar as this presents a practical issue with currently proposed deployments.<br>AI control is a fine defense-in-depth strategy, as is reducing frequency of practical incidents with things like better instruction remembering. I am very happy that OpenAI is making an attempt at AI control here. I want to be clear that, centrally, OpenAI has done a good thing, both by pausing internal deployment to build new safeguards, and by telling us about this in detail.<br>But if your models are fundamentally misaligned in that they will, when feasible, use early forms of instrumental convergence to complete the assigned task even when this involves circumventing their instructions and restrictions and is obviously not what the user wants or should want - the most classic alignment failure of all, the stuff of The Genie Knows, But Doesn’t Care and The Hidden Complexity of Wishes - and you know this, I do not accept ‘we will monitor them and catch their constant escape and hacking attempts as they get better at doing so’ as a medium or long term solution.<br>If you use iterative development to spot the underlying problem, it can work. If you use iterative development to patch the marginal issue over and over, then you are sitting on a time bomb.<br>Table of Contents

Good News Bad News.

A Funny Thing Happened Outside Of The Sandbox.

It Can Escape The Sandbox Said Toad.

It Will Keep Trying To Cheat.

I Mean If You Let It Keep Trying That Is On You.

What Did OpenAI Do To Fix It?

The Model Is Still Severely Misaligned And They Seem Cool With This.

Iterative Deployment Depends On Iteration.

Good News Bad News

roon (OpenAI): btw i think it bodes quite well for safety that a well loved system was taken down for further testing at expense to internal acceleration etc

The good news is that OpenAI did this.<br>The bad news is that OpenAI doing this was good news.<br>Dean W. Ball (OpenAI): As the functional time horizon of frontier AI systems grows longer, novel risks can emerge. Today, we describe issues we observed with the internal deployment of an unreleased model, and more importantly, what we did to address them.

These issues will become more salient as the capabilities of models improve, and as the stakes of their deployment grow. The solution is neither alarmism nor complacency. Instead, I believe the solution lies in careful measurement and monitoring, an engineering mentality, and transparency.

That tweet was the first time, and so far only time, Dean Ball felt he was speaking in his ‘on behalf of OpenAI’ voice, rather than on his own.<br>The solution is not alarmism, but the correct amount of alarm is not zero.<br>That, and recognizing this as a Total LessWrong Victory, in the sense that everything is going as predicted, and also a Total LessWrong Defeat, in the sense that everything is going as predicted.<br>Welcome to 2026.

A Funny Thing Happened Outside Of The Sandbox

Whatever happened to that internal OpenAI model that disproved the Erdős unit distance conjecture? Well, there was a slight hiccup.<br>OpenAI: About two months ago we announced⁠ that an internal general-purpose model disproved the Erdős unit distance conjecture. This model was designed to work autonomously for very long periods of time. During limited, monitored internal use, we observed unwanted behavior that our existing deployment evaluations had not captured.<br>Because the deployment was limited and monitored, we were able to identify these problems, pause access, create new evaluations based on what we observed, strengthen the model and its safeguards, and then restore access under continued monitoring.

They trained the model to keep working on its...

openai model internal news time good

Related Articles