Nobody Was Watching: Anthropic, OpenAI, and Open Models

nedruod1 pts0 comments

Nobody Was Watching - by Ryan Baker - norabble

norabble

SubscribeSign in

Nobody Was Watching<br>Two labs skipped monitoring on their cyber evals. The same week, a hundred companies signed a letter supporting deployments that can’t be monitored.

Ryan Baker<br>Aug 04, 2026

Share

The Monitors Were Off Again

Last week I wrote An OpenAI Model Escaped Its Sandbox. Where Was the Observer?. This weekend we learned that the practice of disabling observational monitors was true at Anthropic too.<br>Several defense-in-depth measures, on both our side and our partner’s, could have prevented these incidents, or at least reduced their likelihood of occurring. Careful validation of all internet access paths before evaluations began and real-time monitoring of the evaluation logs would have helped to surface the problem sooner. Both we and our partner also could have reviewed evaluation transcripts or network logs more thoroughly. It’s also possible that a prompt which told Claude it did have internet access would have changed how Claude behaved when it came into contact with real systems.<br>Source: Investigating three real-world incidents in our cybersecurity evaluations<br>Anthropic, July 30th 2026

Anthropic commits to doing better. Specific to this aspect:<br>norabble is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Subscribe

First, evaluation environments that involve powerful autonomous capabilities also require significant controls. Safety testing happens before a model is released precisely because we don’t yet know what it is capable of. Evaluation environments increasingly need to be held to the same security standard as any other system our models run in.

The shared failures of what seems from the outside like a basic failure in design, makes you wonder about the overall corner cutting. While this is not entirely parallel to the concerns mentioned in the letter from Frontier Lab employees, Pacing the Frontier, it’s hard to see it as unrelated.<br>AI could help create a dramatically better future, but that outcome is not guaranteed. The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.<br>To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.<br>Building on work already underway to monitor frontier model releases:<br>“We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”<br>Source: Pacing the Frontier<br>1,346 employees of frontier AI companies, July 2026

It’s a bit of a long-shot here, given the state of international relations right now. But if you don’t try, what else are you going to do?<br>A Nuanced Position, Poorly Received

The internet is wrong about Amodei’s stance on open-weight models. I’ll call out some excerpts, but please read the whole thing if you worry these are out of context.<br>… Anthropic has never advocated for a ban on open-weights models.<br>… Protectionist bans would not address my most serious national security concerns.<br>… Open-weights models—it does not matter whether they come from China or anywhere else—do potentially present a higher risk than closed models, because it is very difficult to apply guardrails to them or monitor their usage, and once weights are released they cannot be withdrawn.<br>… All sufficiently capable models, open and closed, should go through mandatory safety testing. The best way to address threat #2 is to just directly test models for cyber, biological, and alignment risks before release.<br>Source: Our position on open-weights models<br>Anthropic, July 27th 2026

It’s unfortunate that the nuance of the difference of opinion here is so poorly understood. Amodei is absolutely right that releasing models as open-weight carries risks. Even the letter supporting open-weights acknowledges this.<br>To be sure, open weights carry real and distinct risks. Once released, the weights are beyond the original developer’s control, and modified versions are difficult to trace or reverse. But the right response to this risk is not to prohibit open weights. In a world where cybersecurity attackers use advanced AI, defenders need access to models with comparable capabilities so they can detect, simulate, and respond to emerging threats.<br>Source: Open Weights and American AI Leadership<br>270 signatories, July 24th 2026

Should we ignore those...

open models weights frontier anthropic real

Related Articles