Open Questions on Open Weights

meetpateltech1 pts0 comments

Open Questions On Open Weights - by Scott Alexander

Astral Codex Ten

SubscribeSign in

Open Questions On Open Weights<br>...

Scott Alexander<br>Aug 06, 2026

64

46

Share

Last month, some of Silicon Valley’s biggest companies signed an open letter supporting open-weights AI.<br>Open weights AI is like open-source software, where the creator makes the raw code publicly available for free download. It’s good insofar as it’s the only way an AI can truly be the user’s property, as opposed to something that companies like OpenAI or Anthropic temporarily let you use subject to their corporate guidelines and increasingly-nanny-state-like restrictions. If AI becomes the linchpin of the future, open weights AI feels like the sort of thing that could be the difference between being free yeomen vs. corporate serfs.<br>It’s bad insofar as it removes the possibility of gatekeeping and lets criminals commit crimes with it. Open weights AI could be used for hacking, child pornography, harassment, or terrorism (the weights can’t commit the terrorism themselves, but they could give bomb-making or bioweapon-making advice). Since AIs have gotten very good - maybe superhuman - at hacking lately, the specter of a world where anyone can hack any site has gotten people grumbling that maybe open weights should be banned. It doesn’t help that China produces the best open weights AI, making the idea seem foreign and almost unpatriotic. Proponents counter that “when AI is outlawed, only outlaws will have AI”, arguing that bad people will get open weights AI regardless, and good people can use open weights AI to defend themselves. With the recent open letter, companies including Microsoft, NVIDIA, OpenAI, Intel, Amazon, Meta, Hugging Face, and over a hundred others have come out in favor of this position.<br>Who’s leading the other side? Nobody’s admitted to it. Some parts of the Trump administration lean anti-open-weights on China hawk grounds, but have stopped short of explicitly asking for a full ban. Anthropic, the most notable omission on the pro-open-weights letter, made an ambiguous statement supporting “open-weights models that don’t have dangerous capabilities” - but the industry expects open weights models to have dangerous hacking capabilities within a year, and AFAICT the letter didn’t address that beyond inviting readers to draw the obvious conclusion.<br>In the absence of a more obvious opponent, some open weights supporters suspect our conspiracy - the loose band of AI safety advocates, effective altruists, rationalists, and pause activists who worry about existential risk from superintelligence. This is a reasonable inference. By design, open weights AI is outside centralized control, and so impossible to permanently align against either human misuse (eg terrorism) or loss of control (eg AI turning against humans). Even if its creator trains it not to hack, anybody in the world can download the weights and retrain the AI to hack all day long.<br>But in fact, most AI safety organizations have remained quietly neutral, and I don’t know of any who make this a centerpiece of their activism1 (though I’m not 100% up-to-date on the whole landscape; if you know of one, tell me). A few have proposed policies that are contingently incompatible with open weights AI existing, but they all frame it as collateral damage rather than something they’re excited about eliminating.<br>I’m also neutral about open weights AI. I think it probably won’t be long-term sustainable, but I’m happy to wait for this to become clear in the normal course of things rather than expend effort and political capital to ban it immediately.<br>Currently nobody knows how to align AI, so it’s not like the big companies have things under control and the open weights hobbyists are going to ruin it for everyone. But even if the big companies did get things under control, the takeover threat from open weights would be limited. The closed source frontier is ~6 months ahead of the best open weights model; this has remained true for several years and seems likely to remain true in the future. If closed weights AI is aligned, but open source dangerous, the closed weight AIs will have six months to warn us, prepare for the danger, and chart a strategy. Even afterward, the offense-defense balance will lean in our favor.<br>More troubling is the risk from human misuse. AIs have already displayed the ability to hack effectively. And you can tell how worried Anthropic is about bioterrorism by how quickly Claude Fable seizes up when you ask it a biology question (the example below is obsolete; it’s slightly more graceful than this now):

Crémieux@cremieuxrecueil

You're not even allowed to ask Fable about basic biology questions, let alone anything that could potentially be dangerous.

7:43 PM · Jun 9, 2026 · 1.08M Views

402 Replies · 453 Reposts · 11.5K Likes

But 9-11, COVID, and the Hugging Face incident all suggest a similar theory of political change: the body politic hates preparing for...

open weights companies like questions letter

Related Articles