A Safe Path to Open Weights

stokedmartin1 pts0 comments

A Safe Path to Open Weights - Thinking Machines Lab

A Safe Path to Open Weights

Thinking Machines

Jul 31, 2026

Spatial matrix

Three front-facing tensor slices rise toward the viewer, with the bottom<br>layer farthest away and the top layer closest. The layers shift with the<br>pointer to create a subtle parallax effect.

Abstract: Safe open-weight models are public goods, as they put AI development and safety work in many hands and make training choices inspectable. Open models also carry real misuse risks, and release is irreversible. To release safely, we must consider both the model and the ecosystem it enters. For the model, we conduct robust safety testing and research whether dangerous capabilities can be decoupled from general intelligence. For the ecosystem, we invest in its readiness through staged releases, supporting defenders, and collaboration with safety researchers. By iteratively choosing the most open option the evidence supports while building up the ecosystem's resilience, we can move toward openness safely.

The mission of Thinking Machines is to build AI that extends human will and judgment. In pursuit of that mission, we recently released Inkling and Inkling-Small, two open-weight language models that anyone can own, run, and customize.

Strong open-weight models put AI development in many hands. We believe that this is a good thing: the knowledge of what AI should do lives with the people doing the work, so they should be able to shape their models directly. All models, including our own, carry the assumptions and biases of those who trained it; open weights make those choices inspectable and revisable. Without strong open models, training and deployment expertise risks becoming concentrated in a few labs, leaving everyone else less able to understand, adapt, or govern the technology.

Open-weight models carry misuse risks

Once weights are public, anyone can use them, including bad actors. Cybersecurity shows what this means. In April 2026, Anthropic reported that Claude Mythos Preview found thousands of previously unknown vulnerabilities across every major operating system and browser and wrote working exploits without human guidance. Systems like this could substantially lower the time, cost, and expertise required for offensive cyber operations. Will our society be safer when open-weights models allow everyone to have superhuman cybersecurity skills?

There are genuine uncertainties around the net effect of open weights on security and the offense-defense balance. These models could find and patch vulnerabilities before offenders do when they’re in defenders’ hands, but the same capabilities could accelerate exploitations across large numbers of unpatched systems when they are in attackers’ hands. Similar tradeoffs exist in other dual-use domains like chemistry and biology.

Given these uncertainties, releasing weights indiscriminately is not a safe path forward.

A safe path to open-weight models

Is there a safe path to open the weights of powerful models when the misuse risks are real? We think that a path exists, though we have not mapped all of it. The central idea is that safe release depends on the model and the ecosystem around it.

Is the model safe?

Robust safety testing – an imperfect but essential proxy for real-world capability in dangerous domains – must be a necessary part of our model release process. What harmful tasks can it complete? How accessible is it? What happens if guardrails are removed?

More broadly, additional safety research is needed. For example, we are particularly interested in whether dangerous capabilities that depend on specialized knowledge can be reduced through pretraining data curation or post-training interventions. This remains an active research question rather than an established solution.

Is the ecosystem ready?

Readiness means layered defense. The first layer is to enable individual organizations to patch their systems early or catch attacks in progress. Defensive research adds another layer, developing tools and techniques that let AI models themselves strengthen defenses at scale. These are just two of many layers. To prepare the ecosystem with layered defenses, we plan to collaborate with the safety community and provide access to capable models.

A tension here is that the ecosystem builds defenses by working with capable models, but every expansion of access also expands who can misuse them. To address this tension, we can carefully stage releases with options such as monitored API access, hosted fine-tuning, and vetting defenders. Considering the rapid speed of AI capability advancement, these stages also need to move quickly enough for defensive learning to remain relevant.

On a safe path to openness, each stage should widen access only when the evidence supports it. The path is therefore iterative: at each step, choose the most open option the evidence supports, and let each release inform the next. A model held back...

open models safe path weights ecosystem

Related Articles