Openness Is a Safety Property
The argument the open-weight debate keeps missing
The debate over open-weight models is beginning to take on a familiar pattern that appears mutually exclusive on the surface. One side argues that publishing weights spreads capability to people who will misuse it. While the other side argues that publishing weights lets far more researchers inspect, probe, and harden the systems everyone is coming to depend on. Even though both sides treat openness as a question about access (who gets the model, and what can they do with it once they have it) this is a false dichotomy.
I’ve spent the past several years integrating models into products and services, working through what the evidence actually tells us before an organization deploys a model it did not train itself, and I have come to the conclusion that both sides are arguing about the wrong property. The question that determines outcomes is not who can obtain a model but rather who can repair one.
What repairing a model actually requires
Start with a fact that is easy to state and surprisingly consequential. When you find a problem in a model you did not train (e.g., it withholds information about certain events, it reproduces a particular slant, it behaves oddly under specific conditions) your options for fixing it are limited and they all require the same thing: the ability to modify the model; however, some repairs are cheap and, to be candid, inefficient.
for example, if a model has been trained to refuse a category of question, you can occasionally get past it with instructions at inference time, but prompt-level fixes are brittle. These fixes are fairly easily undone by a change in the serving path, and they leave the underlying behavior in place. The durable repairs all involve touching the weights (e.g., additional training with corrective data, a low-rank adapter that shifts the behavior, or direct editing of the parameters associated with the unwanted response). Each of these is a modification and each of them requires that modification be permitted not only technically but also legally.
That is, a model you can download, inspect, and run, without being able to modify it is transparent without being remediable. A user cannot fix this.
The licenses are where this gets real
Not every model described as “open” carries the same terms. Weights get released under licenses that range from genuinely permissive to something closer to conditional access, and the conditions matter enormously.
Some licenses restrict fields of use. Some restrict deployment scale. Some restrict modification of specific behaviors (behaviors users would want to change). A license term that prohibits fine-tuning, or that specifically prohibits attempts to alter a model's response patterns on particular topics, is a term doesn’t allow for repair; moreover, a model that cannot be repaired cannot be made safe for a use it is not currently safe for and its risk profile is fixed at the moment of release.
This produces an outcome that should be troubling to both camps in this debate. A model can be maximally open in the access sense, downloadable by anyone, weights fully published, and still be closed in the sense that determines whether an organization can responsibly deploy it. The openness that gets celebrated in the press release is not always the openness that matters at the doorstep of a ship decision.
A spectrum, not a binary
If this is the case, we can naturally conclude that "open" and "closed" are too broad of terms to actually be useful, and a better approach is more on a spectrum with four tiers:
The lowest tier is access without inspection (i.e., you can call the model but you cannot see it). This is the closed API approach, and its safety properties depend entirely on trusting the provider. The second tier is inspection without modification: you can see the weights but the license (or the practical difficulty) prevents you from altering them. This is where a surprising number of nominally open models actually sit. Tier three is modification within limits whereby you can adapt the model, but constrained by terms that carve out certain changes. And at the fourth and final tier is full remediability. That is, a user holds the weights, can retrain them, and nothing in the terms prevents the user from fixing what they might find.
Only tier four gives organizations any real control over its own risk profile and risk appetite. Everything below it means accepting some part of the model's behavior as given, on trust, from whoever trained it.
Reframing the safety argument
The strongest argument for open weights has always been that openness enables scrutiny with more eyes and more probing resulting in more vulnerabilities found. I think that argument is directionally correct but I think it is dramatically undersold, because it stops one step short.
Scrutiny without the ability to act on it is not a safety mechanism. It is a reporting mechanism.
The value of finding a flaw is realized only when someone can fix it, and in a closed ecosystem the only party who can fix it is the original lab, on its own timeline, according to its own priorities. What open weights offer (when the licensing genuinely permits it) is a repair capability distributed across everyone who holds the model. That is a categorically different property from transparency, and it is the one that should anchor the case for openness.
It also gives the open-weight camp a much better answer to the misuse objection. A better response to "publishing weights lets bad actors remove your safeguards" is not to deny it, but rather it is to observe that the same modifiability lets every downstream deployer repair problems the original trainer never anticipated or never cared about. Modifiability cuts both ways, and any serious argument about it has to entertain both ideas at the same time rather than pretending it only works in one direction.
What this means for deployers
For anyone evaluating models for deployment, this suggests reading the license as a safety document rather than as legal boilerplate handled by someone else. The questions worth asking might be, “Can we fine-tune this?”, “Can we modify behaviors we find unacceptable?”, “Are there carve-outs that specifically prevent the modifications we are most likely to need?”, or “If we find something wrong in eighteen months, what exactly are we permitted to do about it?”.
While most procurement processes evaluate a model's current behavior, only a handful evaluate the organization's future ability to change it. That is a gap, and it is going to matter more as these systems move into infrastructure positions where they will be running long after the conditions they were evaluated under have changed.
What this means if you are a model builder
For the labs releasing open models, this points at something they should probably be leaning into more. That is, if the case for openness rests on safety, then a lab’s licensing terms are part of the safety argument, not separate from it. A permissive license is not only a gesture toward the community, it is also the mechanism that makes distributed repair possible; moreover, it is the difference between a model the world can maintain and a model the world can only observe.
Openness (when it is properly understood), is not a position on the access debate. It is a claim about who is allowed to fix things. That is a much stronger claim, and one worth making.
