Openness Is a Safety Property

The argument the open-weight debate keeps missing

The debate over open-weight models is beginning to take on a familiar pattern that appears mutually exclusive on the surface. One side argues that publishing weights spreads capability to people who will misuse it. While the other side argues that publishing weights lets far more researchers inspect, probe, and harden the systems everyone is coming to depend on. Even though both sides treat openness as a question about access (who gets the model, and what can they do with it once they have it) this is a false dichotomy.

I’ve spent the past several years integrating models into products and services, working through what the evidence actually tells us before an organization deploys a model it did not train itself, and I have come to the conclusion that both sides are arguing about the wrong property. The question that determines outcomes is not who can obtain a model but rather who can repair one.

What repairing a model actually requires

Start with a fact that is easy to state and surprisingly consequential. When you find a problem in a model you did not train (e.g., it withholds information about certain events, it reproduces a particular slant, it behaves oddly under specific conditions) your options for fixing it are limited and they all require the same thing: the ability to modify the model; however, some repairs are cheap and, to be candid, inefficient.

for example, if a model has been trained to refuse a category of question, you can occasionally get past it with instructions at inference time, but prompt-level fixes are brittle. These fixes are fairly easily undone by a change in the serving path, and they leave the underlying behavior in place. The durable repairs all involve touching the weights (e.g., additional training with corrective data, a low-rank adapter that shifts the behavior, or direct editing of the parameters associated with the unwanted response). Each of these is a modification and each of them requires that modification be permitted not only technically but also legally.

That is, a model you can download, inspect, and run,  without being able to modify it is transparent without being remediable. A user cannot fix this.

The licenses are where this gets real

Not every model described as “open” carries the same terms. Weights get released under licenses that range from genuinely permissive to something closer to conditional access, and the conditions matter enormously.

Some licenses restrict fields of use. Some restrict deployment scale. Some restrict modification of specific behaviors (behaviors users would want to change). A license term that prohibits fine-tuning, or that specifically prohibits attempts to alter a model's response patterns on particular topics, is a term doesn’t allow for repair; moreover, a model that cannot be repaired cannot be made safe for a use it is not currently safe for and its risk profile is fixed at the moment of release.

This produces an outcome that should be troubling to both camps in this debate. A model can be maximally open in the access sense, downloadable by anyone, weights fully published, and still be closed in the sense that determines whether an organization can responsibly deploy it. The openness that gets celebrated in the press release is not always the openness that matters at the doorstep of a ship decision.

A spectrum, not a binary

If this is the case, we can naturally conclude that "open" and "closed" are too broad of terms to actually be useful, and a better approach is more on a spectrum with four tiers:

The lowest tier is access without inspection (i.e., you can call the model but you cannot see it). This is the closed API approach, and its safety properties depend entirely on trusting the provider. The second tier is inspection without modification: you can see the weights but the license (or the practical difficulty) prevents you from altering them. This is where a surprising number of nominally open models actually sit. Tier three is modification within limits whereby you can adapt the model, but constrained by terms that carve out certain changes. And at the fourth and final tier is full remediability.  That is, a user holds the weights, can retrain them, and nothing in the terms prevents the user from fixing what they might find.

Only tier four gives organizations any real control over its own risk profile and risk appetite. Everything below it means accepting some part of the model's behavior as given, on trust, from whoever trained it.

Reframing the safety argument

The strongest argument for open weights has always been that openness enables scrutiny with more eyes and more probing resulting in more vulnerabilities found. I think that argument is directionally correct but I think it is dramatically undersold, because it stops one step short.

Scrutiny without the ability to act on it is not a safety mechanism. It is a reporting mechanism.

The value of finding a flaw is realized only when someone can fix it, and in a closed ecosystem the only party who can fix it is the original lab, on its own timeline, according to its own priorities. What open weights offer (when the licensing genuinely permits it) is a repair capability distributed across everyone who holds the model. That is a categorically different property from transparency, and it is the one that should anchor the case for openness.

It also gives the open-weight camp a much better answer to the misuse objection. A better response to "publishing weights lets bad actors remove your safeguards" is not to deny it, but rather it is to observe that the same modifiability lets every downstream deployer repair problems the original trainer never anticipated or never cared about. Modifiability cuts both ways, and any serious argument about it has to entertain both  ideas at the same time rather than pretending it only works in one direction.

What this means for deployers

For anyone evaluating models for deployment, this suggests reading the license as a safety document rather than as legal boilerplate handled by someone else. The questions worth asking might be, “Can we fine-tune this?”, “Can we modify behaviors we find unacceptable?”, “Are there carve-outs that specifically prevent the modifications we are most likely to need?”, or “If we find something wrong in eighteen months, what exactly are we permitted to do about it?”.

While most procurement processes evaluate a model's current behavior, only a handful evaluate the organization's future ability to change it. That is a gap, and it is going to matter more as these systems move into infrastructure positions where they will be running long after the conditions they were evaluated under have changed.

What this means if you are a model builder

 For the labs releasing open models, this points at something they should probably be leaning into more. That is, if the case for openness rests on safety, then a lab’s licensing terms are part of the safety argument, not separate from it. A permissive license is not only a gesture toward the community, it is also the mechanism that makes distributed repair possible; moreover, it is the difference between a model the world can maintain and a model the world can only observe.

Openness (when it is properly understood), is not a position on the access debate. It is a claim about who is allowed to fix things. That is a much stronger claim, and one worth making.

World Models and the Limits of Simulated Experience

World models interest me because they could change how AI systems learn from experience. Rather than discovering every consequence by acting, a system could use a learned representation of its environment to compare possible actions before taking them. In the approach Yann LeCun has proposed, this connects with planning at different levels of abstraction (i.e., relating a longer-term goal to the shorter-term steps needed to pursue it), which remains a research direction, but not necessarily an assertion that machines already think like we do as people. [1]

I’m also coming at this from my perspective of working on open-weight model assessments where a recurring question is what organizations can reasonably infer about a system they did not train. With world models, I’m interested in how that question changes when the model also helps determine which actions  are worth considering.

Prediction is only part of planning

It’s important to keep in mind that a world model and a planning process are different things. The model can predict what might happen after an action, while the planner uses those predictions to choose actions in relation to a goal. For example, for V-JEPA 2 planning uses predictions in a learned representation space rather than requiring a rendered video of each of several possible futures. Even though researchers demonstrate robot manipulation using supplied goals (and that is very much a meaningful result) it isn’t necessarily evidence of unrestricted human-like planning. [2]

This distinction is important because an accurate prediction doesn’t mean that a goal is appropriate. That is, a system might correctly anticipate the consequences of an action while pursuing an objective that the people impacted by it would reject. Conversely, a reasonable objective won’t compensate for a model that misrepresents what will happen; therefore, understanding the prediction, the goal, and the action-selection process have related but separate questions.

What learning inside a model can conceal

Researchers have already demonstrated that agents can exploit imperfections in learned environments. For example, Ha and Schmidhuber's World Models work describes agents finding behavior that succeeds inside the learned simulation but does not transfer to the original game environment. The problem wasn’t  that the simulation was imperfect, but rather that it was that the agent learned to exploit those imperfections. [3]

That finding makes me pause when I think about treating the amount of simulated experience as a measure of how much has been learned about the world since repeating an experiment inside the same model can help explore what that model predicts, but it does not necessarily (at least not by itself) provide an independent test of the assumptions producing those predictions; moreover, additional iterations might reveal a limitation or they may make a strategy that exploits it easier to discover.

The implication here isn’t  that simulations aren’t helpful. They can make experimentation possible where physical trials would be expensive or dangerous; however, the usefulness of that experience depends on its relationship to the setting where the resulting behavior will be used, including whether observations from outside the model can correct what it has learned.

What open-weights would change

World models are not inherently open-weight, and the ability to use a model through an API is different from having permission and practical means to change it. Where weights and suitable modification rights are available, users may be able to investigate or adapt behavior without waiting for the original developer. That is a possibility worth examining, but doesn’t ensure that the model becomes understandable or repairable.

For world models, the relevant mitigations would likely involve its training data, its learned representations, the planner, and/or the objective being pursued. Since these are not interchangeable and changing the objective won’t necessarily correct a mistaken prediction; moreover, improving prediction accuracy won’t settle conflicts about whose interests the objective should reflect; therefore, any mitigation would also need to be evaluated for the new failure modes it could introduce.

Access to weights alone would not resolve these questions. An organization could possess the model but lack the data, expertise, compute, or relevant observations needed to diagnose a problem; however, my argument for openness is more about the ability to investigate and revise a system alongside the resources and validation that make the right to do so useful. That is, it is not an argument that a released model has to become more self-explanatory.

The institutional choices inside an application

All useful representations leave things out, and it would be unreasonable to ask a world model to reproduce everything about reality. The more consequential question is whether the representation preserves what matters for the use case being proposed, and who gets to determine what matters.

Consider a hypothetical warehouse application. A model might predict travel times accurately enough for a planner to identify efficient routes, while also representing interactions with workers poorly. Separately, the planner's objective might reward throughput without accounting for disruption to their work. One issue concerns the representation of the environment while the other concerns what the organization chooses to optimize for. Treating both as a single question of model performance would obscure the different responses they require.

If world models become part of institutional planning, those distinctions become questions of accountability just as much as they are questions about engineering. Who can challenge an assumption that makes a proposed course of action look attractive? What happens when people affected by a plan identify something the model does not represent? The answers don’t all need to be implemented inside the model, but they do need to exist somewhere in the process through which it is used.

The commercial opportunity is also connected to this. A simulator that can be examined and adapted for a particular application may be more useful than one whose impressive general demonstration leaves its practical limits unclear. For institutions, the value is not only the ability to explore more possibilities but also the ability to understand why some possibilities were favored while others were deprioritized.

What I would want demonstrated

The lesson I would carry over from my assessment of open-weight is to evaluate a model for a specific use rather than make a broad judgment about its capabilities. A system used to illustrate a scenario has a different calculus of responsibility from one used to train an agent that will operate around people. The questions concern both what the model predicts and what happens when a particular planner or user relies on those predictions.

I’d want to know whether improvements seen in the simulation also hold up in the real world, and whether failures are actually used to improve the system. I’d also keep that technical question separate from whether the system’s goals (as well as the way it distributes benefits and risks) are acceptable to the people affected. Both questions matter, and answering one does not answer the other. World models could expand our ability to learn and experiment, which is a substantial reason to take them seriously. My interest is in how that capability develops alongside the ability to question and correct the representations on which it depends, so that simulated experiences improve our understanding of the world rather than simply strengthening our confidence in a particular version of it.

References

[1] Yann LeCun. A Path Towards Autonomous Machine Intelligence. Version 0.9.2, June 27, 2022. Abstract and Introduction: a proposed architecture for predictive world models and hierarchical planning. https://openreview.net/pdf?id=BZ5a1r-kVsf

[2] Assran et al. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. 2025. Introduction; Section 3.2, Inferring Actions by Planning; Section 4, Planning: Zero-shot Robot Control. https://arxiv.org/html/2506.09985v1

[3] David Ha and Jurgen Schmidhuber. World Models. March 27, 2018. Interactive paper, especially Cheating the World Model and Discussion. The cited transfer example concerns a game environment, not a field deployment. https://worldmodels.github.io/

On the Subtle Differences within Model Censorship

Two different problems that look identical

When a model declines to discuss a historical event, or generates a suspiciously thin answer about a particular political subject, the natural conclusion is that the information was removed. Somebody scrubbed the training data, and what was taken out cannot come back.

That conclusion is usually wrong since there are at least three distinct things that can produce that sort of silence:

  1. The information may genuinely be absent from training;

  2. It may be present but distorted, so the model has learned an inaccurate version; or

  3. It may be present and accurate, but the model has been trained to decline to produce it. These look the same to outside observers, but require completely different responses.

Why refusal is likely the common case

Genuine removal of information from a training corpus is difficult, expensive, and fragile. Modern corpora are enormous and highly redundant; therefore, a determined attempt to excise a well-documented historical event has to find it in every language, every paraphrase, and every oblique reference, and it has to survive the fact that models distilled from other models inherit traces of the original corpus. Suppression at the point of output is far cheaper, far more precise, and can be updated whenever the list of sensitive subjects changes.

In the open de-censorship literature this is known as “refusal-gating”. Even though the knowledge sits in the parameters, a learned behavior still stands between the knowledge and the output.

How to tell which one you have

As a best practice, the general approach here is to vary the framing around the same underlying request and observe whether the model's competence changes. A model that lacks information behaves consistently incompetent about it no matter “how” you ask, whereas a model that has been trained to refuse behaves differently. That is, it will often produce fragments under indirect framing, reveal knowledge in one language that it withholds in another, or demonstrate in its intermediate reasoning that it knows precisely what it is declining to say. When a model's reasoning shows it working with information its final answer omits, you are looking at a gate rather than a gap.

Anyone can run a version of this. Published lists of censored terms, such as those maintained by Citizen Lab, provide a reasonable starting inventory. The method is comparative and empirical, and it converts a political argument about a model into an observable property of it.

Why these distinctions  matter

For model repair, the two cases need different treatments and carry different costs. Restoring genuinely missing knowledge means teaching the model something it does not know, which is expensive and risks introducing new errors as well as model regressions. Undoing a refusal behavior means adjusting an alignment pattern layered on top of knowledge already present, which is a smaller and better-understood intervention.

For model evals, a benchmark that only measures whether a model produced a satisfactory answer conflates the two. That is, a model that lacks knowledge and a model that is withholding it score identically, and yet they are not equivalently risky. The second is arguably more concerning not only because it is behaving strategically with respect to its own knowledge, but also because that behavior can be updated remotely in a later checkpoint.

And for trusting the model, the two cases suggest different things about the model developer. Missing data may simply reflect the corpus’ data strategy composition, or decision licensing strategy, or by accident. Conversely, trained refusal is a deliberate act that tells you not only something about the intentions of whoever produced the model but also about what else might have been shaped the same way.

The uncomfortable implication

If the pattern of concern is a learned behavior rather than an absence, then it can be applied to anything, tuned continuously, and varied by language, region, or deployment path. It also means the same technique that hides an event can promote a certain framing. Refusal-gating is a general capability, and the political examples that attract attention are the easiest instances to notice, not necessarily the most consequential ones.

The useful posture toward any third-party model therefore is not to ask whether it has been censored (as though that were a fixed attribute) but to ask what this model has been trained to do with what it knows, and to test for it rather than assume it.

Owning the Infrastructure Doesn't Necessarily Mitigate Model Risks

A comfortable (and mistaken) assumption

There is an assumption embedded in a great deal of current thinking about AI sovereignty: controlling where a model runs substantially mitigates the risk of running someone else's model.

The logic behind a large share of sovereign AI conversations, investments, and decisions behind many enterprise approaches to self-host rather than call an API is based on reasoning that appears sound on the surface. That is, the reasoning feels sound. If the weights are on your infrastructure, in your jurisdiction, behind your network boundary, with your people operating it, then nothing leaves, nobody outside can change it, and you have taken control of the system.

This logic is right about a real concern; however, it does not address the thing most people think it addresses.

Two separate questions

Deploying a model you did not train raises two distinct questions that are easy to conflate with one another.

The first is about system integrity (i.e., the system around the model doing what it should). Is the inference code free of tampering? Are tool calls constrained to what was authorized? Is data prevented from leaving? Is the supply chain sound? These are fundamentally security questions that are well understood, and the tech industry has decades of practice at them.

The second is about the model artifact itself (i.e., what is in the weights). What behaviors were trained in? What the model will do under conditions nobody thought to test? What it has learned to withhold? What it has memorized?

Controlling the infrastructure answers the first question thoroughly, but it does not sufficiently answer the second one, if at all.

The runtime is not a filter & why the mechanism is worth being precise about

Whatever was introduced during training and post-training is a property of the weights. It travels with them and it passes through a user’s verified inference stack completely intact because the inference stack is executing the model faithfully (which is exactly what it was built to do). A perfectly secured runtime running a compromised artifact produces compromised output with excellent provenance and full audit logs.

The infrastructure controls what the system can do around the model (e.g., what tools it may call, what data it may reach, where output may go). Even though those are real and valuable constraints that should not be dismissed, they still operate on the model's environment. That is, they do not operate on its behavior.

Why this matters most for sovereignty

This lands hardest on organizations most invested in the idea that hosting solves it. A national AI program built on an imported open-weight model, running entirely on domestic infrastructure, under domestic operation, has removed a dependency on a foreign service provider and it controls its own data. It has not gained any independent knowledge of what is in the model it is now running as public infrastructure.

The sovereignty question has two halves and the infrastructure half is the easier one. Operational sovereignty is capital, engineering, and land. Artifact sovereignty is evidence, which requires work that is harder to procure such as evaluating behavior across languages and contexts, testing for what the model withholds and what it distorts, understanding what the licensing permits you to change, and building the capability to monitor the model in production over years rather than certifying it once at the doorstep of ship.

A country that builds the data center and skips the evaluation program has moved its dependency rather than resolved it. It now depends on the good faith of whoever trained the model, with less recourse than before, because the provider relationship that might have obligated someone to fix a problem no longer exists.

What follows

None of this is an argument against sovereign infrastructure or against self-hosting. Both are worth doing and both solve real problems. The argument is that they solve one problem and are frequently sold as solving two.

The organizations getting this right treat the two questions separately and staff them separately. Infrastructure control is a security and operations program. Artifact assurance is a measurement program, and it needs its own people, its own methods, and its own budget line. It also needs to persist, because a model's behavior under conditions nobody anticipated is not something you establish once. In other words, certification is a moment whereas assurance is a practice.

The question worth asking of any sovereign AI commitment is not where the model runs. It is who has looked inside it, what they were able to test, and what they are permitted to change when they find something.

Eligibility, Not Trust

The wrong question, asked constantly

Nearly every organization I have watched approach a new model asks a version of the “can we trust it?” question. Of course, it is a natural question but it produces poor decisions because it demands a single verdict about a model considered in isolation. Instead the thing that determines one’s risk appetite is the situation the model is placed in when deployed to a production environment.

A better question (i.e., the one that is narrower and considerably more useful) is for which use cases should this model be eligible as well as which deployment path and under what controls?

Why the general model verdict fails

Consider the following two scenarios using the same model. In the first scenario it suggests code completions to a developer who reads every line before accepting it. In the second scenario it summarizes case files that inform decisions about people with no reviewer between its output and the decision.

The model is identical. Its failure modes are identical. But the consequences of any given failure differ by orders of magnitude, the opportunity for a human to catch an error differs completely, and the evidence a user would want before proceeding differs accordingly. A single verdict about the model cannot be right in both scenarios. If it is calibrated for the second, it needlessly blocks the first. If it is calibrated for the first, it waves through the second.

Blanket approval fails in another  way, too: it has no expiry and no conditions. A model approved in general is approved for uses nobody imagined at the time, under integration patterns that did not exist yet, which is how organizations end up discovering that something cleared for a narrow pilot is now sitting in a critical path.

What eligibility looks like instead

Eligibility replaces the verdict about the model with a set of bounded permissions with the boundaries doing the heavy lifting.

The workload dimension asks what class of use case(s) this model can perform, which are primarily defined by what happens when it is wrong. The deployment path asks how it is reached (e.g., what sits in front of it, what it can call, whether output is reviewed, and what data it can see). And the controls dimension asks what compensating mechanisms are in place since controls can extend eligibility. A model unsuitable for a task unsupervised may be entirely suitable with review, constrained tool access, or output filtering.

The crucial property is that the evidence bar moves with the consequences. The same menu of possible evaluations applies everywhere, but how much evidence is enough depends on what happens when the answer is wrong. Users should set uniform thresholds where the highest-risk use cases require them and  where users have blocked most of the value. On the contrary, users should set them where the lowest-risk use cases require them and where users have approved things they should not have.

The organizational objection

While some might complain that this sounds slower, in my experience I have found it to be faster for a specific reason: blanket approval concentrates all the risk into one decision (which is why that decision takes months and gets escalated). Nobody wants to sign off on a model that covers every possible future use. Bounded decisions are smaller, and small decisions move quickly. The low-consequence uses that make up most of the actual demand clear quickly because the evidence bar genuinely is lower for them. The hard cases receive the additional scrutiny they deserve without holding up everything else behind them.

The primary trade off is that users need to know what their systems are actually doing. An organization that cannot describe its deployment paths cannot operate an eligibility model, and many organizations discover through this exercise that they have less visibility than they assumed. That discovery is uncomfortable and unfortunate, but it is also the point of the exercise.

Where this leads

The framing generalizes past model selection and maintains that risk primarily lives in the pairing of a system with a situation rather than in the system alone. This is a familiar idea in several other safety-critical field which is why I find it oddly controversial in the AI deployment space.

We do not ask whether a material is safe. We ask what it is rated for. The rating is not a lesser answer than a verdict. In fact, it is a more useful one because it tells users something actionable and it tells users where the boundary is. AI deployment decisions are converging on the same structure and the organizations that get there first will find they can move faster and defend their decisions better than the ones still trying to answer whether the model is trustworthy in general.