I work on whether the evidence about AI actually supports the conclusions people want to draw from it.
There is no shortage of evidence anymore. Models are evaluated constantly. Red teams probe for failures nobody anticipated. Benchmarks appear faster than anyone can keep up with. Producing evidence is no longer the hard part. Understanding what it means, and whether it is strong enough to support a decision, is.
Much of my work sits in the space between measurement and decision-making. Other teams run evaluations, conduct experiments, and generate evidence. I spend my time looking across those different sources, understanding what each one does and does not tell us, and determining what conclusions can reasonably be drawn from them.
A result may be statistically significant and still be a weak basis for action. Another may be imperfect but point to a genuine risk. Most of the judgment lies in distinguishing between the two.
The technical evidence is only part of the story. The questions that interest me most are often the ones that emerge after the metrics have been calculated: who benefits from a system, who bears the risk when it fails, what assumptions are being embedded into institutions, and what dependencies are created when organizations build on technologies they do not control.
A finding can be technically correct and still support a bad decision. Likewise, a system can perform well against its evaluations while creating risks those evaluations were never designed to measure. Understanding the difference is a large part of what I do.
Recently, much of that work has focused on open-weight AI models, where the technical and institutional questions become difficult to separate. Evaluating capabilities, safeguards, and failure modes is important, but so is understanding what organizations are actually acquiring when they adopt someone else’s model, what they become dependent on over time, and what risks are transferred rather than eliminated.
Those questions are often debated extensively but tested surprisingly little. I am interested in bringing evidence to questions that are usually argued from intuition, ideology, or optimism.
At its core, my work is about helping people avoid becoming more certain than the evidence allows.
Writing
Openness Is a Safety Property
Everyone argues about who can download a model. The question that decides outcomes is who can fix one.
Your Model Isn't Censored the Way You Think
Three very different problems produce exactly the same silence. Telling them apart changes everything about what you can do next.
Owning the Infrastructure Doesn't Decontaminate the Model
Hosting a model on your own soil settles one real question and leaves the harder one completely untouched.
Eligibility, Not Trust
We never ask whether steel is safe. We ask what load it's rated to carry. Models deserve the same question.
Selected work
Open-weight models and ecosystem strategy
Microsoft · current
The question that sits above any single model: which ecosystems an organization should build on, on what terms, and what it accepts commercially and geopolitically when it commits to one. Most of that gets decided on instinct. It doesn’t have to be.
EarthTime
Carnegie Mellon University CREATE Lab
Using decades of satellite imagery to show what past decisions did to the present, from how redlining still shapes American cities to the global rise of China’s technology firms. I ran the same approach forward: global temperatures projected to 2100, and which coastlines go under at 0 to 4°C of warming. The stories are public at earthtime.org.
Seeing the signals through the noise
World Economic Forum
How do you spot a global problem before it looks like one? I built the data science methods the Forum used to answer that, which turned out to be a question about how much faith a faint pattern deserves.
If you're working on any of this, I'd like to hear from you. andrew@andrewberkley.com
