Skip to content

Maike Osborne

Position
Professor
Organisation
University of Oxford
Biography

Why do you care about AI Existential Safety?

I believe AI presents a real existential threat, and one that I, as an AI researcher, have a duty to address. Nor is the threat confined to a distant future: as AI systems are deployed within ever more sensitive applications, from healthcare to defence, the need for them to be safer is with us today.
What concerns me most is that our capacity to measure these systems is falling behind our capacity to build them. Capable observers disagree about what today’s models can do, where they are heading, and how they ought to be governed, and that disagreement is rarely settled by appeal to evidence, because the evidence is thin and poorly calibrated. Decisions of real consequence are being taken on the basis of point estimates from static benchmarks, with no account of how uncertain those estimates are or when they cease to hold.
I founded the Bayesian Governance Lab at Oxford to work on this. We ask how much can really be known about an AI system, and how to act well under what remains uncertain. Uncertainty is something to be measured and reasoned about rather than asserted, and the mathematics for doing so exists: Bayesian inference is the discipline of drawing calibrated conclusions from limited and expensive evidence. Turning it on AI systems themselves strikes me as among the more tractable contributions a machine learning researcher can make to existential safety.

Please give one or more examples of research interests relevant to AI existential safety:

Perhaps the most consequential open question in AI governance is whether, and how fast, AI systems will come to improve AI itself. Recursive self-improvement, in which systems accelerate the research and development that produces their successors, is a regime in which capability growth could outpace human-speed oversight, and one that the frontier labs’ own safety frameworks designate as critical. My interest is in whether we could tell, and how soon, and with what confidence. The questions that matter here seem to me inferential ones: where is a system on its capability trajectory now, how sure are we, and for how long does that judgement hold?
I am therefore interested in treating the tracking and forecasting of AI capabilities as a problem of sequential Bayesian inference rather than of benchmarking. The systems we wish to govern are not static objects sampled independently: they learn, act in environments, and may modify the processes that build them, which undermines the statistical assumptions that make conventional evaluation scores meaningful. Nor can their behaviour be characterised exhaustively, the space being far too large, so evaluation becomes a question of experimental design: what is worth measuring next, given what we already believe and what measurement costs. My background in Bayesian optimisation, active learning and changepoint detection is what I bring to this.
A second interest is in the limits of what evaluation can establish at all. Systems may behave differently when they are aware of being tested, which raises the question of whether deployment behaviour is recoverable from test data even in principle, and under what conditions. Underlying both is a view about what safety claims should look like. A useful claim about an AI system is calibrated, auditable, and explicit about its assumptions and about when it expires, since any claim decays as the system it describes is updated. Commitments of the if-then form that now appear in frontier safety policies depend entirely on this: such a commitment is only as sound as the monitoring that decides whether its trigger condition has been met.
And because the recipients of these claims are human overseers, I am interested in monitoring whose outputs people can interrogate and safely overrule, which continues a line of work I have pursued on collaboration between human experts and inference machinery. This all descends from a longer-standing interest in probabilistic numerics, the treatment of computation itself as inference, which yields algorithms that estimate their own error rather than assuming it away. The instinct is the one I bring to evaluation: a system should know, and say, how much it does not know.

Sign up for the Future of Life Institute newsletter

Join 70,000+ others receiving periodic updates on our work and focus areas.
cloudmagnifiercrossarrow-up linkedin facebook pinterest youtube rss twitter instagram facebook-blank rss-blank linkedin-blank pinterest youtube twitter instagram