Perhaps the most consequential open question in AI governance is whether, and how fast, AI systems will come to improve AI itself. Recursive self-improvement, in which systems accelerate the research and development that produces their successors, is a regime in which capability growth could outpace human-speed oversight, and one that the frontier labs’ own safety frameworks designate as critical. My interest is in whether we could tell, and how soon, and with what confidence. The questions that matter here seem to me inferential ones: where is a system on its capability trajectory now, how sure are we, and for how long does that judgement hold?
I am therefore interested in treating the tracking and forecasting of AI capabilities as a problem of sequential Bayesian inference rather than of benchmarking. The systems we wish to govern are not static objects sampled independently: they learn, act in environments, and may modify the processes that build them, which undermines the statistical assumptions that make conventional evaluation scores meaningful. Nor can their behaviour be characterised exhaustively, the space being far too large, so evaluation becomes a question of experimental design: what is worth measuring next, given what we already believe and what measurement costs. My background in Bayesian optimisation, active learning and changepoint detection is what I bring to this.
A second interest is in the limits of what evaluation can establish at all. Systems may behave differently when they are aware of being tested, which raises the question of whether deployment behaviour is recoverable from test data even in principle, and under what conditions. Underlying both is a view about what safety claims should look like. A useful claim about an AI system is calibrated, auditable, and explicit about its assumptions and about when it expires, since any claim decays as the system it describes is updated. Commitments of the if-then form that now appear in frontier safety policies depend entirely on this: such a commitment is only as sound as the monitoring that decides whether its trigger condition has been met.
And because the recipients of these claims are human overseers, I am interested in monitoring whose outputs people can interrogate and safely overrule, which continues a line of work I have pursued on collaboration between human experts and inference machinery. This all descends from a longer-standing interest in probabilistic numerics, the treatment of computation itself as inference, which yields algorithms that estimate their own error rather than assuming it away. The instinct is the one I bring to evaluation: a system should know, and say, how much it does not know.

