Skip to content

Francesco Quinzan

Organisation
The University of Oxford
Biography

Why do you care about AI Existential Safety?

I care about AI existential safety because AI systems are becoming increasingly capable and deeply integrated into society, while our ability to reliably anticipate their behaviour still lags behind their rapid progress. I believe ensuring advanced AI systems remain safe, reliable, and beneficial is both an important scientific challenge and a societal responsibility. As these systems become more widely deployed, failures or unintended behaviour could affect areas such as healthcare, scientific discovery, and public decision-making at scale. Addressing these questions early is therefore essential if we want increasingly powerful AI systems to contribute positively to society and remain aligned with human interests over the long term.

Please give at least one example of your research interests related to AI existential safety:

A major focus of my research is the problem of AI alignment: ensuring that increasingly capable AI systems behave in ways that are helpful, reliable, and consistent with human intentions. This problem is becoming increasingly important as AI systems are integrated into areas such as healthcare, scientific discovery, education, and public decision-making, where failures could have large-scale societal consequences.

Modern AI systems are often adapted using human feedback, where humans compare different model behaviours and indicate which ones they prefer. In practice, however, human feedback is rarely perfect. People can disagree with one another, provide inconsistent judgments, or give feedback that is intentionally misleading. This is particularly problematic, since learning from unreliable feedback can lead to unstable or unintended behaviour, making alignment substantially more difficult. My work studies how to make this process more reliable. In particular, I develop methods that allow AI systems to learn more robustly from imperfect human feedback, reducing the extent to which noisy or misleading evaluations distort model behaviour. Rather than assuming that human judgments are always fully consistent and reliable, my work explicitly accounts for ucertainty in human feedback.

This research is timely, since future AI systems will likely become increasingly adaptive and influential, while continuing to rely on human feedback and interaction during deployment. If we do not understand how to make these feedback-driven systems reliable and robust, there is a risk that increasingly capable models could develop behaviours that are difficult to predict, evaluate, or correct. Developing methods that make alignment more stable under realistic conditions is therefore an important step toward ensuring that advanced AI systems remain safe, trustworthy, and beneficial to society over the long term.

Sign up for the Future of Life Institute newsletter

Join 70,000+ others receiving periodic updates on our work and focus areas.
cloudmagnifiercrossarrow-up linkedin facebook pinterest youtube rss twitter instagram facebook-blank rss-blank linkedin-blank pinterest youtube twitter instagram