
Zekai Wang
Why do you care about AI Existential Safety?
I care about AI existential safety because frontier systems are becoming more capable and more agentic, yet we still lack reliable ways to ensure they stay aligned under distribution shift or strategic pressure. In my own work on robustness, fairness, and safety for generative models, I’ve seen how small objective mismatches can create large harms; with highly autonomous models that gap could scale to societal or even catastrophic levels. I’m motivated by a simple conviction: we shouldn’t move the limits of intelligence faster than our ability to align it with human values. Doing the technical work now, such as transparent alignment, uncertainty calibration, and robust cooperation, is one of the most responsible paths to keeping powerful AI beneficial and controllable.
Please give at least one example of your research interests related to AI existential safety:
My research related to AI existential safety grows out of a long-term focus on making powerful generative and representation models reliable in the kinds of worst-case situations that matter for real deployment. So far, my work has centered on three pillars: certified robustness, empirical robustness, and fairness.
Certified Robustness: I have worked on building verification tools that can provide provable guarantees about a model’s behavior under bounded perturbations. The motivation is that when models become more capable and more widely used, safety cannot rely only on average-case performance or on a single attack method. Certified robustness offers a way to reason about worst-case failures, to know what a system will not do, and to make safety claims that remain valid under distribution shift. This kind of guarantee is a basic ingredient for existential safety, because it helps prevent brittle behavior from scaling into large-impact failures.
Empirical Robustness: Alongside guarantees, I study robustness in practice under strong, adaptive adversaries. I have investigated how adversarial training and related defenses behave when moving from training to real-world testing, and how to close gaps that appear only at deployment time. The reason this matters for existential safety is that high-stakes systems fail in realistic settings where attackers adapt and environments change. Empirical robustness work helps ensure that safety mechanisms actually hold up under real pressures, not just in controlled evaluations.
Fairness: I also study how robustness interacts with fairness, because a system that is “robust on average” can still fail systematically for specific groups or classes. In my past work, I examined these robustness–fairness tensions and developed approaches to reduce disparities while preserving robustness. I see this as part of existential safety in a broader sense: preventing scalable harms requires that safety improvements do not concentrate risk onto particular populations or create hidden brittle pockets in deployment.
Overall, these three directions reflect a single aim: to make safety properties measurable, stress-tested, and, when possible, provable, so that scaling model capability does not outpace our ability to keep systems reliable and aligned with human interests.
