
Muhao Chen
Why do you care about AI Existential Safety?
I care about AI existential safety because my research is fundamentally about making advanced AI trustworthy under real-world incentives and adversarial pressure. The same failure modes that show up today as jailbreaks, data leakage, or unsafe outputs can scale dramatically as models become more capable and agentic. On my homepage and in my lab’s mission, I emphasize accountability, security, robustness, and safety for large (multi-modal) language models and foundation-model agents. As these systems are increasingly used in high-stakes settings (e.g., healthcare, legal workflows, critical infrastructure), small alignment or security gaps can cascade into large, systemic harms. That’s why I focus on safety/agentic AI research: to turn “it works” into “it remains controllable, auditable, and resilient,” even as autonomy and deployment scale.
Please give at least one example of your research interests related to AI existential safety:
My defense-oriented line of work focuses on building guardrails that keep increasingly agentic systems controllable, auditable, and robust under distribution shift and adversarial pressure. In this direction, we develop practical safety layers such as ThinkGuard for deliberative, critique-augmented safety decisions, OmniGuard for unified guardrails across modalities, PolicyGuard for detecting and even anticipating policy violations in long-horizon agent trajectories, and AGrail for adaptive guardrails that can evolve as policies, environments, and failure modes change.
In parallel, I work on automated red teaming and evaluation to systematically surface dangerous behaviors before they show up in the wild and to keep safety measurement meaningful as models improve. This includes AutoDAN-style automated jailbreak generation to probe hidden vulnerabilities and ArenaBencher-style benchmark evolution to reduce leakage, strengthen comparisons, and continuously pressure-test models against newly emerging failure modes.
