
Su Hyeong Lee
Why do you care about AI Existential Safety?
AI systems are becoming increasingly capable, and I think there is a real possibility that future systems could acquire strategically important abilities before we understand them well enough to evaluate or control them. My motivation for AI existential safety comes from the gap between the scale of possible impact and the maturity of our scientific understanding. Modern neural networks can display useful, surprising, and sometimes opaque behavior; as capabilities increase, opacity, goal misgeneralization, deception, and emergent agency become increasingly serious concerns.
My background is in mathematics and machine learning, so I am drawn to safety work that turns these concerns into tractable technical questions. I care about developing tools that let us see what models have learned, whether learned structure generalizes safely, and when models may contain latent agentic substructures that are not captured by surface-level evaluations. I see existential safety as a scientific and moral responsibility: if advanced AI could shape the long-run future, then understanding and reducing its failure modes should be a central research priority.
Please give at least one example of your research interests related to AI existential safety:
One of my central research interests is developing mathematical and empirical tools for identifying latent agentic structure in neural networks. In particular, my recent work on probabilistic modeling of latent agentic substructures asks whether a network may contain internal mechanisms that behave like goal-directed subagents, even when this is not obvious from the model’s input-output behavior. I see this as relevant to inner alignment and mesa-optimization: if powerful models can learn internal optimization processes, then safety evaluations need ways to detect, characterize, and monitor those processes.
A second related interest is understanding what neural representations encode, and when our measurement tools can mislead us. My work on neural probing studies what can and cannot be inferred from simple probes of deep representations. This matters for AI safety because interpretability methods often rely on probes, classifiers, or other readouts; we need principled accounts of when such methods reveal genuine internal structure and when they merely impose an external explanation.
I am also interested in neural averaging and model comparison as ways to understand the geometry of learned solutions. If different training runs, architectures, or fine-tuning processes implement related capabilities through different representations, then studying their common structure could help us identify robust features, capability-relevant circuits, and safety-relevant changes introduced by training.
More broadly, I want to contribute to an eventual science of AI evaluation and interpretability that connects mathematical foundations with empirical work on frontier systems. The goal is to build tools that help us answer questions such as: What has a model learned internally? Are there hidden optimization processes or goal-like structures? Do safety interventions change the relevant internal mechanisms, or only surface behavior? And how can we evaluate these questions before systems become too capable to safely study by trial and error?
