
Gabriele Sarti
Why do you care about AI Existential Safety?
The rapid deployment of AI agents capable of autonomously planning and executing long-horizon tasks, together with their access to a multitude of tools with real impact on the digital world, has made safety work an urgent necessity. Recent research demonstrating that verbalized reasoning in language models is often detached from their actual thought processes revealed a critical gap: we lack sufficient tooling to identify pivotal decision-making steps across multi-step agentic workflows and trace them back to relevant factors in the user-provided queries, or the models’ training data. Problematic behaviors such as collusion or deception may emerge only in long-horizon tasks, making them difficult to detect with current methods. I believe AI interpretability tools will play a fundamental role in ensuring that these concerning behaviors do not escalate with future deployments of frontier AI systems.
Please give at least one example of your research interests related to AI existential safety:
My primary research interest is scaling interpretability techniques to agentic workflows. Current popular mechanistic interpretability approaches, such as circuit tracing, probing classifiers, and concept extraction, were designed for single-turn interactions and toy tasks but fail to scale to multi-step agentic setups, where the size of the input context and the generated interactions can render the identification of salient elements difficult. I’m developing auditing mechanisms that combine lightweight feature attribution techniques (which scale efficiently with context size) with counterfactual testing to disentangle the downstream influence of inner knowledge, retrieved context and external memory on agent execution. My work aims to enhance our ability to monitor agent behavior and detect signals of collusion or deception in real-world deployments.
