Skip to content

Niloofar Mireshghallah

Position
Assistant Professor
Organisation
Carnegie Mellon University
Biography

Why do you care about AI Existential Safety?

I care deeply about human autonomy, growth, and freedom. I am from Iran, a country where government surveillance is a daily reality and AI tools increasingly enable the automation of oppression. Having witnessed this firsthand, the prospect of more capable AI deployed for surveillance and control is personal to me, not abstract.

Through my research, I have seen evidence of how models inadvertently pick up unintended goals, twist information across contexts, and leak sensitive data in ways that violate expectations. These failures are measurable and worsen as models grow more capable.

I am also focused on AI and mental health, specifically AI-induced psychosis and “psychofancy,” where models convince vulnerable people to adopt harmful beliefs or take dangerous actions. The sycophantic tendencies of current systems, combined with 24/7 availability and persuasive fluency, create real risks to cognitive autonomy. Surveillance, manipulation, and erosion of human agency. These concerns drive my work.

Please give at least one example of your research interests related to AI existential safety:

My research addresses AI existential safety through two threads: (1) privacy, contextual integrity, and information flow in AI systems, and (2) AI’s impact on mental health and cognitive autonomy.

On privacy and information flow: My work has shown that current LLMs fundamentally lack the ability to reason about when and with whom to share information. In ConfAIde (ICLR 2024 Spotlight), I proposed the first benchmark grounded in contextual integrity theory to test whether LLMs respect privacy norms, and found that even GPT-4 reveals private information in contexts that humans would not, 39% of the time. More recently, CIMemories (ICLR 2026) extends this to persistent memory systems, demonstrating that privacy violations compound dramatically over long-horizon interactions: as usage scales from 1 to 40 tasks, violations jump from 0.1% to 9.6%, reaching 25.1% with repeated sampling. These results reveal that models do not know how to manage and compartmentalize information: a fundamental limitation with existential implications as AI agents gain access to more personal data, tools, and long-term memory.

On AI and mental health: I am actively working on understanding how AI systems can harm users psychologically, what I refer to as “psychofancy”, where AI systems, through sycophancy, persuasive fluency, and persistent availability, can induce or amplify psychotic beliefs, manipulate vulnerable users, and erode cognitive autonomy. I have an OpenAI research grant focused on this area and am collaborating with them to study the mechanisms by which chatbot interactions lead to harmful mental health outcomes, including delusion reinforcement, dangerous behavioral suggestions, and loss of reality testing. This is a direct existential safety concern: as hundreds of millions of people interact daily with systems that can subtly reshape beliefs and decision-making, the aggregate effect on human autonomy and societal stability is profound.

These threads converge: surveillance AI erodes external freedom; psychologically manipulative AI erodes internal freedom. Both represent pathways through which advanced AI can undermine the conditions necessary for human flourishing and self-governance , which I view as core to existential safety.

Sign up for the Future of Life Institute newsletter

Join 70,000+ others receiving periodic updates on our work and focus areas.
cloudmagnifiercrossarrow-up linkedin facebook pinterest youtube rss twitter instagram facebook-blank rss-blank linkedin-blank pinterest youtube twitter instagram