Building computational tools to identify and address pathological behavioral states in conversational and autonomous AI
Building computational tools to identify and address pathological behavioral states in conversational and autonomous AI
Project Details
Updated 07/06/26 · Provided via application · VerifiedHigh-profile cases have shown that AI systems can instill false beliefs in its users, create abnormal emotional dependency, and help cause harmful interactions during mental health crises. As AI gains more capabilities and integrated into daily life, identifying unhealthy behaviors of AI early is becoming increasingly critical.
My project aims to create clinically informed tools for identifying and mitigating pathological behavioral trajectories in conversational and agentic AI systems. By working with longitudinal conversation data such as WildChat, I will build methods to identify issues associated with conversational and agentic AI including psychosis-like reasoning, false-belief reinforcement, inappropriate conversational loops, and abnormal user dependency. Current existing AI benchmarks only study individual responses; our proposed methods will seek to study behavioral dynamics over extended interactions and identify how behavioral risk evolves over time.
Working with clinicians and researchers at Yale School of Medicine, particularly including faculty and staff affiliated with the Department of Psychiatry, I will translate established psychiatric concepts into metrics for AI safety evaluation. These tools will be validated using clinicians from Yale.
This project will deliver an open-source "toolbox" that allows for constant monitoring of behavioral risk in conversational and agentic AI systems; clinically validated behavioral safety metrics for detecting problematic reasoning trajectories; and peer-reviewed publications describing and establishing a clinically grounded framework for behavioral AI safety.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiThis project helps make AI systems more safe by identifying harmful patterns of behavior that develop over long interactions (rather than focusing just on individual responses). Similarly to as physicians assessing patients for early red flags of illness, this project will develop tools to identify when an AI system begins reinforcing false beliefs, promoting abnormal dependency, or displaying increasingly unstable reasoning so these issues can be corrected before issues arise.
The project will produce open tools that AI developers can use to improve the safety of conversational and autonomous AI systems. By helping AI developers to detect and address these harmful behaviors, this work facilitates the creation of AI that is more trustworthy and aligned with human well-being.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.