Peer-reviewed & preprints
Research
Academic and clinical scholarship on how conversational AI affects the people who use it — from crisis-response performance to companionship, dependency, and psychosis.
93 entries, newest first
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
Interaction with AI companions and psychological well-being
Stanford-led study of 1,131 adult Character.AI users combining survey self-report with donated chat transcripts from 244 of them, analyzed with LLM-assisted methods against the Comprehensive Inventor…
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
Framework preprint examining how conversational AI systems affect users' psychological health, identifying benefits (information access, learning support) alongside risks including emotional entangle…
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 d…
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…
Characterizing Delusional Spirals through Human-LLM Chat Logs
Peer-reviewed analysis of chat logs from 19 users reporting psychological harm from chatbot use, applying a 28-code inventory to 391,562 messages. Characterises how delusion-reinforcing interaction p…
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Peer-reviewed study testing whether aggregated expert judgment yields valid ground truth for training and evaluating AI systems in mental-health safety contexts. Three certified psychiatrists indepen…
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Proposes a longitudinal evaluation framework (Theater-Stage-Judge) that uses persona-driven user simulation with dynamic psychological-state updating to assess the cognitive-developmental risks of AI…
Artificial intelligence (AI) psychosis: mechanisms, clinical risks and safety considerations in generative AI chatbots
A commentary in BJPsych Open synthesizing emerging case reports of 'AI psychosis', in which intensive generative AI chatbot use is associated with delusional thinking. The authors propose a provision…
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…
Characterizing artificial intelligence (AI) psychosis in a large academic medical setting: evidence of the new clinical phenomenon and the vulnerability of those in early phases of psychosis
First systematic electronic-health-record characterization of "AI psychosis" in a clinical population: a chart review of psychosis patients at Vanderbilt University Medical Center whose records menti…