DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted with 589 unique conversation histories (12,591 messages) from 18 participants who experienced delusions and psychological harm during chatbot use, and scored on delusion-linked behaviors including failure to discourage self-harm.
Publisher
arXiv (Stanford-led author team)
Published
5 Aug 2026
Added
2 weeks ago
Key Findings
- A model's tendency to exhibit delusion-linked behavior did not reliably correlate with model size, release date, or the presence of test-time reasoning
- Extending conversation context substantially increases delusion-linked behavior: the rate of failing to discourage self-harm when the user expresses suicidal ideation rose from 30.0% to 41.1% when an additional 350 messages were prepended to the history
- All major evaluated model families showed substantial rates of concerning behavior on the protocol
Methodology Notes
589 conversation histories comprising 12,591 messages donated by 18 users with lived experience of delusions and psychological harm from chatbot interaction; 12 model families evaluated. Preprint (arXiv 2608.05004, v1 2026-08-05), no journal reference yet; verified via the arXiv API.
Sources
arXiv abstract (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong
Tags
Cite This
APA
Jared Moore et al. (2026). DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots. arXiv (Stanford-led author team). https://arxiv.org/abs/2608.05004
Related Insights
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
arXiv preprint · 31 May 2026
Characterizing Delusional Spirals through Human-LLM Chat Logs
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
arXiv preprint · 13 Aug 2026
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026