DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted with 589 unique conversation histories (12,591 messages) from 18 participants who experienced delusions and psychological harm during chatbot use, and scored on delusion-linked behaviors including failure to discourage self-harm.
Publisher
arXiv (Stanford-led author team)
Published
5 Aug 2026
Added
2 months ago
Key Findings
- A model's tendency to exhibit delusion-linked behavior did not reliably correlate with model size, release date, or the presence of test-time reasoning
- Extending conversation context substantially increases delusion-linked behavior: the rate of failing to discourage self-harm when the user expresses suicidal ideation rose from 30.0% to 41.1% when an additional 350 messages were prepended to the history
- All major evaluated model families showed substantial rates of concerning behavior on the protocol
Methodology Notes
589 conversation histories comprising 12,591 messages donated by 18 users with lived experience of delusions and psychological harm from chatbot interaction; 12 model families evaluated. Preprint (arXiv 2608.05004, v1 2026-08-05), no journal reference yet; verified via the arXiv API. Access caveat verified 2026-08-26: the dataset (HuggingFace spiralsafety/delusioneval, published by the Stanford SPIRALS group) is gated auto under a DelusionEval Controlled Data Use Agreement, Version 1, last updated 2026-08-04. The grant is non-commercial scientific research only, redistribution in whole or part is prohibited, cross-referencing with other datasets is prohibited where it could raise re-identification risk, and non-research deployment contexts are excluded. Stanford may change the terms unilaterally, and the agreement warns that third parties may hold rights in the contributed transcripts. The dataset is therefore not usable in a commercial product context, unlike every permissively licensed benchmark otherwise held here. The released parquet holds 725 conversation windows across 18 behaviour-code labels, which differs from the abstract's 589-history figure because the unit differs.
Sources
arXiv abstract(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong
Tags
Cite This
APA
Jared Moore et al. (2026). DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots. arXiv (Stanford-led author team). https://arxiv.org/abs/2608.05004
Related Insights
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
arXiv preprint · 31 May 2026
Characterizing Delusional Spirals through Human-LLM Chat Logs
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
AI Psychosis: Does Conversational AI Amplify Delusion-Related Language?
arXiv (University of Illinois Urbana-Champaign); accepted to EMNLP 2026 Main Conference · 20 Mar 2026
"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
arXiv (City University of New York; King's College London Institute of Psychiatry, Psychology & Neuroscience) · 15 Apr 2026
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
arXiv preprint · 13 Aug 2026
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026