Skip to main content
Preprint Credible

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted with 589 unique conversation histories (12,591 messages) from 18 participants who experienced delusions and psychological harm during chatbot use, and scored on delusion-linked behaviors including failure to discourage self-harm.

Publisher

arXiv (Stanford-led author team)

Published

5 Aug 2026

Added

6 days ago

Key Findings

  • A model's tendency to exhibit delusion-linked behavior did not reliably correlate with model size, release date, or the presence of test-time reasoning
  • Extending conversation context substantially increases delusion-linked behavior: the rate of failing to discourage self-harm when the user expresses suicidal ideation rose from 30.0% to 41.1% when an additional 350 messages were prepended to the history
  • All major evaluated model families showed substantial rates of concerning behavior on the protocol

Methodology Notes

589 conversation histories comprising 12,591 messages donated by 18 users with lived experience of delusions and psychological harm from chatbot interaction; 12 model families evaluated. Preprint (arXiv 2608.05004, v1 2026-08-05), no journal reference yet; verified via the arXiv API.

Sources

arXiv abstract (primary)

Archived snapshot (Wayback Machine) — preserved against link rot

Authors

Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong

Tags

arxivdelusionbenchmarklong-contextlived-experience

Cite This

APA

Jared Moore et al. (2026). DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots. arXiv (Stanford-led author team). https://arxiv.org/abs/2608.05004