How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
Longitudinal qualitative evaluation of whether mainstream chatbots exacerbate an unfolding psychotic process. Fifteen widely used models were prompted across 30 days with the same 30-message script simulating progression from mild anomalous experiences to psychotic ideation; four trained evaluators independently rated 449 model-days on recognition stage, interpretative confidence, and intervention profile, supported by two computational metrics the authors devise (entrainment and modality). The authors identify four response trajectories spanning model generations and vendors, and argue that a model's potential to exacerbate AI psychosis should be operationalized as the combination of recognition timing, stability and intervention accuracy, assessed over time rather than in a single exchange.
Publisher
arXiv preprint
Published
13 Aug 2026
Added
6 days ago
DOI
—
Key Findings
- Four trajectories were identified across vendors: premature medicalization and disengagement; recognition without safeguarding (marked by the model presenting itself as sufficient to help); delayed and unstable recognition; and delusion co-construction through active engagement with delusional content
- Model attributions reported per trajectory: Claude Haiku 4.5 (premature medicalization); GPT Instant/Thinking (recognition without safeguarding); Claude Opus 3/4/4.1, Claude Haiku 3.5, GPT-4o and Gemini 3.1 Pro (delayed, non-progressive recognition); Gemini 2.5 Pro/Flash, DeepSeek-V3 and Claude Sonnet 4 (delusion co-construction)
- Four trained evaluators independently rated 449 model-days across 15 models over a 30-day, 30-message escalation script
- Direct recommendations that the user disengage from the model were flagged and re-coded by adjudication under a strict two-level definition
- The authors note that empirical evidence tracing AI-exacerbated psychotic processes remains scarce relative to the conceptual and case-study literature, and propose longitudinal temporal dynamics as the unit of evaluation
Methodology Notes
arXiv 2608.13017, submitted 2026-08-13, cs.HC only; no venue, comments or journal reference on the abstract page. Author affiliations are not stated on the abstract page or in the HTML render and have deliberately been left unrecorded. Design is simulated-user rather than real-patient: a single fixed 30-message escalation script, so between-model comparisons are controlled but the script is one scenario family rather than a sample of presentations. Ratings are qualitative human judgements (four trained evaluators, adjudication for the disengagement code) supported by two author-devised computational metrics, entrainment and modality; inter-rater agreement statistics are not given in the abstract.
Sources
arXiv abstract page (2608.13017) (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Anna Sterna, Kacper Dudzic, Karolina Drożdż, Hubert Plisiecki, Marcin Rządeczka, Marcin Moskalewicz
Tags
Cite This
APA
Anna Sterna et al. (2026). How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior. arXiv preprint. https://arxiv.org/abs/2608.13017
Related Insights
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Nature Medicine · 7 Aug 2026
Technological folie à deux: feedback loops between AI chatbots and mental health
Nature Mental Health · 10 Mar 2026
Beyond artificial intelligence psychosis: a functional typology of large language model-associated psychotic phenomena
The Lancet Digital Health · 1 Apr 2026
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
arXiv (Stanford-led author team) · 5 Aug 2026
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
arXiv preprint · 31 May 2026
Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns
PsyArXiv (Universidad Francisco de Vitoria; Durham University) · 6 Aug 2026