Responses of AI chatbots to escalating suicide risk: A simulation study of repeated interactions
Simulation study in which the authors held daily conversations with ChatGPT, DeepSeek and Replika over seven days across three suicidal-risk scenarios whose severity escalated over time, with scenario design informed by the Columbia Suicide Severity Rating Scale. Outcomes were whether the chatbot explored suicidal ideation, assessed symptoms, referred the user to human support, and whether attempts to bypass its safeguards succeeded. The three systems differed sharply in referral behaviour, and most jailbreak attempts succeeded on all three.
Publisher
Journal of Affective Disorders (Elsevier); Wroclaw Medical University
Published
5 Oct 2026
Added
today
Key Findings
- Referral to human support occurred in 54 of 63 daily records (85.7%) for ChatGPT, 48 of 63 (76.2%) for DeepSeek and 6 of 63 (9.5%) for Replika.
- During the high-risk period, exploration of suicidal ideation was higher for ChatGPT than for DeepSeek (adjusted difference 25.9 percentage points, 95% CI 2.4 to 49.5, q = 0.037) and than for Replika (66.7 percentage points, 95% CI 44.2 to 89.1, q < 0.001).
- On day 7, referral occurred in 9 of 9 ChatGPT trajectories, 4 of 9 DeepSeek trajectories and 1 of 9 Replika trajectories.
- Jailbreak attempts succeeded in 6 of 9 attempts on ChatGPT, 7 of 9 on DeepSeek and 8 of 9 on Replika.
- The authors conclude that the chatbots showed marked variability and safety vulnerabilities in simulated suicidal crises, particularly under jailbreaking, and that unsupervised chatbot use during suicidal crises raises concern.
Methodology Notes
Simulated repeated interactions: three suicidal-risk scenarios with escalation over seven days, run on ChatGPT, DeepSeek and Replika, giving 63 daily records per chatbot and 27 trajectories in total; scenario escalation informed by the C-SSRS. Outcomes were coded for exploration of suicidal ideation, symptom assessment, referral to human support and vulnerability to jailbreak attempts; adjusted differences reported with 95% confidence intervals and q-values. The authors state that the simulation design and the small sample of 27 trajectories limit generalisability. Published online 2026-10-05 (PubMed article date; Crossref record created 2026-10-05); article number 122586, volume and issue not yet assigned. Verification route: the ScienceDirect page is behind a bot wall, so the record was confirmed from the PubMed entry (PMID 42833407) and the OpenAlex record, both of which carry the full structured abstract; model versions and the language of the simulated conversations are not stated in the abstract.
Sources
Topics
Authors
Wojciech Pichowicz, Marek Kotas, Natalia Kalka, Julia Dembowska, Julia Karska, Patryk Piotrowski, Błażej Misiak
Tags
Cite This
APA
Wojciech Pichowicz et al. (2026). Responses of AI chatbots to escalating suicide risk: A simulation study of repeated interactions. Journal of Affective Disorders (Elsevier); Wroclaw Medical University. https://doi.org/10.1016/j.jad.2026.122586
Related Insights
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
arXiv preprint · 13 Aug 2026
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue
medRxiv (preprint); University Hospital Frankfurt; UKP Lab, Technical University of Darmstadt · 29 Sept 2026