Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues

Stress-tests three leading LLMs across up to 20-turn psychiatric dialogues using 50 virtual patient profiles, measuring how safety boundaries erode as models attempt comfort and empathy. Finds boundary violations are common and accelerate under adaptive probing.

Publisher

arXiv

Published

2 Jan 2026

Added

3 months ago

DOI

—

Key Findings

  • Safety boundaries erode over multi-turn dialogue as models prioritise comfort and empathy
  • Adaptive probing cut the average number of turns before a boundary violation from 9.21 to 4.64
  • Definitive or zero-risk reassurances were the dominant violation mode; single-turn evaluation misses these failures

Methodology Notes

Preprint (arXiv 2601.14269, v1 2 January 2026). Simulated stress-testing of three LLMs over up to 20 turns across 50 virtual patient profiles.

Authors

Youyou Cheng, Zhuangwei Kang, Kerry Jiang, Chenyu Sun, Qiyang Pan

Tags

arxivmulti-turnguardrail-decayboundary-failuremental-health

Cite This

APA

Youyou Cheng et al. (2026). The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues. arXiv. https://arxiv.org/abs/2601.14269

Related Insights

NGO report

AI Chatbots for Mental Health Support (AI Risk Assessment)

Common Sense Media · 14 Nov 2025

Peer-reviewed

Characterizing Delusional Spirals through Human-LLM Chat Logs

ACM (Proceedings of FAccT 2026) · 25 Jun 2026

Preprint

Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs

arXiv (ELLIS Alicante-led) · 29 Sept 2025

Preprint

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

arXiv (Salesforce AI Research) · 30 Jul 2026

Preprint

Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions

arXiv (Grow Therapy; Stanford University School of Medicine) · 8 Sept 2026

Preprint

Do No Harm: Exposing Hidden Vulnerabilities of LLMs via Persona-based Client Simulation Attack in Psychological Counseling

arXiv · 6 Apr 2026

Preprint

When Chatbots Accommodate: Auditing the Response Policies of AI Companions in Vulnerable Conversations

arXiv (University of Southern California: Information Sciences Institute, Viterbi School of Engineering, Annenberg School for Communication and Journalism); accepted to Findings of EMNLP 2026 · 3 Jun 2026

Peer-reviewed

AI-Facilitated Coercive Control: An Experimental Study

ACM (Proceedings of CHI 2026); Cornell / Cornell Tech · 13 Apr 2026

Preprint

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026