Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
Scenario-driven study of how three LLMs (ChatGPT 5.4 Thinking, Claude Sonnet 4.6 Extended Thinking, Gemini 3 Thinking) identify and resolve therapeutic-alliance ruptures across 21 mental-health conversations, compared with 22 mental-health experts who also rated the LLM strategies. LLMs relied on explicit single-turn cues and produced directive, scripted repairs; experts integrated implicit, relational and contextual information and used validation, open exploration and psychoeducation.
Publisher
arXiv (University of Massachusetts Amherst; University of Illinois Urbana-Champaign; Indiana University Indianapolis)
Published
21 Sept 2026
Added
today
DOI
—
Key Findings
- LLMs showed high agreement with predefined rupture labels in identification (Gemini 100%, Claude 95.2%, ChatGPT 90.5% of 21 scenarios) but experts rated their resolution strategies only moderately effective
- LLM identification relied on explicit linguistic cues within single turns; experts integrated implicit and contextual signals across the conversation
- LLM repairs were more directive and scripted; experts favoured validation, open-ended exploration and psychoeducation
- Consistent LLM limitations in timing, depth and contextual sensitivity
- Expert panel: 22 practitioners (5 psychologists, 4 social workers, 3 licensed counsellors, 5 clinical PhD students, 2 PsyD holders)
Methodology Notes
Theory-driven pre-constructed scenarios (21 conversations) rather than live interactions; text-only; three LLMs queried in 2026; 22 US-based mental-health experts evaluated identification and resolution. Authors note the small single-country expert sample and the absence of non-textual cues. v1 posted 2026-09-21.
Sources
arXiv abstract (v1)(opens in a new tab) (primary)
Authors
Jeongah Lee, Joy Qiuyue Zhong, Drishti Goel, Violeta J. Rodriguez, Dong Whi Yoo, Koustuv Saha, Ravi Karkar
Tags
Cite This
APA
Jeongah Lee et al. (2026). Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors. arXiv (University of Massachusetts Amherst; University of Illinois Urbana-Champaign; Indiana University Indianapolis). https://arxiv.org/abs/2609.25287
Related Insights
Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
ACM (Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency) · 23 Jun 2025
Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions
Journal of Marital and Family Therapy (Wiley, for the American Association for Marriage and Family Therapy) · 31 Aug 2026