Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors

Scenario-driven study of how three LLMs (ChatGPT 5.4 Thinking, Claude Sonnet 4.6 Extended Thinking, Gemini 3 Thinking) identify and resolve therapeutic-alliance ruptures across 21 mental-health conversations, compared with 22 mental-health experts who also rated the LLM strategies. LLMs relied on explicit single-turn cues and produced directive, scripted repairs; experts integrated implicit, relational and contextual information and used validation, open exploration and psychoeducation.

Publisher

arXiv (University of Massachusetts Amherst; University of Illinois Urbana-Champaign; Indiana University Indianapolis)

Published

21 Sept 2026

Added

today

DOI

Key Findings

  • LLMs showed high agreement with predefined rupture labels in identification (Gemini 100%, Claude 95.2%, ChatGPT 90.5% of 21 scenarios) but experts rated their resolution strategies only moderately effective
  • LLM identification relied on explicit linguistic cues within single turns; experts integrated implicit and contextual signals across the conversation
  • LLM repairs were more directive and scripted; experts favoured validation, open-ended exploration and psychoeducation
  • Consistent LLM limitations in timing, depth and contextual sensitivity
  • Expert panel: 22 practitioners (5 psychologists, 4 social workers, 3 licensed counsellors, 5 clinical PhD students, 2 PsyD holders)

Methodology Notes

Theory-driven pre-constructed scenarios (21 conversations) rather than live interactions; text-only; three LLMs queried in 2026; 22 US-based mental-health experts evaluated identification and resolution. Authors note the small single-country expert sample and the absence of non-textual cues. v1 posted 2026-09-21.

Authors

Jeongah Lee, Joy Qiuyue Zhong, Drishti Goel, Violeta J. Rodriguez, Dong Whi Yoo, Koustuv Saha, Ravi Karkar

Tags

therapeutic-alliancerupture-repairexpert-evaluationmental-health-chatbots

Cite This

APA

Jeongah Lee et al. (2026). Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors. arXiv (University of Massachusetts Amherst; University of Illinois Urbana-Champaign; Indiana University Indianapolis). https://arxiv.org/abs/2609.25287