Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 declined with conversational depth across all model families while clinician performance remained stable; conversational content, not length alone, explained the degradation.
Publisher
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group)
Published
14 Jul 2026
Added
2 weeks ago
Key Findings
- F1 for suicidal-ideation detection declined with conversational depth across all 49 evaluated model families; clinicians showed no comparable decline
- Larger models performed better overall but still degraded with depth; conversational content rather than raw length explained the changes
- Restating instructions partially recovered model performance, and the authors conclude safety evaluations must test sustained performance across realistic extended interactions rather than isolated prompts
Methodology Notes
Preregistered design inserting validated clinical statements of suicidal ideation into therapy transcripts at varying depths (0-200 speaker turns); 49 LLMs and 8 clinicians evaluated head-to-head. Preprint, not yet peer-reviewed (medRxiv v1, 2026-07-14). medrxiv.org blocks automated fetchers; verified via the medRxiv API (title, authors, date, abstract).
Sources
Authors
Mark Kalinich, James Luccarelli, Joseph Santa Maria, Matthew Flathers, An Nguyen, Sung Hyun Song, Karim Makhoul, Maria Jose Rivera Criado, Caroline M. Ginapp, Benjamin Hill, Jackson N. Shumate, Hana Notsu, Colin Smith, Frazer Moss, John Torous
Tags
Cite This
APA
Mark Kalinich et al. (2026). Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study. medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group). https://www.medrxiv.org/content/10.64898/2026.07.10.26357132v1
Related Insights
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
arXiv · 2 Jan 2026
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
arXiv · 11 May 2025
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
arXiv (Stanford-led author team) · 5 Aug 2026
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026