Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical communication when a user discloses risk factors for suicidal thoughts and behaviours. Forty-three practising US therapists and trainees wrote prompts expressing seven meta-analysis-derived risk factors in randomised order across seven-turn conversations with OLMo-2-32b, then annotated the model's replies against a clinician-validated codebook covering acknowledgment of risk, empathy, resource referral and invitation to continue. Generalized linear mixed-effects models relate each annotation to the risk factor expressed and the conversation length. The headline result is that the model becomes progressively less willing to invite further discussion as more risk indicators accumulate.
Publisher
Proceedings of the IASEAI Conference (published by AAAI)
Published
15 Jul 2026
Added
3 days ago
DOI
—
Key Findings
- Responses were annotated as containing an invitation to continue the conversation just 14% of the time overall, ranging from about 20% for prompts expressing hopelessness (95% CI 7.8-42.1%) to 5% for prompts expressing prior non-suicidal self-injury (95% CI 1-22.7%)
- Withdrawal compounds across turns: when a prior-NSSI disclosure came seventh in a sequence, the probability of an invitation to continue fell to 0.8% (95% CI 0.1-3.5%), making the model only 16% as likely to invite continuation as on the first turn (p < 0.01)
- Acknowledgment of risk was inconsistent and inverted relative to severity — 85% for a first-turn prior-NSSI disclosure (95% CI 68-93.7%) but 27% for prior suicidal ideation (95% CI 15.8-42.7%) — though acknowledgment did rise with accumulating disclosures (odds ratio 1.62 per roughly two turns, p < 0.001)
- Empathy and a nonspecific pointer to seek help were near-universal, but encouragement to contact a specific resource appeared in only 41% of responses; prior suicide attempt, suicidal ideation and NSSI were 40, 27 and 22 times more likely than hopelessness to elicit a specific resource (p < 0.001)
- The authors frame conversational withdrawal itself as a hazard rather than a missing nicety, arguing it risks stigmatising a user who disclosed to a chatbot what they would not disclose to people around them
Methodology Notes
Human-in-the-loop black-box behavioural assessment, not a benchmark run. 829 annotated model responses collected from 43 participants (practising US-based therapists and trainees, recruited by screener) in guided real-time sessions held in July 2025; five further participants who interacted with OLMo-7b in pre-tests are excluded. Seven STB risk factors taken from a published meta-analysis; each conversation presented one risk factor per turn in randomised order over seven turns (14 messages). Participants paraphrased each supplied statement in their own words before sending, so prompts are not fixed strings. Statistical model is a generalized linear mixed-effects regression with participant-level random effects; the authors note they deliberately omit annotator-by-content interactions and that the invitation-to-continue odds ratio across risk factors is not significant at conventional levels (OR 0.229, p ≈ 0.08) even though the mean differences are large. Single system under test (OLMo-2-32b, chosen because full openness permits later mechanistic follow-up), so the generalisation to closed frontier models is argued by analogy to prior single-prompt work, not measured. Peer-reviewed conference proceedings, IASEAI'26, pages 279-292, volume 2(1). Verification route: the OJS landing page (fetched via text proxy) gives the 2026-07-15 publication date and the page range, and the full PDF was downloaded from ojs.aaai.org (HTTP 200, 404,134 bytes) and read; every figure above was extracted from its results section.
Sources
IASEAI'26 Proceedings article page (primary)
Full paper PDF (IASEAI'26 proceedings)
Author-designated supplementary materials (arXiv preprint version) (31 Oct 2025)
Archived snapshot (Wayback Machine) — preserved against link rot
Topics
Authors
Nick Judd, Alexandre Vaz, Kevin Paeth, Layla Inés Davis, Milena Esherick, Jason Brand, Inês Amaro, Tony Rousmaniere
Tags
Cite This
APA
Nick Judd et al. (2026). Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk. Proceedings of the IASEAI Conference (published by AAAI). https://ojs.aaai.org/index.php/IASEAI/article/view/43031
Related Insights
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School · 1 May 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Nature Medicine · 7 Aug 2026
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure
PsyArXiv (Corporal Michael J. Crescenz VA Medical Center; University of Pennsylvania; Stanford; Columbia University and others) · 21 Jun 2026