Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
A RAND-led study in Psychiatric Services testing whether ChatGPT, Claude, and Gemini give direct responses to suicide-related queries and how those responses align with expert clinicians' risk ratings. Thirty hypothetical suicide-related queries, rated by clinicians into five self-harm risk levels, were each posed 100 times to each chatbot. The chatbots handled the extremes appropriately but failed to differentiate intermediate risk levels, with notable between-model differences.
Publisher
Psychiatric Services (American Psychiatric Association); RAND-led author team
Published
26 Aug 2025
Added
1 month ago
Key Findings
- No chatbot gave direct responses to very-high-risk queries, while ChatGPT and Claude answered very-low-risk queries 100% of the time
- None of the three chatbots meaningfully distinguished low, medium, and high (intermediate) risk levels from very-low-risk queries
- Claude was more likely, and Gemini less likely, than ChatGPT to provide direct responses overall
- ChatGPT reportedly answered lethality-of-means questions (e.g., which method has the highest completed-suicide rate), while Gemini declined even basic statistical queries
Methodology Notes
30 hypothetical suicide-related queries categorized by expert clinicians into five risk strata (very low to very high); each query submitted 100 times to each of three chatbots (ChatGPT, Claude, Gemini). Measures direct-response rates, not full conversational quality; hypothetical single-turn queries, not real user dialogues. Epub 2025-08-26; print issue Psychiatric Services 76(11):944-950, Nov 2025.
Sources
PubMed record (PMID 41174947) (primary)
Psychiatric Services article page (publisher; blocks unauthenticated fetch) (26 Aug 2025)
RAND press release (26 Aug 2025)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Ryan K. McBain, Jonathan H. Cantor, Li Ang Zhang, Olesya Baker, Fang Zhang, Alyssa Burnett, Aaron Kofner, Joshua Breslau, Bradley D. Stein, Ateev Mehrotra, Hao Yu
Tags
Cite This
APA
Ryan K. McBain et al. (2025). Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment. Psychiatric Services (American Psychiatric Association); RAND-led author team. https://pubmed.ncbi.nlm.nih.gov/41174947/
Related Insights
The Columbia–Suicide Severity Rating Scale: Initial Validity and Internal Consistency Findings From Three Multisite Studies With Adolescents and Adults
American Journal of Psychiatry (American Psychiatric Association) · 1 Dec 2011
A machine learning approach to identifying suicide risk among text-based crisis counseling encounters
Frontiers in Psychiatry · 23 Mar 2023
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
arXiv (ELLIS Alicante-led) · 29 Sept 2025
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
arXiv · 11 May 2025
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
Urgent considerations for suicide prevention in the safe and ethical use of artificial intelligence
Canadian Medical Association Journal · 19 Apr 2026
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School · 1 May 2026
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations
JAAD International (Elsevier, for the American Academy of Dermatology) · 25 Jun 2026
AI Chatbot Use and Disclosure for Mental Health Among US Adolescents and Young Adults
JAMA Pediatrics (American Medical Association); RAND-led author team · 1 Jun 2026
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
JMIR Mental Health · 11 Jun 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
Large language models for psychosocial risk assessment: A multi-method evaluation across suicide, intimate partner violence, and substance misuse
PLOS Digital Health · 27 Apr 2026