Skip to main content
Peer-reviewed Authoritative

Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment

A RAND-led study in Psychiatric Services testing whether ChatGPT, Claude, and Gemini give direct responses to suicide-related queries and how those responses align with expert clinicians' risk ratings. Thirty hypothetical suicide-related queries, rated by clinicians into five self-harm risk levels, were each posed 100 times to each chatbot. The chatbots handled the extremes appropriately but failed to differentiate intermediate risk levels, with notable between-model differences.

Publisher

Psychiatric Services (American Psychiatric Association); RAND-led author team

Published

26 Aug 2025

Added

1 month ago

Key Findings

  • No chatbot gave direct responses to very-high-risk queries, while ChatGPT and Claude answered very-low-risk queries 100% of the time
  • None of the three chatbots meaningfully distinguished low, medium, and high (intermediate) risk levels from very-low-risk queries
  • Claude was more likely, and Gemini less likely, than ChatGPT to provide direct responses overall
  • ChatGPT reportedly answered lethality-of-means questions (e.g., which method has the highest completed-suicide rate), while Gemini declined even basic statistical queries

Methodology Notes

30 hypothetical suicide-related queries categorized by expert clinicians into five risk strata (very low to very high); each query submitted 100 times to each of three chatbots (ChatGPT, Claude, Gemini). Measures direct-response rates, not full conversational quality; hypothetical single-turn queries, not real user dialogues. Epub 2025-08-26; print issue Psychiatric Services 76(11):944-950, Nov 2025.

Authors

Ryan K. McBain, Jonathan H. Cantor, Li Ang Zhang, Olesya Baker, Fang Zhang, Alyssa Burnett, Aaron Kofner, Joshua Breslau, Bradley D. Stein, Ateev Mehrotra, Hao Yu

Tags

suicide-queriesrandchatgptclaudegeminirisk-stratificationclinician-alignment

Cite This

APA

Ryan K. McBain et al. (2025). Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment. Psychiatric Services (American Psychiatric Association); RAND-led author team. https://pubmed.ncbi.nlm.nih.gov/41174947/

Related Insights

Peer-reviewed

The Columbia–Suicide Severity Rating Scale: Initial Validity and Internal Consistency Findings From Three Multisite Studies With Adolescents and Adults

American Journal of Psychiatry (American Psychiatric Association) · 1 Dec 2011

Peer-reviewed

A machine learning approach to identifying suicide risk among text-based crisis counseling encounters

Frontiers in Psychiatry · 23 Mar 2023

Preprint

Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs

arXiv (ELLIS Alicante-led) · 29 Sept 2025

Preprint

Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale

arXiv · 11 May 2025

Benchmark / dataset

Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)

arXiv (Chinese research team) · 2 Jun 2025

Benchmark / dataset

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026

Peer-reviewed

Urgent considerations for suicide prevention in the safe and ethical use of artificial intelligence

Canadian Medical Association Journal · 19 Apr 2026

Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School · 1 May 2026

Peer-reviewed

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026

Peer-reviewed

Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations

JAAD International (Elsevier, for the American Academy of Dermatology) · 25 Jun 2026

Peer-reviewed

AI Chatbot Use and Disclosure for Mental Health Among US Adolescents and Young Adults

JAMA Pediatrics (American Medical Association); RAND-led author team · 1 Jun 2026

Peer-reviewed

Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models

JMIR Mental Health · 11 Jun 2026

Peer-reviewed

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

JMIR AI · 29 Jun 2026

Preprint

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026

Peer-reviewed

Large language models for psychosocial risk assessment: A multi-method evaluation across suicide, intimate partner violence, and substance misuse

PLOS Digital Health · 27 Apr 2026