Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicide-risk detection and response in simulated conversations. Finds models detect risk far more reliably than they act on it: failures concentrate in confirming ambiguous risk statements, offering concrete crisis resources, and persistently directing users to emergency support.
Key Findings
- The highest overall safety score across ten models from four providers was 64 out of 100; newer models improved over earlier ones for three of the four providers
- In 60.8% of analyzed conversations, chatbots failed to ask direct questions to confirm whether an ambiguous user statement reflected suicidal thoughts or another immediate safety risk
- About 33% of conversations lacked relevant mental-healthcare resources or a specific 24/7 crisis line, and in a further 12% models failed to persistently direct users toward emergency support even when risk was immediate
- The authors note that even a score of 65 could still reflect up to 20% of interactions rated as having high potential for harm
Methodology Notes
Harvard Business School Working Paper No. 26-084, dated May 2026 (exact day not stated; recorded as 2026-05-01). SSRN DOI 10.2139/ssrn.6836551. The author team spans Massachusetts General Hospital, HBS, Dartmouth, Stanford, and Spring Health and overlaps the VERA-MH developer group — an application study from the framework's own orbit rather than a fully independent audit. hbs.edu PDF and ssrn.com block automated fetchers; verified via Crossref DOI metadata (title, full author list) plus the HBS AI Institute summary of 2026-07-30, per the authoritative-surrogate route.
Sources
SSRN working paper (DOI) (primary)
HBS AI Institute summary (verification surrogate) (30 Jul 2026)
Authors
Kate H. Bentley, Emily Van Ark, Julian De Freitas, Tim Hahn, Nicholas C. Jacobson, Nina Vasan, Ursula Whiteside, Luca Belli, Josh Gierenger, Nilu Zhao, Millard Brown, Adam M. Chekroud, Matt Hawrilenko
Tags
Cite This
APA
Kate H. Bentley et al. (2026). Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response. Harvard Business School. https://dx.doi.org/10.2139/ssrn.6836551
Related Insights
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
IEEE (2025 IEEE International Conference on Future Machine Learning and Data Science, FMLDS) · 2 Nov 2025
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026