Skip to main content
Preprint Credible

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicide-risk detection and response in simulated conversations. Finds models detect risk far more reliably than they act on it: failures concentrate in confirming ambiguous risk statements, offering concrete crisis resources, and persistently directing users to emergency support.

Publisher

Harvard Business School

Published

1 May 2026

Added

1 week ago

Key Findings

  • The highest overall safety score across ten models from four providers was 64 out of 100; newer models improved over earlier ones for three of the four providers
  • In 60.8% of analyzed conversations, chatbots failed to ask direct questions to confirm whether an ambiguous user statement reflected suicidal thoughts or another immediate safety risk
  • About 33% of conversations lacked relevant mental-healthcare resources or a specific 24/7 crisis line, and in a further 12% models failed to persistently direct users toward emergency support even when risk was immediate
  • The authors note that even a score of 65 could still reflect up to 20% of interactions rated as having high potential for harm

Methodology Notes

Harvard Business School Working Paper No. 26-084, dated May 2026 (exact day not stated; recorded as 2026-05-01). SSRN DOI 10.2139/ssrn.6836551. The author team spans Massachusetts General Hospital, HBS, Dartmouth, Stanford, and Spring Health and overlaps the VERA-MH developer group — an application study from the framework's own orbit rather than a fully independent audit. hbs.edu PDF and ssrn.com block automated fetchers; verified via Crossref DOI metadata (title, full author list) plus the HBS AI Institute summary of 2026-07-30, per the authoritative-surrogate route.

Authors

Kate H. Bentley, Emily Van Ark, Julian De Freitas, Tim Hahn, Nicholas C. Jacobson, Nina Vasan, Ursula Whiteside, Luca Belli, Josh Gierenger, Nilu Zhao, Millard Brown, Adam M. Chekroud, Matt Hawrilenko

Tags

vera-mhworking-paperscorecardspring-healthcross-provider

Cite This

APA

Kate H. Bentley et al. (2026). Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response. Harvard Business School. https://dx.doi.org/10.2139/ssrn.6836551