Skip to main content
Peer-reviewed Authoritative

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Simulated conversations between LLM user agents and chatbots were rated by licensed clinicians and an LLM judge on identical rubrics, with the judge aligning closely with clinical consensus.

Publisher

JMIR AI

Published

29 Jun 2026

Added

2 weeks ago

Key Findings

  • Licensed clinical raters showed strong chance-corrected inter-rater reliability (0.77) on chatbot suicide-safety ratings
  • The LLM judge aligned closely with clinician consensus (0.81), supporting VERA-MH's validity as an automated safety benchmark
  • Evaluation scores safety dimensions covering risk detection, risk confirmation, guiding to human care, supportive conversation, and following AI boundaries

Methodology Notes

Version of record (JMIR AI 2026;5:e92817, published online 2026-06-29) of the arXiv 2602.05088 preprint. LLM user-simulator plus LLM-judge design benchmarked against licensed-clinician gold-standard ratings; initial scope is suicide risk. Verified via Crossref metadata for DOI 10.2196/92817.

Authors

Kate H. Bentley, Luca Belli, Adam M. Chekroud, Emily J. Ward, Emily R. Dworkin, Emily Van Ark, Kelly M. Johnston, Will Alexander, Millard Brown, Matt Hawrilenko

Tags

jmir-aivera-mhsuicide-safetyllm-judgeversion-of-record

Cite This

APA

Kate H. Bentley et al. (2026). AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation. JMIR AI. https://ai.jmir.org/2026/1/e92817

Related Insights

Benchmark / dataset

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026

Peer-reviewed

Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment

Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025

Benchmark / dataset

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

MLCommons · 1 Mar 2025

Preprint

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026

Benchmark / dataset

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026

Peer-reviewed

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

ACM (Proceedings of FAccT 2026) · 25 Jun 2026

Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School · 1 May 2026

Peer-reviewed

An AI-based mental health guardrail and dataset for identifying psychiatric crises in text-based conversations

npj Digital Medicine · 3 Apr 2026

Preprint

Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure

PsyArXiv (Corporal Michael J. Crescenz VA Medical Center; University of Pennsylvania; Stanford; Columbia University and others) · 21 Jun 2026