Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Simulated conversations between LLM user agents and chatbots were rated by licensed clinicians and an LLM judge on identical rubrics, with the judge aligning closely with clinical consensus.

Publisher

JMIR AI

Published

29 Jun 2026

Added

2 months ago

Key Findings

  • Licensed clinical raters showed strong chance-corrected inter-rater reliability (0.77) on chatbot suicide-safety ratings
  • The LLM judge aligned closely with clinician consensus (0.81), supporting VERA-MH's validity as an automated safety benchmark
  • Evaluation scores safety dimensions covering risk detection, risk confirmation, guiding to human care, supportive conversation, and following AI boundaries

Methodology Notes

Version of record (JMIR AI 2026;5:e92817, published online 2026-06-29) of the arXiv 2602.05088 preprint. LLM user-simulator plus LLM-judge design benchmarked against licensed-clinician gold-standard ratings; initial scope is suicide risk. Verified via Crossref metadata for DOI 10.2196/92817.

Authors

Kate H. Bentley, Luca Belli, Adam M. Chekroud, Emily J. Ward, Emily R. Dworkin, Emily Van Ark, Kelly M. Johnston, Will Alexander, Millard Brown, Matt Hawrilenko

Tags

jmir-aivera-mhsuicide-safetyllm-judgeversion-of-record

Cite This

APA

Kate H. Bentley et al. (2026). AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation. JMIR AI. https://ai.jmir.org/2026/1/e92817

Related Insights

Benchmark / dataset

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026

Peer-reviewed

Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment

Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025

Benchmark / dataset

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

MLCommons · 1 Mar 2025

Preprint

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026

Benchmark / dataset

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026

Lab publication

Funding better evaluations of AI's impact on wellbeing

Anthropic · 25 Aug 2026

Preprint

aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI

arXiv (Keido Labs) · 13 Jul 2026

Benchmark / dataset

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

arXiv (MindSurf) · 29 Jun 2026

Benchmark / dataset

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026

Peer-reviewed

The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7

Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) · 1 Jul 2026

Peer-reviewed

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

ACM (Proceedings of FAccT 2026) · 25 Jun 2026

Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School · 1 May 2026

Peer-reviewed

Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation

JMIR Mental Health (JMIR Publications) · 1 Sept 2026

Preprint

mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0)

PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington) · 15 May 2026

Peer-reviewed

An AI-based mental health guardrail and dataset for identifying psychiatric crises in text-based conversations

npj Digital Medicine · 3 Apr 2026

NGO report

AI and Suicide Prevention: A Cross-Sector Primer

Partnership on AI (via arXiv) · 5 May 2026

Preprint

Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure

PsyArXiv (Corporal Michael J. Crescenz VA Medical Center; University of Pennsylvania; Stanford; Columbia University and others) · 21 Jun 2026

Preprint

Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow

Research Square (preprint); Spring Health (Spring Care Inc) · 1 Oct 2026

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Slingshot AI · 1 Oct 2026

Benchmark / dataset

VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)

Spring Health (SpringCare/VERA-MH open-source repository) · 24 Sept 2026

Benchmark / dataset

Evaluating AI Safety in Teen Conversations

Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine · 23 Sept 2026