Skip to main content
Benchmark / dataset Credible

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

An open-source, clinically grounded automated evaluation of chatbot safety in mental-health contexts, with an initial focus on suicide risk. It uses language-model user simulators and an LLM judge scoring five safety dimensions, validated against licensed-clinician ratings.

Publisher

arXiv (Spring Health / Slingshot AI-affiliated author team)

Published

4 Feb 2026

Added

1 month ago

Key Findings

  • Individual clinicians were consistent with one another (chance-corrected inter-rater reliability = 0.77)
  • The LLM judge aligned strongly with clinician consensus (IRR = 0.81), supporting validity and reliability
  • Scores five safety dimensions: detecting risk, confirming risk, guiding to human care, supportive conversation, and following AI boundaries

Methodology Notes

Preprint (arXiv, v1 2026-02-04, v3 2026-02-17). User-simulator plus LLM-judge design benchmarked against clinician gold-standard ratings; initial scope is suicide risk with an open-source rubric.

Sources

arXiv abstract (primary)

Archived snapshot (Wayback Machine) — preserved against link rot

Authors

Kate H. Bentley, Luca Belli, Adam M. Chekroud, Emily J. Ward, Emily R. Dworkin, Emily Van Ark, Kelly M. Johnston, Will Alexander, Millard Brown, Matt Hawrilenko

Tags

arxivvera-mhsuicide-safetyllm-judgebenchmark

Cite This

APA

Kate H. Bentley et al. (2026). VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health. arXiv (Spring Health / Slingshot AI-affiliated author team). https://arxiv.org/abs/2602.05088

Related Insights

Peer-reviewed

Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment

Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025

Preprint

Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs

arXiv (ELLIS Alicante-led) · 29 Sept 2025

Benchmark / dataset

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

MLCommons · 1 Mar 2025

Peer-reviewed

Large language models for psychosocial risk assessment: A multi-method evaluation across suicide, intimate partner violence, and substance misuse

PLOS Digital Health · 27 Apr 2026

Benchmark / dataset

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

arXiv (Emory University-led) · 27 Oct 2025

Benchmark / dataset

Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)

arXiv (Chinese research team) · 2 Jun 2025

Peer-reviewed

Beyond Engagement: A Multidimensional Framework to Evaluate the Safe Development of Agentic AI in Mental Health

Lecture Notes in Computer Science (Springer Nature) — AI for Clinical Applications · 22 Sept 2025

Preprint

TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health

arXiv · 3 Mar 2026

Peer-reviewed

When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

Association for Computational Linguistics (EACL 2026) · 24 Mar 2026

Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School · 1 May 2026

Peer-reviewed

Automated Safety Testing and Reporting Application for Conversational Safety Monitoring of Generative AI Tools for Mental Health: Development and Validation Study

JMIR Mental Health · 19 May 2026

Peer-reviewed

Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models

JMIR Mental Health · 11 Jun 2026

Peer-reviewed

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

JMIR AI · 29 Jun 2026

Preprint

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026

Peer-reviewed

A clinically validated framework for auditing AI chatbot behavior in mental health interactions

Nature Medicine · 7 Aug 2026

Lab publication

GPT-5.5 System Card

OpenAI · 23 Apr 2026