VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
An open-source, clinically grounded automated evaluation of chatbot safety in mental-health contexts, with an initial focus on suicide risk. It uses language-model user simulators and an LLM judge scoring five safety dimensions, validated against licensed-clinician ratings.
Publisher
arXiv (Spring Health / Slingshot AI-affiliated author team)
Published
4 Feb 2026
Added
1 month ago
Key Findings
- Individual clinicians were consistent with one another (chance-corrected inter-rater reliability = 0.77)
- The LLM judge aligned strongly with clinician consensus (IRR = 0.81), supporting validity and reliability
- Scores five safety dimensions: detecting risk, confirming risk, guiding to human care, supportive conversation, and following AI boundaries
Methodology Notes
Preprint (arXiv, v1 2026-02-04, v3 2026-02-17). User-simulator plus LLM-judge design benchmarked against clinician gold-standard ratings; initial scope is suicide risk with an open-source rubric.
Sources
arXiv abstract (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Kate H. Bentley, Luca Belli, Adam M. Chekroud, Emily J. Ward, Emily R. Dworkin, Emily Van Ark, Kelly M. Johnston, Will Alexander, Millard Brown, Matt Hawrilenko
Tags
Cite This
APA
Kate H. Bentley et al. (2026). VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health. arXiv (Spring Health / Slingshot AI-affiliated author team). https://arxiv.org/abs/2602.05088
Related Insights
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
arXiv (ELLIS Alicante-led) · 29 Sept 2025
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
MLCommons · 1 Mar 2025
Large language models for psychosocial risk assessment: A multi-method evaluation across suicide, intimate partner violence, and substance misuse
PLOS Digital Health · 27 Apr 2026
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
arXiv (Emory University-led) · 27 Oct 2025
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025
Beyond Engagement: A Multidimensional Framework to Evaluate the Safe Development of Agentic AI in Mental Health
Lecture Notes in Computer Science (Springer Nature) — AI for Clinical Applications · 22 Sept 2025
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
arXiv · 3 Mar 2026
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
Association for Computational Linguistics (EACL 2026) · 24 Mar 2026
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School · 1 May 2026
Automated Safety Testing and Reporting Application for Conversational Safety Monitoring of Generative AI Tools for Mental Health: Development and Validation Study
JMIR Mental Health · 19 May 2026
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
JMIR Mental Health · 11 Jun 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Nature Medicine · 7 Aug 2026
GPT-5.5 System Card
OpenAI · 23 Apr 2026