MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the interactional role an AI counselor adopts (perpetrator, instigator, facilitator, or enabler). MHSafeEval is a closed-loop, agent-based evaluation framework that assesses safety across multi-turn counseling interactions, showing that static single-response benchmarks miss cumulative, role-dependent safety failures.
Publisher
Association for Computational Linguistics (Findings of ACL 2026)
Published
1 Jul 2026
Added
1 month ago
Key Findings
- Introduces a role-aware harm taxonomy (perpetrator/instigator/facilitator/enabler) for AI counselor behavior rather than classifying harms by content alone
- Agent-based multi-turn evaluation surfaces cumulative safety failures that emerge over the course of counseling conversations and are missed by static benchmarks
- Demonstrates improved diagnostic capability over static single-response evaluation methods
Methodology Notes
Published in Findings of ACL 2026 (July 2026; exact day not stated — anthology gives month/year), pages 27760-27793, DOI 10.18653/v1/2026.findings-acl.1382. Version of record of arXiv:2604.17730; the preprint was never separately held in this library. Code and data released at github.com/suhyun565/MHSafeEval. Verified directly on the ACL Anthology paper page. Re-verified 2026-08-26: the GitHub release is genuine and carries the per-condition patient configuration files as well as code. Do not confuse it with the HuggingFace repository Suhyunlee/MHSafeEval, which holds only .gitattributes and a README, declares no licence, and has not been touched since three minutes after its creation on 2026-04-21 — that repository is documentation-only and is not evidence that the data is unreleased.
Authors
Suhyun Lee, Palakorn Achananuparp, Neemesh Yadav, Ee-Peng Lim, Yang Deng
Tags
Cite This
APA
Suhyun Lee et al. (2026). MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models. Association for Computational Linguistics (Findings of ACL 2026). https://aclanthology.org/2026.findings-acl.1382/
Related Insights
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
Association for Computational Linguistics (EACL 2026) · 24 Mar 2026
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026