Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Peer-reviewed study testing whether aggregated expert judgment yields valid ground truth for training and evaluating AI systems in mental-health safety contexts. Three certified psychiatrists independently rated LLM-generated responses to mental-health scenarios using a calibrated rubric; inter-rater reliability was consistently poor, with disagreement greatest on the most safety-critical items.
Publisher
ACM (Proceedings of FAccT 2026)
Published
25 Jun 2026
Added
2 months ago
Key Findings
- Inter-rater reliability among three certified psychiatrists was consistently poor (ICC 0.087-0.295), below thresholds considered acceptable for consequential assessment
- Expert disagreement was highest on the most safety-critical items; suicide and self-harm responses produced greater divergence than any other category
- Challenges the assumption that averaging expert ratings yields reliable ground truth for mental-health AI safety training and evaluation
Methodology Notes
Published in the Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (DOI 10.1145/3805689.3812332, 2026-06-25); Stanford-led team including Nina Vasan and Max Lamparth; also presented at the American Psychiatric Association Annual Meeting 2026. ACM Digital Library blocks automated fetchers; verified via Crossref DOI metadata, the arXiv preprint (2601.18061), and Stanford Report coverage.
Authors
Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy, Robbie Fraser, Nina Vasan, Darja Djordjevic, Akanksha Dadlani, Max Lamparth
Tags
Cite This
APA
Kiana Jafari et al. (2026). Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing. ACM (Proceedings of FAccT 2026). https://dl.acm.org/doi/10.1145/3805689.3812332
Related Insights
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Characterizing Delusional Spirals through Human-LLM Chat Logs
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
arXiv preprint · 8 Aug 2026
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
arXiv (preprint) · 21 Aug 2026
Beyond the Response: Examining Reasoning and Execution Fidelity in Large Language Models for Mental Health
Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26) (ACM); Vanderbilt University · 25 Jun 2026
When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Seeking Late Night Life Lines: Experiences of Conversational AI Use in Mental Health Crisis
Association for Computing Machinery (Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency, FAccT 2026) · 25 Jun 2026
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions
Journal of Marital and Family Therapy (Wiley, for the American Association for Marriage and Family Therapy) · 31 Aug 2026