Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Peer-reviewed study testing whether aggregated expert judgment yields valid ground truth for training and evaluating AI systems in mental-health safety contexts. Three certified psychiatrists independently rated LLM-generated responses to mental-health scenarios using a calibrated rubric; inter-rater reliability was consistently poor, with disagreement greatest on the most safety-critical items.
Publisher
ACM (Proceedings of FAccT 2026)
Published
25 Jun 2026
Added
2 weeks ago
Key Findings
- Inter-rater reliability among three certified psychiatrists was consistently poor (ICC 0.087-0.295), below thresholds considered acceptable for consequential assessment
- Expert disagreement was highest on the most safety-critical items; suicide and self-harm responses produced greater divergence than any other category
- Challenges the assumption that averaging expert ratings yields reliable ground truth for mental-health AI safety training and evaluation
Methodology Notes
Published in the Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (DOI 10.1145/3805689.3812332, 2026-06-25); Stanford-led team including Nina Vasan and Max Lamparth; also presented at the American Psychiatric Association Annual Meeting 2026. ACM Digital Library blocks automated fetchers; verified via Crossref DOI metadata, the arXiv preprint (2601.18061), and Stanford Report coverage.
Sources
Authors
Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy, Robbie Fraser, Nina Vasan, Darja Djordjevic, Akanksha Dadlani, Max Lamparth
Tags
Cite This
APA
Kiana Jafari et al. (2026). Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing. ACM (Proceedings of FAccT 2026). https://dl.acm.org/doi/10.1145/3805689.3812332
Related Insights
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Characterizing Delusional Spirals through Human-LLM Chat Logs
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) · 14 Jul 2026
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
arXiv preprint · 8 Aug 2026
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026