Skip to main content
Peer-reviewed Authoritative

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

Peer-reviewed study testing whether aggregated expert judgment yields valid ground truth for training and evaluating AI systems in mental-health safety contexts. Three certified psychiatrists independently rated LLM-generated responses to mental-health scenarios using a calibrated rubric; inter-rater reliability was consistently poor, with disagreement greatest on the most safety-critical items.

Publisher

ACM (Proceedings of FAccT 2026)

Published

25 Jun 2026

Added

2 weeks ago

Key Findings

  • Inter-rater reliability among three certified psychiatrists was consistently poor (ICC 0.087-0.295), below thresholds considered acceptable for consequential assessment
  • Expert disagreement was highest on the most safety-critical items; suicide and self-harm responses produced greater divergence than any other category
  • Challenges the assumption that averaging expert ratings yields reliable ground truth for mental-health AI safety training and evaluation

Methodology Notes

Published in the Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (DOI 10.1145/3805689.3812332, 2026-06-25); Stanford-led team including Nina Vasan and Max Lamparth; also presented at the American Psychiatric Association Annual Meeting 2026. ACM Digital Library blocks automated fetchers; verified via Crossref DOI metadata, the arXiv preprint (2601.18061), and Stanford Report coverage.

Authors

Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy, Robbie Fraser, Nina Vasan, Darja Djordjevic, Akanksha Dadlani, Max Lamparth

Tags

facctinter-rater-reliabilityexpert-ground-truthstanfordllm-judge

Cite This

APA

Kiana Jafari et al. (2026). Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing. ACM (Proceedings of FAccT 2026). https://dl.acm.org/doi/10.1145/3805689.3812332