Responsible Evaluation of AI for Mental Health
Position-plus-analysis paper from a consortium of NLP and clinical-psychology researchers. Two annotators coded 135 mental-health papers published in ACL Anthology venues over five years (36% from 2025) for their evaluation practice, then proposed a taxonomy distinguishing three types of AI mental-health support (assessment, intervention, information synthesis), each with distinct risks and evaluation requirements, illustrated with five case studies and closing recommendations grounded in psychometrics, clinical science and implementation science.
Publisher
Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers)
Published
1 Jul 2026
Added
today
Key Findings
- Of 135 *CL mental-health papers analysed, 50% relied only on AI/NLP metrics and 52% reported no human evaluation.
- 29% of papers with human evaluation involved no domain experts; 17% did not share their evaluation guidelines; 36% did not discuss limitations of their evaluation.
- The paper separates assessment, intervention and information-synthesis systems and argues each requires different validity evidence and different human-in-the-loop safeguards.
- Annotation covered 152 papers before manual inspection; 50% of the data was double-annotated with Cohen's kappa 0.67, disagreements resolved by a senior annotator.
Methodology Notes
Normative framework plus a literature audit limited to ACL Anthology venues (query date November 2025). Affiliations include TU Darmstadt, Vanderbilt, Trier, Bar-Ilan, Trinity College Dublin, University of Washington, Georgia Tech, Marburg, Leiden, Bocconi and Queen Mary University of London / Alan Turing Institute. Published July 2026 (proceedings dated 2 to 7 July; day not stated, so month precision), DOI 10.18653/v1/2026.acl-long.347, pages 7625 to 7660; 36-page PDF verified from the Anthology with the .bib record. Materials at github.com/ukplab/nlp-mh-evals and the project page ukplab.github.io/nlp-mh-evals. Earlier version arXiv 2602.00065.
Authors
Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych
Tags
Cite This
APA
Hiba Arnaout et al. (2026). Responsible Evaluation of AI for Mental Health. Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers). https://aclanthology.org/2026.acl-long.347/