Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Responsible Evaluation of AI for Mental Health

Position-plus-analysis paper from a consortium of NLP and clinical-psychology researchers. Two annotators coded 135 mental-health papers published in ACL Anthology venues over five years (36% from 2025) for their evaluation practice, then proposed a taxonomy distinguishing three types of AI mental-health support (assessment, intervention, information synthesis), each with distinct risks and evaluation requirements, illustrated with five case studies and closing recommendations grounded in psychometrics, clinical science and implementation science.

Publisher

Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers)

Published

1 Jul 2026

Added

today

Key Findings

  • Of 135 *CL mental-health papers analysed, 50% relied only on AI/NLP metrics and 52% reported no human evaluation.
  • 29% of papers with human evaluation involved no domain experts; 17% did not share their evaluation guidelines; 36% did not discuss limitations of their evaluation.
  • The paper separates assessment, intervention and information-synthesis systems and argues each requires different validity evidence and different human-in-the-loop safeguards.
  • Annotation covered 152 papers before manual inspection; 50% of the data was double-annotated with Cohen's kappa 0.67, disagreements resolved by a senior annotator.

Methodology Notes

Normative framework plus a literature audit limited to ACL Anthology venues (query date November 2025). Affiliations include TU Darmstadt, Vanderbilt, Trier, Bar-Ilan, Trinity College Dublin, University of Washington, Georgia Tech, Marburg, Leiden, Bocconi and Queen Mary University of London / Alan Turing Institute. Published July 2026 (proceedings dated 2 to 7 July; day not stated, so month precision), DOI 10.18653/v1/2026.acl-long.347, pages 7625 to 7660; 36-page PDF verified from the Anthology with the .bib record. Materials at github.com/ukplab/nlp-mh-evals and the project page ukplab.github.io/nlp-mh-evals. Earlier version arXiv 2602.00065.

Authors

Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych

Tags

acl-2026evaluation-practicemental-health-nlpposition-paperliterature-audit

Cite This

APA

Hiba Arnaout et al. (2026). Responsible Evaluation of AI for Mental Health. Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers). https://aclanthology.org/2026.acl-long.347/