Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review
A JBI-methodology, PRISMA-ScR scoping review of how purpose-built generative-AI mental health chatbot interventions actually implement safety, based on a prospectively registered and peer-reviewed protocol and a July 2025 search of seven databases. Twenty-one studies across eleven countries were included and synthesised across technical safeguards, pre-deployment safety work and delivery-phase risk mitigation. The review reports that crisis referral protocols were mostly underdeveloped, human oversight limited, and systematic adverse-event monitoring sparse.
Publisher
Healthcare (MDPI); University of British Columbia
Published
20 May 2026
Added
today
Key Findings
- Twenty-one studies across 11 countries met inclusion; two reviewers screened and extracted independently.
- Most interventions used at least one technical safety mechanism, most commonly fine-tuning and prompt engineering; only a smaller subset used layered architectures combining retrieval, content filters or risk classifiers, and rule-based algorithms.
- Pre-deployment safeguards that did appear were clinical-expert and user co-design, research-ethics procedures and data-privacy measures.
- At delivery, detailed onboarding with role clarification was common but human oversight was limited; crisis referral protocols varied in rigour and were mostly underdeveloped, and systematic adverse-event monitoring was sparse.
- Documented safety failures across the included studies were missed suicidal ideation and provision of inaccurate clinical information.
- The authors call for regulatory oversight proportional to the risk these systems carry and for standardised safety-outcome measurement.
Methodology Notes
Scoping review following Joanna Briggs Institute methodology and PRISMA-ScR, protocol prospectively registered and peer-reviewed; search of MEDLINE, Scopus, PsycINFO, ACM Digital Library, IEEE Xplore, Google Scholar and Consensus in July 2025; narrative synthesis across three pre-specified domains. A synthesis of published intervention studies, so it maps what developers report rather than what deployed products do, and the July 2025 cut-off predates a year of rapid product change. Published in Healthcare 14(10):1395, 2026-05-20, CC BY 4.0, DOI 10.3390/healthcare14101395; the earlier PsyArXiv preprint (osf.io/preprints/psyarxiv/g8q5v, 2026-04-06) is superseded by this version.
Topics
Authors
Lotenna Olisaeloka, Chris G. Richardson, Angel Y. Wang, Richard J. Munthali, Daniel V. Vigo
Tags
Cite This
APA
Lotenna Olisaeloka et al. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare (MDPI); University of British Columbia. https://doi.org/10.3390/healthcare14101395
Related Insights
Charting the evolution of artificial intelligence mental health chatbots from rule-based systems to large language models: a systematic review
World Psychiatry · 15 Sept 2025
Large Language Model–Based Chatbots and Agentic AI for Mental Health Counseling: Systematic Review of Methodologies, Evaluation Frameworks, and Ethical Safeguards
JMIR AI · 13 Mar 2026
Ethical evaluation of AI-supported mental health applications
Frontiers in Psychiatry; Dr. Abdurrahman Yurtaslan Ankara Oncology Training and Research Hospital; TOBB ETU Faculty of Medicine · 28 Aug 2026