Skip to main content
Peer-reviewed Credible — Major labs, established NGOs, reputable named-author preprints

Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review

A JBI-methodology, PRISMA-ScR scoping review of how purpose-built generative-AI mental health chatbot interventions actually implement safety, based on a prospectively registered and peer-reviewed protocol and a July 2025 search of seven databases. Twenty-one studies across eleven countries were included and synthesised across technical safeguards, pre-deployment safety work and delivery-phase risk mitigation. The review reports that crisis referral protocols were mostly underdeveloped, human oversight limited, and systematic adverse-event monitoring sparse.

Publisher

Healthcare (MDPI); University of British Columbia

Published

20 May 2026

Added

today

Key Findings

  • Twenty-one studies across 11 countries met inclusion; two reviewers screened and extracted independently.
  • Most interventions used at least one technical safety mechanism, most commonly fine-tuning and prompt engineering; only a smaller subset used layered architectures combining retrieval, content filters or risk classifiers, and rule-based algorithms.
  • Pre-deployment safeguards that did appear were clinical-expert and user co-design, research-ethics procedures and data-privacy measures.
  • At delivery, detailed onboarding with role clarification was common but human oversight was limited; crisis referral protocols varied in rigour and were mostly underdeveloped, and systematic adverse-event monitoring was sparse.
  • Documented safety failures across the included studies were missed suicidal ideation and provision of inaccurate clinical information.
  • The authors call for regulatory oversight proportional to the risk these systems carry and for standardised safety-outcome measurement.

Methodology Notes

Scoping review following Joanna Briggs Institute methodology and PRISMA-ScR, protocol prospectively registered and peer-reviewed; search of MEDLINE, Scopus, PsycINFO, ACM Digital Library, IEEE Xplore, Google Scholar and Consensus in July 2025; narrative synthesis across three pre-specified domains. A synthesis of published intervention studies, so it maps what developers report rather than what deployed products do, and the July 2025 cut-off predates a year of rapid product change. Published in Healthcare 14(10):1395, 2026-05-20, CC BY 4.0, DOI 10.3390/healthcare14101395; the earlier PsyArXiv preprint (osf.io/preprints/psyarxiv/g8q5v, 2026-04-06) is superseded by this version.

Authors

Lotenna Olisaeloka, Chris G. Richardson, Angel Y. Wang, Richard J. Munthali, Daniel V. Vigo

Tags

scoping-reviewprisma-scrubccrisis-referraladverse-eventsmdpi

Cite This

APA

Lotenna Olisaeloka et al. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare (MDPI); University of British Columbia. https://doi.org/10.3390/healthcare14101395