Charting the evolution of artificial intelligence mental health chatbots from rule-based systems to large language models: a systematic review
A systematic review in World Psychiatry, the official journal of the World Psychiatric Association, of 160 studies (2020-2024) classifying mental-health chatbot architectures (rule-based, machine-learning, and large language model based) and proposing a three-tier evaluation framework: foundational bench testing, pilot feasibility testing, and clinical efficacy testing. It documents a validation gap in which LLM chatbots surged to 45% of new studies in 2024 but only 16% underwent clinical-efficacy testing.
Key Findings
- LLM-based chatbots rose to 45% of new studies in 2024, yet only 16% of LLM studies underwent clinical-efficacy testing and most (77%) remained in early validation.
- Only 47% of the 160 studies focused on clinical-efficacy testing, exposing a gap in robust validation of therapeutic benefit.
- Documents discrepancies between marketed claims ('AI-powered') and actual architectures, with many interventions relying on simple rule-based scripts, and proposes a three-tier evaluation framework aligned with medical-AI certification.
Methodology Notes
Systematic review of 160 studies published 2020-2024. World Psychiatry 2025;24(3):383-394. Verified via PubMed (PMID 40948070) and Crossref (DOI 10.1002/wps.21352); the Wiley page blocks automated fetchers. Published 2025-09-15 per Crossref.
Sources
World Psychiatry Article (primary)
Authors
Yining Hua, Steve Siddals, Zilin Ma, Isaac Galatzer-Levy, Winna Xia, Christine Hau, Hongbin Na, Matthew Flathers, Jake Linardon, Cyrus Ayubcha, John Torous
Tags
Cite This
APA
Yining Hua et al. (2025). Charting the evolution of artificial intelligence mental health chatbots from rule-based systems to large language models: a systematic review. World Psychiatry. https://onlinelibrary.wiley.com/doi/10.1002/wps.21352