Large Language Model–Based Chatbots and Agentic AI for Mental Health Counseling: Systematic Review of Methodologies, Evaluation Frameworks, and Ethical Safeguards
A systematic review synthesizing the methodologies, evaluation practices, and ethical/governance frameworks reported in studies of large language model chatbots and agentic AI used for mental-health counseling, and identifying recurring gaps in validation and safety reporting.
Key Findings
- GPT-based models featured in roughly 45% of reviewed studies; around 90% used fine-tuned or domain-adapted models.
- External validation of systems was frequently missing across the reviewed literature.
- Reporting of safety, ethics, and governance practices was inconsistent across studies.
Methodology Notes
Peer-reviewed systematic review (JMIR AI 2026;5:e80348; DOI 10.2196/80348; published 2026-03-13). Covers LLM chatbots and agentic AI in mental-health counseling. The JMIR HTML rendered empty to the automated fetcher; title, venue, and date verified via Crossref DOI metadata.
Sources
JMIR AI (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Ha Na Cho, Jiayuan Wang, Di Hu, Kai Zheng
Tags
Cite This
APA
Ha Na Cho et al. (2026). Large Language Model–Based Chatbots and Agentic AI for Mental Health Counseling: Systematic Review of Methodologies, Evaluation Frameworks, and Ethical Safeguards. JMIR AI. https://ai.jmir.org/2026/1/e80348
Related Insights
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
Association for Computational Linguistics (EACL 2026) · 24 Mar 2026
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
arXiv · 3 Mar 2026