Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis

PRISMA systematic review of quantitative studies evaluating general-purpose large language models for direct mental health care tasks, searched across PubMed, Embase, ACM Digital Library, IEEE Xplore and Google Scholar from November 2022 to March 2026. Methodological quality was appraised with the Mixed Methods Appraisal Tool and certainty of evidence with GRADE; random-effects meta-analyses with Hartung-Knapp adjustment and prediction intervals were run on screening and diagnosis studies of GPT-4, GPT-3.5 and GPT-3.

Publisher

JMIR AI (JMIR Publications)

Published

31 Aug 2026

Added

today

Key Findings

  • 66 studies were included: 37 vignette or simulation studies, 22 retrospective studies and 7 prospective studies.
  • Applications: screening and diagnosis (29 studies), clinical decision support (14), treatment support (10), documentation and monitoring (6), patient education (4) and patient engagement (3).
  • Meta-analysis of 8 screening and diagnosis studies: GPT-4 pooled sensitivity 0.83 (95% CI 0.38 to 0.97; 95% prediction interval 0.02 to 1.00) versus GPT-3.5 0.70 and GPT-3 0.61; GPT-4 specificity 0.77 (95% CI 0.52 to 0.91; prediction interval 0.10 to 0.99).
  • Certainty of evidence was generally low across domains; only documentation and monitoring reached moderate certainty.
  • The authors conclude the evidence is too heterogeneous, indirect and uncertain to support routine unsupervised use, particularly for diagnosis, risk assessment, crisis response or therapeutic interaction, and that broad accessibility should not be mistaken for clinical readiness.

Methodology Notes

Systematic review and meta-analysis; quantitative synthesis restricted to OpenAI models and to studies reporting sensitivity and specificity; prediction intervals span nearly 0 to 1, so pooled estimates are unstable. Limitations named by the authors: indirectness from vignette designs, uncertain representativeness, inconsistent outcome reporting, sparse prospective real-world evaluation. Published 2026-08-31 in JMIR AI volume 5. The publisher's article page returns an anti-bot interstitial (HTTP 202) from the sweep host, so verification used the Crossref record, which carries the publisher-deposited structured abstract; funding and competing-interest statements were not visible through that route.

Authors

Janni Leung, Benjamin Johnson, Kelsey McRae, Stephanie Fong, Paula Cardona Gonzalez, Caitlin McClure-Thomas, Tianze Sun, Naomi Kern, Yuen Ming Wong, Richard Liu, Gary Chung Kai Chan

Tags

systematic-reviewmeta-analysisgradeclinical-readinessjmir-ai

Cite This

APA

Janni Leung et al. (2026). Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis. JMIR AI (JMIR Publications). https://ai.jmir.org/2026/1/e87730