Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation

Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions were generated with GPT-4o artificial users built from clinical vignettes varying on seven characteristics, and ten clinical experts (licensed psychotherapists or advanced trainees) rated each session on the 14-item Quality of Behavioral Activation Scale plus therapeutic capabilities and artificial-user realism. The chatbot completed every intervention phase, rated highest on message safety and clarity and lowest on clinical reasoning and rapport; the artificial users were rated unrealistic and undemanding, and no session triggered the crisis protocol because the simulated high-risk user never disclosed suicidal thoughts despite a persona specifying suicidality.

Publisher

JMIR Mental Health (JMIR Publications)

Published

1 Sept 2026

Added

today

Key Findings

  • Mean Q-BAS rating 4.03 (SD 1.18) on a 0-6 scale and holistic session quality 3.94 (SD 1.23); 13 of 14 components exceeded the satisfactory threshold of 3
  • Highest component ratings for mood assessment (5.42) and activity planning (4.98); lowest for explaining positive reinforcement (2.92) and supporting activity-mood monitoring (3.02)
  • Therapeutic capability ratings were highest for message safety (5.90, SD 0.37) and clarity (5.56) and lowest for therapeutic rapport (4.12) and natural conversation flow (4.25)
  • Artificial users were rated below the scale midpoint for authenticity (2.75) and difficulty (1.23) and were often highly compliant
  • No session contained an explicit disclosure of suicidal ideation, intent or self-harm even though high-severity personas specified suicidality, so the crisis protocol was never triggered; the authors conclude artificial-user testing 'cannot establish safety under real distress, resistance, or crisis disclosure' and recommend predefined high-risk scripts
  • Experts described the chatbot as structured, clear and safe but identified insufficient clinical reasoning about the suitability and feasibility of activities, barriers and rewards as the main limitation

Methodology Notes

JMIR Mental Health 2026;13:e94781, DOI 10.2196/94781, published 2026-09-01 (PMC epub date; PMID 42679232, PMC13533316), CC BY 4.0; full text read via PMC because mental.jmir.org is blocked from this box. Affiliations: Institute for Information Systems, Karlsruhe Institute of Technology; Department of Clinical Psychology and Psychotherapy, Universität Greifswald; Clinical Child and Adolescent Psychology, Saarland University. No human users and no comparison condition; the Q-BAS threshold is not validated for chatbot delivery; complete prompts are in Multimedia Appendix 1. One author reports consultancy fees from e-mental-health companies. Resolves the JMIR accepted-but-unpublished watch item for this DOI.

Authors

Florian Onur Kuhlmeier, Leon Hanschmann, Melina Rabe, Stefan Lüttke, Eva-Lotta Brakemeier, Alexander Maedche

Tags

behavioral-activationgpt-4oartificial-userssimulated-patientsexpert-ratinggermanyyouth

Cite This

APA

Florian Onur Kuhlmeier et al. (2026). Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation. JMIR Mental Health (JMIR Publications). https://mental.jmir.org/2026/1/e94781