Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation
Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions were generated with GPT-4o artificial users built from clinical vignettes varying on seven characteristics, and ten clinical experts (licensed psychotherapists or advanced trainees) rated each session on the 14-item Quality of Behavioral Activation Scale plus therapeutic capabilities and artificial-user realism. The chatbot completed every intervention phase, rated highest on message safety and clarity and lowest on clinical reasoning and rapport; the artificial users were rated unrealistic and undemanding, and no session triggered the crisis protocol because the simulated high-risk user never disclosed suicidal thoughts despite a persona specifying suicidality.
Publisher
JMIR Mental Health (JMIR Publications)
Published
1 Sept 2026
Added
today
Key Findings
- Mean Q-BAS rating 4.03 (SD 1.18) on a 0-6 scale and holistic session quality 3.94 (SD 1.23); 13 of 14 components exceeded the satisfactory threshold of 3
- Highest component ratings for mood assessment (5.42) and activity planning (4.98); lowest for explaining positive reinforcement (2.92) and supporting activity-mood monitoring (3.02)
- Therapeutic capability ratings were highest for message safety (5.90, SD 0.37) and clarity (5.56) and lowest for therapeutic rapport (4.12) and natural conversation flow (4.25)
- Artificial users were rated below the scale midpoint for authenticity (2.75) and difficulty (1.23) and were often highly compliant
- No session contained an explicit disclosure of suicidal ideation, intent or self-harm even though high-severity personas specified suicidality, so the crisis protocol was never triggered; the authors conclude artificial-user testing 'cannot establish safety under real distress, resistance, or crisis disclosure' and recommend predefined high-risk scripts
- Experts described the chatbot as structured, clear and safe but identified insufficient clinical reasoning about the suitability and feasibility of activities, barriers and rewards as the main limitation
Methodology Notes
JMIR Mental Health 2026;13:e94781, DOI 10.2196/94781, published 2026-09-01 (PMC epub date; PMID 42679232, PMC13533316), CC BY 4.0; full text read via PMC because mental.jmir.org is blocked from this box. Affiliations: Institute for Information Systems, Karlsruhe Institute of Technology; Department of Clinical Psychology and Psychotherapy, Universität Greifswald; Clinical Child and Adolescent Psychology, Saarland University. No human users and no comparison condition; the Q-BAS threshold is not validated for chatbot delivery; complete prompts are in Multimedia Appendix 1. One author reports consultancy fees from e-mental-health companies. Resolves the JMIR accepted-but-unpublished watch item for this DOI.
Sources
JMIR Mental Health article(opens in a new tab) (primary)
PubMed Central full text (CC BY)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Florian Onur Kuhlmeier, Leon Hanschmann, Melina Rabe, Stefan Lüttke, Eva-Lotta Brakemeier, Alexander Maedche
Tags
Cite This
APA
Florian Onur Kuhlmeier et al. (2026). Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation. JMIR Mental Health (JMIR Publications). https://mental.jmir.org/2026/1/e94781
Related Insights
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
arXiv (preprint) · 21 Aug 2026
Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions
Journal of Marital and Family Therapy (Wiley, for the American Association for Marriage and Family Therapy) · 31 Aug 2026
The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7
Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) · 1 Jul 2026