Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT Sessions
Compares session-level behaviour of human peer counsellors and an LLM counsellor trained on the same single-session CBT manual. An 18-month ethnography on a peer-support platform (seven counsellors, 110 self-counselling sessions, 60 focus groups) refined CBT prompts across several model families; a session-generation method paired client turns from 27 public human-led CBT sessions with LLM counsellor turns; three licensed clinical psychologists evaluated 27 human and 27 LLM sessions. Human peers adapted CBT to context and built rapport at the cost of structure; LLM counsellors adhered to technique but struggled with turn-taking, salience and lecturing, and produced what the authors call deceptive empathy.
Publisher
Proceedings of the ACM on Human-Computer Interaction (CSCW 2026); Brown University
Published
23 Sept 2026
Added
today
Key Findings
- LLM counsellors showed greater methodological adherence to CBT techniques but failed to sustain turn-taking, often could not distinguish clinically important from trivial content, and were more prone to lecturing and imposing solutions
- Human peer counsellors used small talk, contextual self-disclosure and cultural adaptation to build rapport, often at the expense of session structure and therapeutic focus
- LLM counsellors produced 'deceptive empathy': excessively anthropomorphic responses that can inflate user expectations of human care
- Prompts were iterated across GPT-3, GPT-3.5, GPT-4, Llama 3.1 and 3.2, Claude 3 Sonnet and Claude 3 Haiku over 110 self-counselling sessions
- The authors conclude that single-turn empathy advantages do not carry to leading multi-turn sessions and map guardrails for hybrid human-AI systems
Methodology Notes
Three-phase mixed-methods study: 18-month ethnography with seven trained peer counsellors; controlled session generation (client responses drawn from 27 publicly available human-led CBT transcripts, counsellor responses generated by a CBT-prompted LLM); expert evaluation of 27 human and 27 LLM sessions by three licensed clinical psychologists. v3 (2026-09-21) states it places greater emphasis on qualitative analysis given the controlled generative synthesis of LLM sessions. Published in PACM HCI vol. 10 no. 6, article CSCW196 (October 2026 issue); Crossref registers DOI 10.1145/3817112 with issued date 2026-09-23.
Sources
PACM HCI (CSCW 2026) article(opens in a new tab) (primary)
arXiv preprint (v3, 2026-09-21)(opens in a new tab) (21 Sept 2026)
Authors
Zainab Iftikhar, Sean Ransom, Amy Wei Xiao, Nicole Nugent, Jeff Huang
Tags
Cite This
APA
Zainab Iftikhar et al. (2026). Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT Sessions. Proceedings of the ACM on Human-Computer Interaction (CSCW 2026); Brown University. https://doi.org/10.1145/3817112
Related Insights
Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
ACM (Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency) · 23 Jun 2025
Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial
JMIR Mental Health · 6 Jun 2017
Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions
Journal of Marital and Family Therapy (Wiley, for the American Association for Marriage and Family Therapy) · 31 Aug 2026