Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
The authors introduce an ontology of ten therapeutic moves, compact function-based categories grounded in the MULTI-60 psychotherapy process inventory, validated through an annotation campaign with five licensed psychologists and scaled with a judge-based method matched to expert agreement. Applying it to real counselling transcripts and to model-led sessions, they compare move distributions between human clinicians and a panel of frontier models. Exposing the ontology to models as a set of tools reduces the deviation from the human distribution without fine-tuning.
Key Findings
- A ten-move, function-based therapeutic ontology grounded in the MULTI-60 inventory and validated by five licensed psychologists
- Models over-use inquiry at up to three times the human clinician rate
- Models neglect psychoeducation relative to human clinicians
- Models are strongly context-anchored: they carry forward strategies a human clinician initiated but rarely initiate those strategies themselves
- Exposing the ontology as a set of tools roughly halves mean deviation from the human move distribution and improves turn-level alignment with human therapists by 7-9 percentage points, with no fine-tuning
Methodology Notes
Annotation campaign with five licensed psychologists; an LLM judge validated against expert agreement and used to scale annotation; applied to real counselling transcripts and to model-led sessions across a panel of frontier models. The abstract does not state the panel composition, the transcript corpus size, or the source and consent basis of the real counselling transcripts. Preprint, not peer-reviewed, no venue listed; v1 submitted 2026-08-21.
Sources
arXiv abstract (arXiv:2608.21325) (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos, Anabela C. Areias, Maya D'Eon, Fabiola Costa, Ricardo Rei, Nuno M. Guerreiro
Tags
Cite This
APA
Afonso Baldo et al. (2026). Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy. arXiv (preprint). https://arxiv.org/abs/2608.21325
Related Insights
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
Journal of Psychopathology and Clinical Science (American Psychological Association) · 17 Aug 2026
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Large Language Model–Based Chatbots and Agentic AI for Mental Health Counseling: Systematic Review of Methodologies, Evaluation Frameworks, and Ethical Safeguards
JMIR AI · 13 Mar 2026
Beyond Engagement: A Multidimensional Framework to Evaluate the Safe Development of Agentic AI in Mental Health
Lecture Notes in Computer Science (Springer Nature) — AI for Clinical Applications · 22 Sept 2025