Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions
Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing on 352,649 conversations, the authors argue that clinician review of every AI output fails at scale because of vigilance decay, and describe the three-layer human-on-the-loop framework they arrived at: preventive design (scope restrictions, adversarial and over-trigger benchmark scenarios, therapist review before launch, an external advisory committee), real-time monitoring by a second independent safety model on every client message, and continuous clinician quality assurance combining an LLM-as-judge on a clinician rubric with random review of unflagged conversations.
Publisher
arXiv (Grow Therapy; Stanford University School of Medicine)
Published
8 Sept 2026
Added
today
Key Findings
- The safety layer screens every client message in parallel with the coaching model for suicidal ideation, self-harm, psychotic symptoms, intimate-partner violence and medical emergencies; on a serious signal the conversation is paused, the 988 Suicide and Crisis Lifeline is displayed and the treating therapist is notified in the provider portal.
- Over 50% of providers in the 25,000-plus network have at least one client using the tool; the tool is positioned as between-session support, not a substitute for therapy.
- Clinician review found that the conversation memory context window was causing the coach to accumulate the client's framing over many turns and increasingly mirror their emotional stance instead of gently challenging it; the fix was to shorten the prior-conversation history available to the model.
- A second review found safety triggers firing disproportionately to the situation, leading to a recalibration of risk thresholds to reduce over-triggering without losing detection.
- Automated quality metrics show a shift to more positive or hopeful language by the end of the conversation in 61.9% of conversations and greater stated motivation or readiness to act in 64.9%; no safety-event rates or controlled outcomes are reported.
Methodology Notes
Retrospective account of a live deployment; no control condition, no denominator of safety events, no disclosure of the underlying language models. Three authors are Grow Therapy employees, one is an advisor and one collaborates with the company (declared). arXiv 2609.09533 version 1, 8 September 2026 (announced 10 September); US; not peer reviewed.
Sources
arXiv preprint(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Matthew A. Scult, John L. Havlik, Kevin Ramotar, Ethan Goh, Manoj Kanagaraj
Tags
Cite This
APA
Matthew A. Scult et al. (2026). Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions. arXiv (Grow Therapy; Stanford University School of Medicine). https://arxiv.org/abs/2609.09533
Related Insights
Automated Safety Testing and Reporting Application for Conversational Safety Monitoring of Generative AI Tools for Mental Health: Development and Validation Study
JMIR Mental Health · 19 May 2026
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Nature Medicine · 7 Aug 2026
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
arXiv · 2 Jan 2026