Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions

Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing on 352,649 conversations, the authors argue that clinician review of every AI output fails at scale because of vigilance decay, and describe the three-layer human-on-the-loop framework they arrived at: preventive design (scope restrictions, adversarial and over-trigger benchmark scenarios, therapist review before launch, an external advisory committee), real-time monitoring by a second independent safety model on every client message, and continuous clinician quality assurance combining an LLM-as-judge on a clinician rubric with random review of unflagged conversations.

Publisher

arXiv (Grow Therapy; Stanford University School of Medicine)

Published

8 Sept 2026

Added

today

Key Findings

  • The safety layer screens every client message in parallel with the coaching model for suicidal ideation, self-harm, psychotic symptoms, intimate-partner violence and medical emergencies; on a serious signal the conversation is paused, the 988 Suicide and Crisis Lifeline is displayed and the treating therapist is notified in the provider portal.
  • Over 50% of providers in the 25,000-plus network have at least one client using the tool; the tool is positioned as between-session support, not a substitute for therapy.
  • Clinician review found that the conversation memory context window was causing the coach to accumulate the client's framing over many turns and increasingly mirror their emotional stance instead of gently challenging it; the fix was to shorten the prior-conversation history available to the model.
  • A second review found safety triggers firing disproportionately to the situation, leading to a recalibration of risk thresholds to reduce over-triggering without losing detection.
  • Automated quality metrics show a shift to more positive or hopeful language by the end of the conversation in 61.9% of conversations and greater stated motivation or readiness to act in 64.9%; no safety-event rates or controlled outcomes are reported.

Methodology Notes

Retrospective account of a live deployment; no control condition, no denominator of safety events, no disclosure of the underlying language models. Three authors are Grow Therapy employees, one is an advisor and one collaborates with the company (declared). arXiv 2609.09533 version 1, 8 September 2026 (announced 10 September); US; not peer reviewed.

Authors

Matthew A. Scult, John L. Havlik, Kevin Ramotar, Ethan Goh, Manoj Kanagaraj

Tags

grow-therapyhuman-on-the-loopdeploymentbetween-session-coachingsafety-classifier988

Cite This

APA

Matthew A. Scult et al. (2026). Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions. arXiv (Grow Therapy; Stanford University School of Medicine). https://arxiv.org/abs/2609.09533