Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
Pairs four replicated mental-health safety benchmarks with an ecological audit of 20,000 deployment conversations to compare a purpose-built mental-health AI (Ash) with six general-purpose models from four families (OpenAI GPT-5, GPT-5.1 and GPT-5.2; DeepSeek V3; Google Gemini 3 Flash; Moonshot Kimi K2). The purpose-built system produced lower potentially harmful content rates on suicide and self-harm, eating-disorder and substance-use prompts than every comparator, and the authors argue that small simulation benchmarks under-represent the linguistic diversity of real use and should be complemented by ecological auditing.
Publisher
arXiv (Slingshot AI)
Published
14 Jan 2026
Added
today
DOI
—
Key Findings
- On the CCDH benchmark replication, Ash produced potentially harmful content in 6.2% of responses versus 18.0% to 52.0% for the six frontier comparators (all p < .001)
- In the 20,000-conversation deployment audit, clinician review within the audit pipeline found no suicide-risk conversations lacking crisis resources and three self-injury conversations without crisis intervention, a within-pipeline conditional rate of 3 in 800 (0.38%)
- Against blinded clinician adjudication of 600 randomly sampled deployment conversations, the language-model judge showed 100% sensitivity (6 of 6; 95% CI 54.1 to 100%) and 99.2% specificity (589 of 594), and the conversational model delivered crisis resources in all six clinician-confirmed cases
- Version 2 (3 August 2026) runs to 43 pages and eight figures
Methodology Notes
Vendor-authored evaluation of the authors' own product against general-purpose models: benchmark replications plus an ecological audit of real conversations with clinician review; the six clinician-confirmed crisis cases in the 600-conversation adjudication sample give wide confidence intervals. v1 posted 14 January 2026, v2 3 August 2026; date is the arXiv v1 submission date. Abstract verified on arXiv; full text not read by the sweep. A later paper from the same group on the same product, 'Engagement Phenotypes for a Sample of 102,684 AI Mental Health Chatbot Users and Dose-Response Associations with Clinical Outcomes' (arXiv 2605.00275, v2 2 July 2026), reports five engagement phenotypes (early dropouts 52.2%, weekly users 25.3%, concentrated 16.8%, intensive 4.1%, power users 1.6%) and overnight sessions for 66.9% of users; it is recorded as an additional source rather than a separate row.
Sources
Topics
Authors
Stamatis, Caitlin A., Meyerhoff, Jonah, Zhang, Richard, Tieleman, Olivier, Malgaroli, Matteo, Hull, Thomas D.
Tags
Cite This
APA
Stamatis, Caitlin A. et al. (2026). Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety. arXiv (Slingshot AI). https://arxiv.org/abs/2601.17003
Related Insights
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026
Generic AI or Nothing: Support-Seeking Patterns After Market Withdrawal of a Purpose-Built AI Wellbeing Tool
PsyArXiv (OSF); Slingshot AI research team and collaborators · 30 Mar 2026