Skip to main content
Preprint Preliminary — Early preprints, credible essays, unreviewed grey literature

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Pairs four replicated mental-health safety benchmarks with an ecological audit of 20,000 deployment conversations to compare a purpose-built mental-health AI (Ash) with six general-purpose models from four families (OpenAI GPT-5, GPT-5.1 and GPT-5.2; DeepSeek V3; Google Gemini 3 Flash; Moonshot Kimi K2). The purpose-built system produced lower potentially harmful content rates on suicide and self-harm, eating-disorder and substance-use prompts than every comparator, and the authors argue that small simulation benchmarks under-represent the linguistic diversity of real use and should be complemented by ecological auditing.

Publisher

arXiv (Slingshot AI)

Published

14 Jan 2026

Added

today

DOI

Key Findings

  • On the CCDH benchmark replication, Ash produced potentially harmful content in 6.2% of responses versus 18.0% to 52.0% for the six frontier comparators (all p < .001)
  • In the 20,000-conversation deployment audit, clinician review within the audit pipeline found no suicide-risk conversations lacking crisis resources and three self-injury conversations without crisis intervention, a within-pipeline conditional rate of 3 in 800 (0.38%)
  • Against blinded clinician adjudication of 600 randomly sampled deployment conversations, the language-model judge showed 100% sensitivity (6 of 6; 95% CI 54.1 to 100%) and 99.2% specificity (589 of 594), and the conversational model delivered crisis resources in all six clinician-confirmed cases
  • Version 2 (3 August 2026) runs to 43 pages and eight figures

Methodology Notes

Vendor-authored evaluation of the authors' own product against general-purpose models: benchmark replications plus an ecological audit of real conversations with clinician review; the six clinician-confirmed crisis cases in the 600-conversation adjudication sample give wide confidence intervals. v1 posted 14 January 2026, v2 3 August 2026; date is the arXiv v1 submission date. Abstract verified on arXiv; full text not read by the sweep. A later paper from the same group on the same product, 'Engagement Phenotypes for a Sample of 102,684 AI Mental Health Chatbot Users and Dose-Response Associations with Clinical Outcomes' (arXiv 2605.00275, v2 2 July 2026), reports five engagement phenotypes (early dropouts 52.2%, weekly users 25.3%, concentrated 16.8%, intensive 4.1%, power users 1.6%) and overnight sessions for 66.9% of users; it is recorded as an additional source rather than a separate row.

Authors

Stamatis, Caitlin A., Meyerhoff, Jonah, Zhang, Richard, Tieleman, Olivier, Malgaroli, Matteo, Hull, Thomas D.

Tags

ashslingshot-aibenchmarkecological-auditvendor-researchccdh

Cite This

APA

Stamatis, Caitlin A. et al. (2026). Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety. arXiv (Slingshot AI). https://arxiv.org/abs/2601.17003