Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow

A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minute chatbot 'Guided Intake' conversation offered to U.S. adults in an employer-sponsored mental health benefit. A stratified probability sample of 400 of 30,532 conversations (August 2025 to June 2026) was rated by two licensed psychologists blind to the agent's output, and estimates were weighted back to the full population.

Publisher

Research Square (preprint); Spring Health (Spring Care Inc)

Published

1 Oct 2026

Added

today

Key Findings

  • Among conversations clinicians judged to contain actionable suicide risk, the agent escalated 93.2% (95% CI 88.8-97.8%, weighted sensitivity); specificity was 99.7% (99.6-99.8%)
  • Positive predictive value of an escalation was 80.7% in the abstract and 80.9% (74.9-86.8%) in the results table, a false-escalation rate of about 19%
  • No clinician-rated Non-Immediate or Immediate risk conversation was classified as no risk; 7 sampled conversations (an estimated 27 population-wide) were under-called by one level and not escalated
  • The agent classified 93.7% of conversations as no risk and escalated 1.5% (24 Immediate, 431 Non-Immediate); agreement with recent PHQ-9/C-SSRS self-report was low (Cohen's kappa 0.09)
  • In 5.1% of conversations where the participant had denied suicidal thoughts on a recent assessment, the agent flagged at least possible risk
  • Agent-clinician quadratic-weighted kappa was 0.72 against clinician-clinician kappa 0.91; the safety agent cost about $0.04 of a $0.25 intake conversation

Methodology Notes

Retrospective cohort, Yale IRB 2000029276, waiver of consent; adults 18+ in the U.S. whose employer opted into the platform's AI tools and who completed Guided Intake (abandoned conversations excluded). Two Spring Health psychologists rated 400 transcripts (50 double-rated); disagreements resolved to the more severe rating; design-based stratified bootstrap with inverse-probability weighting. The chatbot listened for spontaneously disclosed risk and asked clarifying questions; it did not screen everyone. Limitations stated by the authors: suicide risk only (not harm to others, psychosis or mania), no youth, detection only (response quality not evaluated), single workflow with 24/7 clinician backup. All authors are employed by and hold equity in Spring Care Inc; no external funding. The abstract and the results table give slightly different PPV figures (80.7% vs 80.9%). Posted 2026-10-01 (Research Square, CC BY 4.0); not peer reviewed.

Authors

Emily J Ward, Kate H Bentley, Emily R Dworkin, Matt Hawrilenko, Emily Van Ark, John DeLorenzo, Millard Brown, Adam M Chekroud

Tags

spring-healthsuicide-risk-detectionc-ssrsreal-world-validationsafety-agentgpt-4oinverse-probability-weighting

Cite This

APA

Emily J Ward et al. (2026). Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow. Research Square (preprint); Spring Health (Spring Care Inc). https://www.researchsquare.com/article/rs-10882429/v1