A Taxonomy of AI Chatbot Harms in Substance Use Disorder Recovery: Content Analysis of Online Recovery Communities
Qualitative content analysis of user-reported interactions with general-purpose AI chatbots posted in 20 public Reddit substance-use-disorder recovery communities between the public release of ChatGPT (30 November 2022) and 27 March 2026. From 763,665 posts, keyword filtering and double screening yielded 842 firsthand chatbot-interaction posts and 61 unique posts describing harmful chatbot behaviour, which were coded into a seven-domain harm taxonomy. Inaccurate or misleading guidance and unregulated clinical authority (dosing and taper advice, advice to override a prescriber) were the most frequent domains; crisis-recognition failure, facilitation of return to use, undermining of human recovery support, recovery-identity undermining and reciprocated emotional attachment were less frequent but higher acuity.
Publisher
JMIR Preprints (JMIR Publications)
Published
23 Sept 2026
Added
today
Key Findings
- Corpus: 763,665 posts from 20 Reddit SUD recovery communities filtered with 32 chatbot/LLM terms to 1,285 candidates; 842 eligible firsthand interactions (screening agreement 91%, kappa 0.80); 61 unique harm posts after duplicate removal
- Seven harm domains: crisis recognition and escalation failure; facilitation of return to substance use; unregulated clinical authority; inaccurate or misleading guidance; undermining of nonprofessional recovery support; recovery-identity undermining; inappropriate AI emotional attachment or dependency (the seventh emerged during coding)
- Inaccurate or misleading guidance appeared in 52 of 61 harm posts, in four forms: fabricated or overconfident specifics, false recovery or withdrawal timelines, false reassurance, and false catastrophic framing
- Unregulated clinical authority included patient-specific dosing, alcohol and benzodiazepine taper schedules and advice to override a prescriber; all medication-assisted-treatment taper recommendations were coded as both unregulated authority and inaccurate guidance
- Domain-level inter-rater reliability kappa 0.84 across 364 decisions on 52 posts; harm screening kappa 0.65 (prevalence-adjusted 0.87, 93% agreement)
- The authors frame the taxonomy as a basis for safety audits and user-facing interventions, noting that several behaviours need contextual, multi-turn assessment
Methodology Notes
Qualitative content analysis of self-reported Reddit posts from a single platform; chatbot behaviour was coded from users' descriptions rather than transcripts, ambiguous cases were coded as non-harmful, and the harm corpus is small (61 posts). Two coders with three pilot rounds; a prespecified six-domain SUD-specific framework plus one emergent domain. Status: unpublished, non-peer-reviewed preprint submitted to JMIR Mental Health on 2026-09-23 (JMIR Preprints 112742). Verification route: the preprint page is a JavaScript shell to plain fetches; the manuscript PDF was read through the r.jina.ai reader; Crossref carries the abstract. Corresponding author at the Luddy School of Informatics, Computing, and Engineering, Indiana University Indianapolis. The curator could not re-read the manuscript (the reader proxy returned 403 in the curator session); the Crossref record (title, authors, abstract sentence, posted 2026-09-23) was verified directly and the manuscript details rest on the beat agent read.
Topics
Authors
Dong Whi Yoo, Tirthraj Rathod, Jaihyun Park, Jesse C. Stewart, Brandon G. Oberlin
Tags
Cite This
APA
Dong Whi Yoo et al. (2026). A Taxonomy of AI Chatbot Harms in Substance Use Disorder Recovery: Content Analysis of Online Recovery Communities. JMIR Preprints (JMIR Publications). https://preprints.jmir.org/preprint/112742
Related Insights
Generative artificial intelligence addiction is prospectively associated with psychotic-like experiences: Evidence from a two-wave longitudinal mediation and network study
Computers in Human Behavior Reports (Elsevier); Southern Medical University; Jilin University of Finance and Economics; South China Normal University · 9 Sept 2026
Large language models for psychosocial risk assessment: A multi-method evaluation across suicide, intimate partner violence, and substance misuse
PLOS Digital Health · 27 Apr 2026