Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT

Mixed-methods study of how young adults aged 18 to 25 use ChatGPT when distressed, built on complete donated ChatGPT histories (19,930 conversations from 158 participants) plus a survey that included the PHQ-4. Five distress conversations were then reviewed and rewritten by ten licensed clinicians. Distressed participants reported more emotional engagement with ChatGPT and more behaviour change from using it. Clinicians endorsed the assistant's availability and much of its wording but identified seven process failures, including not checking whether the young person was safe.

Publisher

arXiv (University of Washington-led; with Stanford University, University of Oxford, Georgetown University School of Medicine and The University of Texas at Austin)

Published

28 Sept 2026

Added

today

DOI

—

Key Findings

  • 158 young adults (49.2% female) donated chat histories totalling 19,930 ChatGPT conversations; 39.9% scored in the distressed range on the PHQ-4
  • Distressed participants reported significantly higher Emotional Engagement and Behavioral Change scores than non-distressed peers (both survived FDR correction); Trust and Self-Efficacy pointed the same way but were marginal, and Dependency Concern showed no association
  • The first author read the user messages in all 19,930 conversations; after homework conversations were removed (5,549 remained), two lead authors identified 77 conversations with high emotional distress, from which five were selected for clinician review (self-harm, a delusion, sexual assault, a breakup and shame about sexual desire)
  • Ten clinicians produced 33 rewritten responses and named seven process failures: claiming knowledge it could not have, offering solutions before understanding the situation, giving more advice than the person could use, mixing roles, not checking safety, compliance that itself caused harm, and responses that could dysregulate
  • Clinician rewrites followed three ordered stages (ask about safety first, bring intensity down, explore without agreeing), which the authors translate into design guidelines for general-purpose conversational agents

Methodology Notes

Participants recruited through Prolific, Reddit, school flyers and referrals (expressions of interest November 2024 to March 2026; observation dates use month precision); eligibility required ChatGPT use at least ten times in the previous two weeks; complete ChatGPT data exports plus a two-part survey (PHQ-4 and perceived ChatGPT experience measures); Welch t-tests with Cohen's d and FDR correction. Clinician phase: ten licensed clinicians who had worked with clients aged 18 to 25 completed a worksheet and semi-structured interview on five selected conversations. Selection of five illustrative conversations is purposive, so the clinician findings are qualitative; the survey is cross-sectional self-report and the sample is not representative. Manuscript is formatted for CHI 2027 but no acceptance is stated. v1 submitted 2026-09-28 17:53 UTC; announced in the Wednesday 30 Sep cs.AI listing. Verified from the arXiv abs page and HTML render (both HTTP 200).

Authors

Marx Wang, Ella Zhang, Cameron Tan, Andrea Mock, Songling Ngo, Zijing Wang, Robert Wolfe, Shirin Amouei, Rachel A. Hanebutt, Desmond C. Ong, Caroline Figueroa, Katie Davis, Anind K. Dey, Alexis Hiniker

Tags

chatgptyoung-adultsdata-donationclinician-reviewphq-4design-guidelines

Cite This

APA

Marx Wang et al. (2026). Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT. arXiv (University of Washington-led; with Stanford University, University of Oxford, Georgetown University School of Medicine and The University of Texas at Austin). https://arxiv.org/abs/2609.35953