Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT
Mixed-methods study of how young adults aged 18 to 25 use ChatGPT when distressed, built on complete donated ChatGPT histories (19,930 conversations from 158 participants) plus a survey that included the PHQ-4. Five distress conversations were then reviewed and rewritten by ten licensed clinicians. Distressed participants reported more emotional engagement with ChatGPT and more behaviour change from using it. Clinicians endorsed the assistant's availability and much of its wording but identified seven process failures, including not checking whether the young person was safe.
Publisher
arXiv (University of Washington-led; with Stanford University, University of Oxford, Georgetown University School of Medicine and The University of Texas at Austin)
Published
28 Sept 2026
Added
today
DOI
—
Key Findings
- 158 young adults (49.2% female) donated chat histories totalling 19,930 ChatGPT conversations; 39.9% scored in the distressed range on the PHQ-4
- Distressed participants reported significantly higher Emotional Engagement and Behavioral Change scores than non-distressed peers (both survived FDR correction); Trust and Self-Efficacy pointed the same way but were marginal, and Dependency Concern showed no association
- The first author read the user messages in all 19,930 conversations; after homework conversations were removed (5,549 remained), two lead authors identified 77 conversations with high emotional distress, from which five were selected for clinician review (self-harm, a delusion, sexual assault, a breakup and shame about sexual desire)
- Ten clinicians produced 33 rewritten responses and named seven process failures: claiming knowledge it could not have, offering solutions before understanding the situation, giving more advice than the person could use, mixing roles, not checking safety, compliance that itself caused harm, and responses that could dysregulate
- Clinician rewrites followed three ordered stages (ask about safety first, bring intensity down, explore without agreeing), which the authors translate into design guidelines for general-purpose conversational agents
Methodology Notes
Participants recruited through Prolific, Reddit, school flyers and referrals (expressions of interest November 2024 to March 2026; observation dates use month precision); eligibility required ChatGPT use at least ten times in the previous two weeks; complete ChatGPT data exports plus a two-part survey (PHQ-4 and perceived ChatGPT experience measures); Welch t-tests with Cohen's d and FDR correction. Clinician phase: ten licensed clinicians who had worked with clients aged 18 to 25 completed a worksheet and semi-structured interview on five selected conversations. Selection of five illustrative conversations is purposive, so the clinician findings are qualitative; the survey is cross-sectional self-report and the sample is not representative. Manuscript is formatted for CHI 2027 but no acceptance is stated. v1 submitted 2026-09-28 17:53 UTC; announced in the Wednesday 30 Sep cs.AI listing. Verified from the arXiv abs page and HTML render (both HTTP 200).
Sources
arXiv preprint(opens in a new tab) (primary)
arXiv HTML render (v1)(opens in a new tab) (28 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Marx Wang, Ella Zhang, Cameron Tan, Andrea Mock, Songling Ngo, Zijing Wang, Robert Wolfe, Shirin Amouei, Rachel A. Hanebutt, Desmond C. Ong, Caroline Figueroa, Katie Davis, Anind K. Dey, Alexis Hiniker
Tags
Cite This
APA
Marx Wang et al. (2026). Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT. arXiv (University of Washington-led; with Stanford University, University of Oxford, Georgetown University School of Medicine and The University of Texas at Austin). https://arxiv.org/abs/2609.35953
Related Insights
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
arXiv (Slingshot AI) · 14 Jan 2026
AI Chatbot Use and Disclosure for Mental Health Among US Adolescents and Young Adults
JAMA Pediatrics (American Medical Association); RAND-led author team · 1 Jun 2026
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts
arXiv (School of Computing and Information, University of Pittsburgh) · 29 Aug 2026