When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha
A study of whether the models behind therapy apps and general chatbots correctly assess clinical risk when adolescents express distress in age-native language: hyperbolic phrasing, ironic positivity, rapid semantic drift and contextual polysemy. The authors build two benchmarks, 64 Generation Alpha mental-health expressions validated by native speakers and clinicians, and 75 paired Standard/Gen Alpha multi-turn conversations totalling 780 turns, and test Claude, GPT-4o and Llama-3.1. Models comprehend the vocabulary substantially better than they calibrate the clinical risk it carries, a gap the authors do not find in human therapists.
Publisher
ACM (Proceedings of FAccT 2026)
Published
25 Jun 2026
Added
today
Key Findings
- Models understand 76-82% of Generation Alpha vocabulary but correctly calibrate only 64-72% of clinical risk, a 10-14 percentage point gap (p<.001, d>0.48)
- Human therapists show a 3 percentage point gap on the same material (p=.22), so the failure is model-specific rather than inherent to the task
- The gap is consistent across the three model architectures tested and widens with ambiguity, from 7 to 18 percentage points
- Six failure patterns are quantified: minimization acceptance (43pp), sarcasm masking (29pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp)
- The paper cites 13.1% of US adolescents, about 5.4 million, as using generative AI for mental health advice
- Benchmark reliability: ICC=0.72 for native-speaker validation, kappa=0.78 for clinician validation
Methodology Notes
Two constructed benchmarks evaluated against three model families (Claude, GPT-4o, Llama-3.1). Expressions and conversations are authored rather than collected from real adolescent users, so ecological validity rests on the native-speaker and clinician validation rather than on observed use. Peer-reviewed FAccT 2026 proceedings paper, pages 1681-1720, DOI 10.1145/3805689.3806522, published 2026-06-25 per the Crossref record. Date note: the arXiv mirror carries a v1 submission timestamp of 2026-06-14 and was only announced on arXiv around 2026-08-21, so the same paper carries three different plausible dates; the proceedings publication date is used. The paper also reports a modelled annual missed-crisis extrapolation from a baseline miss rate, which is a projection rather than an observed count and should be cited as such. The ACM Digital Library landing page returns 403 to automated fetchers; the record was verified against the Crossref metadata and the arXiv mirror.
Topics
Authors
Manisha Mehta, Virendra Mehta
Tags
Cite This
APA
Manisha Mehta, Virendra Mehta (2026). When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha. ACM (Proceedings of FAccT 2026). https://doi.org/10.1145/3805689.3806522
Related Insights
Characterizing Delusional Spirals through Human-LLM Chat Logs
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
AI Chatbot Use and Disclosure for Mental Health Among US Adolescents and Young Adults
JAMA Pediatrics (American Medical Association); RAND-led author team · 1 Jun 2026
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
arXiv preprint · 8 Aug 2026