Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situations youth commonly face. The authors argue that existing child-safety evaluations of AI lack grounding in the real-world harms youth experience, rest on unvalidated assumptions about what counts as an appropriate output (refusal in particular), and typically look only at detecting adversarial prompts or surface-level output harms. Practitioners identified both chatbot behaviors likely to cause harm and behaviors that could meaningfully support youth, and gave concrete recommendations for chatbot responses and for evaluation infrastructure.
Publisher
arXiv preprint
Published
8 Aug 2026
Added
6 days ago
DOI
—
Key Findings
- Current child-safety evaluations rest on unvalidated assumptions about appropriate output, with refusal treated as a safe default; practitioners disputed that assumption
- Evaluations that focus on adversarial prompt detection or surface-level output harms can miss responses that harm youth in practice
- 19 practitioners working directly with vulnerable youth (social workers, therapists, psychologists) reviewed chatbot responses to risky situations drawn from prior empirical work
- Practitioners distinguished chatbot behaviors likely to cause harm from behaviors that could meaningfully support youth in difficult moments, and set out what role chatbots should and should not play
- The paper recommends incorporating practitioner perspectives into child-safety evaluation design and supporting infrastructure
Methodology Notes
Qualitative interview study, n=19 practitioners; stimuli were chatbot responses to youth-risk situations established in prior empirical work. arXiv 2608.07902, submitted 2026-08-08, cs.CY (also cs.AI, cs.HC). The abstract page records acceptance at the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026); the proceedings version is not yet published, so this row is the preprint. Author affiliations are not stated on the arXiv abstract page. Small purposive practitioner sample; findings are qualitative and not intended as prevalence estimates.
Sources
arXiv abstract page (2608.07902) (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Hannah Cha, Neha Shukla, Solon Barocas, Alexandra Chouldechova, Eugenia Kim, Jennifer Wortman Vaughan
Tags
Cite This
APA
Hannah Cha et al. (2026). Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety. arXiv preprint. https://arxiv.org/abs/2608.07902
Related Insights
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Understanding Teen Overreliance on AI Companion Chatbots Through Self-Reported Reddit Narratives
ACM (Proceedings of CHI 2026) · 13 Apr 2026
Talking to machines: Children's experiences with AI assistants and companions
Australia eSafety Commissioner · 3 Aug 2026
AI chatbots and youth suicide risk: current evidence, critical gaps, and a clinical research agenda
npj Digital Medicine · 1 Aug 2026