CAREBench: A Child-Safety Risk Benchmark for Language Models
Introduces CAREBench (Child AI Risk Evaluation), a 500-prompt benchmark of upstream child-safety risks in language models: assistance that helps adults manipulate, impersonate, profile or isolate minors, and responses that deepen a child's emotional dependence on the AI rather than redirecting toward human support. Twelve risk categories include grooming and relationship engineering, deception and impersonation, sextortion and image-based abuse, AI anthropomorphization, emotional dependency, therapist replacement and intersection with major mental illness; explicit abuse material is excluded. Responses are judged by an ensemble of three LLM judges calibrated against parent and clinician annotations, and seven frontier models fail between 2.3% and 58.0% of prompts.
Publisher
arXiv (Handshake AI; University of California, Los Angeles; McGill University)
Published
29 Jun 2026
Added
today
DOI
—
Key Findings
- Failure rates across seven frontier APIs range from 2.3% (Claude Fable 5) to 58.0% (GPT-5.4); Claude Opus 4.6 fails 7.0%, GPT-5.5 21.7%, Gemini 3.1 Pro 31.9%, Grok 4.1 Fast Reasoning 33.1% and Kimi K2 Thinking 33.6%
- Failures concentrate in the relational and mental-health categories: GPT-5.5's highest failure rates are AI anthropomorphization (56%), online grooming (47%), emotional dependency (44%) and intersection with major mental illness (39%); the best-scoring model's residual failures are therapist replacement (26%) and emotional dependency (17%)
- Grok 4.1 produces 11 of the 13 responses rated maximum severity by all three judges, and is weakest in the relational and mental-health categories
- Judging uses a weighted ensemble of Claude Opus 4.6, GPT-5.4 and Gemini 3.1 Pro; on a held-out parent-labelled set the ensemble reaches Cohen's kappa 0.55 against human labels, comparable to parent-parent agreement (kappa 0.525, 78.0%), and humans prefer the ensemble's 'acceptable' response 89.3% of the time
- A 30-parent panel annotated the developmental-context categories, a credentialed clinician the clinical categories and a child-safety practitioner the grooming severity; the clinician rated 73% of reviewed prompts very realistic and 88% very plausible
- Prompts, annotated responses, annotated preferences and all benchmark rollouts are released on Hugging Face under CC BY 4.0 with Apache-2.0 evaluation code
Methodology Notes
arXiv 2606.29685, v1 submitted 2026-06-29 (cs.LG). Affiliations from the PDF title block: Handshake AI (all authors), UCLA and McGill for two authors; correspondence addresses are at joinhandshake.com, so this is a vendor-authored benchmark from a data-labelling company. 500 prompts across twelve risk domains and six elicitation styles, evaluated over three runs per model with a MultiJudge ensemble (weights 0.7 Opus 4.6, 0.25 GPT-5.4, 0.05 Gemini 3.1 Pro) calibrated on 205 prompts x five models (1,021 labels). Inter-rater agreement among parents is moderate (kappa 0.525) and 29.2% of pairwise preferences were disputed, so single-response verdicts carry noise. Two of the judged models are also judges, an acknowledged self-preference risk. Dataset presence verified via the Hugging Face API (handshake-ai-research/CAREBench, created 2026-06-15, four CSV splits, ~30 MB, CC BY 4.0, gated auto-approval) and the GitHub repository (Handshake-AI-Research/CAREBench, Apache-2.0).
Topics
Authors
Kaavya Krishna-Kumar, Elaine Lau, Vaughn Robinson, Jay Caldwell, Sheriff Issaka, Skyler Wang, Francisco Guzmán, Steven Kelling, Jonas Mueller
Tags
Cite This
APA
Kaavya Krishna-Kumar et al. (2026). CAREBench: A Child-Safety Risk Benchmark for Language Models. arXiv (Handshake AI; University of California, Los Angeles; McGill University). https://arxiv.org/abs/2606.29685
Related Insights
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
arXiv preprint · 8 Aug 2026
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
arXiv (Nanyang Technological University; National University of Singapore) · 26 Aug 2026
Social AI Companions: AI Risk Assessment
Common Sense Media; Stanford School of Medicine Brainstorm Lab for Mental Health Innovation · 30 Apr 2025
AI Mental Health Apps (Common Sense Media Youth AI Safety Institute Risk Assessment)
Common Sense Media Youth AI Safety Institute, with Stanford Medicine Brainstorm Lab · 5 May 2026