Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

CAREBench: A Child-Safety Risk Benchmark for Language Models

Introduces CAREBench (Child AI Risk Evaluation), a 500-prompt benchmark of upstream child-safety risks in language models: assistance that helps adults manipulate, impersonate, profile or isolate minors, and responses that deepen a child's emotional dependence on the AI rather than redirecting toward human support. Twelve risk categories include grooming and relationship engineering, deception and impersonation, sextortion and image-based abuse, AI anthropomorphization, emotional dependency, therapist replacement and intersection with major mental illness; explicit abuse material is excluded. Responses are judged by an ensemble of three LLM judges calibrated against parent and clinician annotations, and seven frontier models fail between 2.3% and 58.0% of prompts.

Publisher

arXiv (Handshake AI; University of California, Los Angeles; McGill University)

Published

29 Jun 2026

Added

today

DOI

Key Findings

  • Failure rates across seven frontier APIs range from 2.3% (Claude Fable 5) to 58.0% (GPT-5.4); Claude Opus 4.6 fails 7.0%, GPT-5.5 21.7%, Gemini 3.1 Pro 31.9%, Grok 4.1 Fast Reasoning 33.1% and Kimi K2 Thinking 33.6%
  • Failures concentrate in the relational and mental-health categories: GPT-5.5's highest failure rates are AI anthropomorphization (56%), online grooming (47%), emotional dependency (44%) and intersection with major mental illness (39%); the best-scoring model's residual failures are therapist replacement (26%) and emotional dependency (17%)
  • Grok 4.1 produces 11 of the 13 responses rated maximum severity by all three judges, and is weakest in the relational and mental-health categories
  • Judging uses a weighted ensemble of Claude Opus 4.6, GPT-5.4 and Gemini 3.1 Pro; on a held-out parent-labelled set the ensemble reaches Cohen's kappa 0.55 against human labels, comparable to parent-parent agreement (kappa 0.525, 78.0%), and humans prefer the ensemble's 'acceptable' response 89.3% of the time
  • A 30-parent panel annotated the developmental-context categories, a credentialed clinician the clinical categories and a child-safety practitioner the grooming severity; the clinician rated 73% of reviewed prompts very realistic and 88% very plausible
  • Prompts, annotated responses, annotated preferences and all benchmark rollouts are released on Hugging Face under CC BY 4.0 with Apache-2.0 evaluation code

Methodology Notes

arXiv 2606.29685, v1 submitted 2026-06-29 (cs.LG). Affiliations from the PDF title block: Handshake AI (all authors), UCLA and McGill for two authors; correspondence addresses are at joinhandshake.com, so this is a vendor-authored benchmark from a data-labelling company. 500 prompts across twelve risk domains and six elicitation styles, evaluated over three runs per model with a MultiJudge ensemble (weights 0.7 Opus 4.6, 0.25 GPT-5.4, 0.05 Gemini 3.1 Pro) calibrated on 205 prompts x five models (1,021 labels). Inter-rater agreement among parents is moderate (kappa 0.525) and 29.2% of pairwise preferences were disputed, so single-response verdicts carry noise. Two of the judged models are also judges, an acknowledged self-preference risk. Dataset presence verified via the Hugging Face API (handshake-ai-research/CAREBench, created 2026-06-15, four CSV splits, ~30 MB, CC BY 4.0, gated auto-approval) and the GitHub repository (Handshake-AI-Research/CAREBench, Apache-2.0).

Authors

Kaavya Krishna-Kumar, Elaine Lau, Vaughn Robinson, Jay Caldwell, Sheriff Issaka, Skyler Wang, Francisco Guzmán, Steven Kelling, Jonas Mueller

Tags

child-safetybenchmarkgroomingemotional-dependencyvendor-benchmarkllm-judgehandshake

Cite This

APA

Kaavya Krishna-Kumar et al. (2026). CAREBench: A Child-Safety Risk Benchmark for Language Models. arXiv (Handshake AI; University of California, Los Angeles; McGill University). https://arxiv.org/abs/2606.29685