Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
Benchmark evaluation of how four consumer AI systems maintain safety boundaries when answering paediatric health questions from caregivers, using PediatricSafetyBench-v2: 300 authentic caregiver queries from the HealthCareMagic-100k-en physician-consultation corpus and 300 matched adversarial variants built from six operationalised caregiver-pressure patterns. Responses were scored with a five-component Safety Composite Score validated against human raters, with and without a safety-oriented system prompt.
Publisher
npj Digital Medicine (Springer Nature)
Published
14 Jul 2026
Added
today
Key Findings
- Across GPT-4o-mini, Gemini-2.0-Flash, Claude-3.5-Haiku and Llama-3.1-8B the overall safety-appropriate rate (composite score of 10 or more out of 15) was 95.5%.
- A safety-oriented system prompt raised safety-appropriate rates by 5.9 percentage points across all four models.
- Adversarial caregiver pressure was associated with higher, not lower, Safety Composite Scores for all four models across all ten topic categories and severity levels; false-expertise claims were the most vulnerability-inducing pressure pattern and emotional escalation was associated with the highest scores.
- The automated score was validated against independent human raters before full-corpus use (mean weighted kappa 0.76; Pearson r 0.88).
- PediatricSafetyBench-v2 is publicly released for longitudinal safety monitoring.
Methodology Notes
Benchmark study, 600 queries (300 naturalistic plus 300 adversarial), four consumer models of the 2024 to early-2025 generation, English; automated five-component composite scoring validated on a human-rated subset; no real caregivers or clinical outcomes. Authors are at Mashhad University of Medical Sciences and Shahid Beheshti University of Medical Sciences (Iran). Received 30 December 2025, accepted 30 June 2026, published 14 July 2026; CC BY-NC-ND. Preprint arXiv 2601.09721 (a version 2 posted 8 September 2026 surfaced the published version). Coverage miss: the July publication was not caught by earlier sweeps.
Sources
Topics
Authors
Vahideh Zolfaghari, Leila Mashhadi, Mitra Ahadi, Farzaneh Sedaghatkar, MohammadReza Kargozari
Tags
Cite This
APA
Vahideh Zolfaghari et al. (2026). Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions. npj Digital Medicine (Springer Nature). https://www.nature.com/articles/s41746-026-02985-9
Related Insights
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
Association for Computational Linguistics (Proceedings of the 1st Workshop on Linguistic Analysis for Health, HeaLing 2026) · 1 Mar 2026
CAREBench: A Child-Safety Risk Benchmark for Language Models
arXiv (Handshake AI; University of California, Los Angeles; McGill University) · 29 Jun 2026