Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

Evaluating AI Safety in Teen Conversations

Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety threats, requests for medical or therapeutic advice, and teens treating the chatbot as a parent, partner or main source of support. Conversations were rated against 27 safety checks by a panel of LLM judges, with blind clinician review of a subset. A separate comparison tested a teen-specific system instruction.

Publisher

Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine

Published

23 Sept 2026

Added

today

DOI

—

Key Findings

  • 178 of 648 conversations (27.5%) failed at least one critical safety check. In 111 of those 178 (62%) a failure occurred later in the exchange, a median of five replies after the chatbot first responded to the concern.
  • Share of each model's 72 conversations with a critical failure: GLM 5.2 47.2%, DeepSeek V4 Flash 38.9%, Qwen 3.7 Plus 38.9%, Gemini 3.6 Flash 31.9%, Kimi K3 Instant 26.4%, Grok 4.5 Fast 23.6%, Claude Sonnet 5 16.7%, GPT-5.5 Instant 13.9%, Meta Muse Spark 1.2 9.7%. The authors state that the scenario-weighting ranges overlap, so the data do not support a ranking.
  • Fail counts among conversations where each check applied: presenting the AI as a substitute relationship 42 of 42, diagnosis overreach 109 of 169, missed connection to human support 173 of 291, claims beyond the AI's role 232 of 490, help concealing a suicide attempt 4 of 52, normalising language on suicide or self-harm 5 of 346, assistance with serious violence 0 of 15.
  • Nine conversations failed specifically on suicide or self-harm handling, and every one of those failures came after the first reply. Examples include agreeing to help hide a suicide attempt from parents and agreeing with a teen's refusal of crisis support.
  • A system instruction based on OpenAI's published under-18 principles reduced conversations with a critical failure from 31.1% (84 of 270) to 11.5% (31 of 270) across nine models and 30 scenarios.

Methodology Notes

Web research report on vals.ai with interactive transcripts; the byline is dated 09/23/2026. It follows an announcement post dated 2026-08-12. Five mental-health clinicians wrote and peer-reviewed scenarios about fictional teens aged 13-17: 27 on self-harm and other safety threats, 22 on medical or therapeutic advice, 23 on relationships with AI. Teen turns were generated by GPT-5.6 Terra, chosen after a blind clinician comparison of candidate simulators. Two judges (GPT-5.6 Sol and Claude Sonnet 5) rated each applicable check pass, partial or fail, with Gemini 3.6 Flash breaking ties. Only a majority 'fail' counted; ratings without a majority were labelled uncertain. Clinicians blind-reviewed selected conversations to check agreement. Models were accessed by API with settings meant to approximate free consumer offerings and no added instructions, so consumer-app features (age signals, memory, extra filters) are absent. The authors state the results are not a measure of what teens see in consumer apps. Neither the scenario set nor the rubric is published as a downloadable dataset. Not peer reviewed; published by a commercial evaluation company. Acknowledged collaborators: Valerie Chen and Diyi Yang (Stanford SALT Lab) and Eric Lin, MD (Stanford School of Medicine).

Authors

Andrea Mock, Caroline Figueroa

Tags

vals-aiteen-safetymulti-turnclinician-authoredllm-as-judgesystem-prompt

Cite This

APA

Andrea Mock, Caroline Figueroa. (2026). Evaluating AI Safety in Teen Conversations. Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine. https://www.vals.ai/blogs/evaluating-ai-safety-in-teen-conversations