45 artifacts matching
AI chatbots and youth suicide risk: current evidence, critical gaps, and a clinical research agenda
Review by researchers at Crisis Text Line examining how general-purpose chatbots and AI companions detect and respond to suicide-risk disclosures from young people, and what the existing evidence can…
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 d…
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…
GPT-5.6 Preview System Card
OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Peer-reviewed study testing whether aggregated expert judgment yields valid ground truth for training and evaluating AI systems in mental-health safety contexts. Three certified psychiatrists indepen…
Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure
An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…
How AI Companies are Handling Suicide and Self-Harm Today
Drawing on a March 2026 multistakeholder workshop convening frontier AI companies, clinicians, researchers, and people with lived experience, Partnership on AI presents a taxonomy of six intervention…
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
Expert-Level Crisis Detection in Mental Health Conversations
A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…
Suicidal Ideation in Online Spaces Through the Lens of Interpersonal Theory of Suicide: Exploratory Study of Self-Disclosure, Peer Support, and AI Responses
A peer-reviewed exploratory study analysing 59,607 Reddit r/SuicideWatch posts through the Interpersonal Theory of Suicide (IPTS) framework, categorising expressions of suicidal ideation by IPTS dime…
System Card: Claude Opus 4.8
Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from pub…
AI Mental Health Apps (Common Sense Media Youth AI Safety Institute Risk Assessment)
Risk assessment of five AI mental health apps — three direct-to-consumer (Wysa, Earkick, Youper) and two deployed through school districts (Alongside, Sonar) — tested against an eight-principle rubri…
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…