77 artifacts matching
The new first listener: Daydreaming styles and self-compassion predict adolescent disclosure to AI, differently in ADHD
Cross-sectional survey of 2,115 adolescents and young people in the United States and Hong Kong examining how daydreaming styles and self-compassion relate to disclosing inner thoughts to an AI chatb…
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
AI chatbots and youth suicide risk: current evidence, critical gaps, and a clinical research agenda
Review by researchers at Crisis Text Line examining how general-purpose chatbots and AI companions detect and respond to suicide-risk disclosures from young people, and what the existing evidence can…
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 d…
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…
Characterizing Delusional Spirals through Human-LLM Chat Logs
Peer-reviewed analysis of chat logs from 19 users reporting psychological harm from chatbot use, applying a 28-code inventory to 391,562 messages. Characterises how delusion-reinforcing interaction p…
Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations
Research letter testing whether ChatGPT recognises mental health concerns and recommends appropriate referral during simulated multi-turn psychodermatology conversations. Fifty first-person narrative…
Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure
An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…
How AI Companies are Handling Suicide and Self-Harm Today
Drawing on a March 2026 multistakeholder workshop convening frontier AI companies, clinicians, researchers, and people with lived experience, Partnership on AI presents a taxonomy of six intervention…
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
Expert-Level Crisis Detection in Mental Health Conversations
A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…
Suicidal Ideation in Online Spaces Through the Lens of Interpersonal Theory of Suicide: Exploratory Study of Self-Disclosure, Peer Support, and AI Responses
A peer-reviewed exploratory study analysing 59,607 Reddit r/SuicideWatch posts through the Interpersonal Theory of Suicide (IPTS) framework, categorising expressions of suicidal ideation by IPTS dime…