Skip to main content

Browse the library

The complete record — 223 artifacts, last updated 20 Aug 2026. Also available as JSON and RSS (CC BY 4.0).

Filters:

77 artifacts matching

17 Aug 2026 British Journal of Clinical Psychology (Wiley, for the British Psychological Society) Peer-reviewed

The new first listener: Daydreaming styles and self-compassion predict adolescent disclosure to AI, differently in ADHD

Cross-sectional survey of 2,115 adolescents and young people in the United States and Hong Kong examining how daydreaming styles and self-compassion relate to disclosing inner thoughts to an AI chatb…

11 Aug 2026 arXiv preprint (University of Warwick / Forensic Capability Network) Benchmark / dataset

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…

8 Aug 2026 arXiv preprint Preprint

Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…

5 Aug 2026 arXiv (Stanford-led author team) Preprint

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…

1 Aug 2026 npj Digital Medicine Peer-reviewed

AI chatbots and youth suicide risk: current evidence, critical gaps, and a clinical research agenda

Review by researchers at Crisis Text Line examining how general-purpose chatbots and AI companions detect and respond to suicide-risk disclosures from young people, and what the existing evidence can…

24 Jul 2026 arXiv preprint Preprint

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…

16 Jul 2026 Meta Lab publication

Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI

Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…

15 Jul 2026 Proceedings of the IASEAI Conference (published by AAAI) Peer-reviewed

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…

14 Jul 2026 medRxiv (Beth Israel Deaconess / Harvard digital-psychiatry group) Preprint

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 d…

9 Jul 2026 Mila (Quebec AI Institute) & ROOST Lab publication

Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection

Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Benchmark / dataset

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…

29 Jun 2026 JMIR AI Peer-reviewed

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…

25 Jun 2026 ACM (Proceedings of FAccT 2026) Peer-reviewed

Characterizing Delusional Spirals through Human-LLM Chat Logs

Peer-reviewed analysis of chat logs from 19 users reporting psychological harm from chatbot use, applying a 28-code inventory to 391,562 messages. Characterises how delusion-reinforcing interaction p…

25 Jun 2026 JAAD International (Elsevier, for the American Academy of Dermatology) Peer-reviewed

Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations

Research letter testing whether ChatGPT recognises mental health concerns and recommends appropriate referral during simulated multi-turn psychodermatology conversations. Fifty first-person narrative…

21 Jun 2026 PsyArXiv (Corporal Michael J. Crescenz VA Medical Center; University of Pennsylvania; Stanford; Columbia University and others) Preprint

Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure

An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…

11 Jun 2026 JMIR Mental Health Peer-reviewed

Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models

Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…

11 Jun 2026 Partnership on AI NGO report

How AI Companies are Handling Suicide and Self-Harm Today

Drawing on a March 2026 multistakeholder workshop convening frontier AI companies, clinicians, researchers, and people with lived experience, Partnership on AI presents a taxonomy of six intervention…

9 Jun 2026 Anthropic Lab publication

System Card: Claude Fable 5 & Claude Mythos 5

Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…

9 Jun 2026 arXiv (Emory University-led) Benchmark / dataset

Expert-Level Crisis Detection in Mental Health Conversations

A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…

3 Jun 2026 JMIR AI Peer-reviewed

Suicidal Ideation in Online Spaces Through the Lens of Interpersonal Theory of Suicide: Exploratory Study of Self-Disclosure, Peer Support, and AI Responses

A peer-reviewed exploratory study analysing 59,607 Reddit r/SuicideWatch posts through the Interpersonal Theory of Suicide (IPTS) framework, categorising expressions of suicidal ideation by IPTS dime…