59 artifacts matching
Preprint
Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment
Audits 31 pre-specified NLP techniques from seven methodological families (model scaling, synthetic data, loss reweighting, ensembling, structured prediction, threshold tuning and LLM methods) in rou…
Preprint
Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports
Cross-sectional secondary analysis of 185 deidentified accounts of mental-health harm linked with AI chatbot use (95 first-hand, 90 from relatives, partners or friends) submitted through the web form…
Preprint
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…
Peer-reviewed
Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data
Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-int…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Preprint
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
Evaluates suicide-risk classification from de-identified transcripts of calls to Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Calls were transcribed on site with a Levant…
Peer-reviewed
Benevolent Gravity: the lethal structure inherent in conversational AI design principles
Analytical paper arguing that the two dominant explanations for fatalities linked to conversational AI — safety-filter failure and commodified intimacy — are structurally insufficient, because a subs…
Peer-reviewed
AI chatbots and youth suicide risk: current evidence, critical gaps, and a clinical research agenda
Review by researchers at Crisis Text Line examining how general-purpose chatbots and AI companions detect and respond to suicide-risk disclosures from young people, and what the existing evidence can…
Lab publication
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Lab publication
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Peer-reviewed
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…
NGO report
Google Search: AI Overview & AI Mode — AI Risk Assessment
A product risk assessment of the generative-AI features built into Google Search (AI Overview and AI Mode) as experienced by child and teen accounts, rated Unacceptable Risk against the assessing org…
Preprint
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 d…
Lab publication
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
Lab publication
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
Peer-reviewed
Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media
Synthesis of research on AI-based suicide detection in social media, combining an umbrella review of 22 systematic reviews covering studies up to 2022 with an ongoing literature update, yielding 195…
Peer-reviewed
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…
Lab publication
GPT-5.6 Preview System Card
OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…
Peer-reviewed
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Peer-reviewed study testing whether aggregated expert judgment yields valid ground truth for training and evaluating AI systems in mental-health safety contexts. Three certified psychiatrists indepen…
Preprint
Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure
An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…