39 artifacts matching
Lab publication
Continuing To Build Upon Our Safety Priorities
A first-party safety update from Character.AI describing safeguards in operation on its platform as of September 2026. It states that self-harm safeguards consider the surrounding conversation, inclu…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Government report
Through Children's Eyes: How Digital Technologies Enable Child Sexual Abuse
Synthesis of two rounds of the Disrupting Harm surveys, nationally representative household surveys of around 21,000 internet-using children aged 12 to 17 across 21 countries in Africa, Asia, Latin A…
Lab publication
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
Lab publication
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Lab publication
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Preprint
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Lab publication
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Lab publication
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
Lab publication
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
Lab publication
GPT-Live System Card
System card for GPT-Live-1 and GPT-Live-1 mini, OpenAI's full-duplex voice models that became the default voice models for paid and free ChatGPT users respectively. The card introduces voice-native s…
Benchmark / dataset
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Benchmark of 31,920 benign 'boundary' health prompts, generated by paraphrasing 2,306 health-related toxic seed prompts across seven categories including self-harm, medical misinformation and unquali…
Peer-reviewed
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts
Large-scale analysis of how people describe using AI for emotional support or therapy in 47 DSM-5-mapped mental-health subreddits between November 2022 and August 2025. From 3.5 million posts, a 146-…
Benchmark / dataset
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Names and measures 'intent legitimation': benign, truthfully accumulated user memories bias a personalized dialogue agent's inference of intent so that an inherently harmful request is treated as con…
Lab publication
GPT-5.6 Preview System Card
OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…
Peer-reviewed
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…
Lab publication
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
Benchmark / dataset
Expert-Level Crisis Detection in Mental Health Conversations
A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…
Preprint
Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback
A controlled prompt-level evaluation of how language models respond to eating-disorder-related requests, built with eating-disorder clinicians. The authors constructed 11,712 prompts that vary four f…