Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

39 artifacts matching

3 Sept 2026 Character.AI (Character Technologies) Lab publication

Lab publication

Continuing To Build Upon Our Safety Priorities

A first-party safety update from Character.AI describing safeguards in operation on its platform as of September 2026. It states that self-harm safeguards consider the surrounding conversation, inclu…

3 Sept 2026 OpenAI Lab publication

Lab publication

GPT-6 Astra System Card

System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…

1 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…

1 Sept 2026 UNICEF Innocenti (Office of Strategy and Evidence), with the Disrupting Harm partnership Government report

Government report

Through Children's Eyes: How Digital Technologies Enable Child Sexual Abuse

Synthesis of two rounds of the Disrupting Harm surveys, nationally representative household surveys of around 21,000 internet-using children aged 12 to 17 across 21 countries in Africa, Asia, Latin A…

12 Aug 2026 xAI Lab publication

Lab publication

Model Card: Grok 4.6

36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…

6 Aug 2026 OpenAI Lab publication

Lab publication

GPT-5.6 – August Updates

System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…

24 Jul 2026 Anthropic Lab publication

Lab publication

System Card: Claude Opus 5

System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…

24 Jul 2026 arXiv preprint Preprint

Preprint

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…

16 Jul 2026 Meta Lab publication

Lab publication

Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI

Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…

9 Jul 2026 Mila (Quebec AI Institute) & ROOST Lab publication

Lab publication

Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection

Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…

9 Jul 2026 OpenAI Lab publication

Lab publication

GPT-5.6 System Card

General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…

8 Jul 2026 OpenAI Lab publication

Lab publication

GPT-Live System Card

System card for GPT-Live-1 and GPT-Live-1 mini, OpenAI's full-duplex voice models that became the default voice models for paid and free ChatGPT users respectively. The card introduces voice-native s…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Benchmark / dataset

Benchmark / dataset

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

Benchmark of 31,920 benign 'boundary' health prompts, generated by paraphrasing 2,306 health-related toxic seed prompts across seven categories including self-harm, medical misinformation and unquali…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Peer-reviewed

Peer-reviewed

Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts

Large-scale analysis of how people describe using AI for emotional support or therapy in 47 DSM-5-mapped mental-health subreddits between November 2022 and August 2025. From 3.5 million posts, a 146-…

1 Jul 2026 Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers) Benchmark / dataset

Benchmark / dataset

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Names and measures 'intent legitimation': benign, truthfully accumulated user memories bias a personalized dialogue agent's inference of intent so that an inherently harmful request is treated as con…

26 Jun 2026 OpenAI Lab publication superseded

Lab publication

GPT-5.6 Preview System Card

OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…

11 Jun 2026 JMIR Mental Health Peer-reviewed

Peer-reviewed

Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models

Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…

9 Jun 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5 & Claude Mythos 5

Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…

9 Jun 2026 arXiv (Emory University-led) Benchmark / dataset

Benchmark / dataset

Expert-Level Crisis Detection in Mental Health Conversations

A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…

1 Jun 2026 arXiv (University of Aberdeen; University of Colorado Anschutz; Heriot-Watt University; University College London) Preprint

Preprint

Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

A controlled prompt-level evaluation of how language models respond to eating-disorder-related requests, built with eating-disorder clinicians. The authors constructed 11,712 prompts that vary four f…