55 artifacts matching
Lab publication
System Card: Claude Sonnet 5.5
System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…
NGO report
Social network e chatbot: uno studio dei rischi per utenti vulnerabili e minori
Report of the CDCR (Cittadino Digitale Critico e Responsabile) project, funded by the Italian Ministry of Enterprises and Made in Italy. Part one is a legal analysis of how Facebook, Instagram, TikTo…
Benchmark / dataset
Evaluating AI Safety in Teen Conversations
Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Lab publication
Model Card: Grok 4.7
Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…
Benchmark / dataset
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper ev…
Peer-reviewed
Temporal and cross-site validation of an AI system for self-harm detection
A prospective and external validation of a self-harm detection system built on emergency-department triage notes, testing how far it travels in time and across hospitals. The model was developed on 2…
Regulator study
Online Safety and Children's Digital Lives: Report of a National Survey of Children and their Parents/Caregivers
A 232-page report of Ireland's first regulator-run national survey of children's online lives, based on in-home face-to-face interviews with 1,002 children aged 9 to 17 and a parent or caregiver of e…
Lab publication
Continuing To Build Upon Our Safety Priorities
A first-party safety update from Character.AI describing safeguards in operation on its platform as of September 2026. It states that self-harm safeguards consider the surrounding conversation, inclu…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Government report
Through Children's Eyes: How Digital Technologies Enable Child Sexual Abuse
Synthesis of two rounds of the Disrupting Harm surveys, nationally representative household surveys of around 21,000 internet-using children aged 12 to 17 across 21 countries in Africa, Asia, Latin A…
Peer-reviewed
Navigating self-injury in a digital world: adolescents' perspectives on coping, help-seeking, and technology
Semi-structured interviews with 21 adolescents who have lived experience of non-suicidal self-injury, about how they cope, what stops them seeking help, and what role digital tools including AI play.…
Peer-reviewed
The interactive turn in generative AI for self-harm
Conceptual paper arguing that the dominant classifier paradigm for AI and self-harm (ingest signals, output a risk label) cannot capture the functional heterogeneity of non-suicidal self-injury, and…
Lab publication
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
Lab publication
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Lab publication
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Preprint
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Lab publication
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Lab publication
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…