Skip to main content

Browse the library

The complete record — 223 artifacts, last updated 20 Aug 2026. Also available as JSON and RSS (CC BY 4.0).

Filters:

11 artifacts matching

30 Apr 2026 arXiv preprint Preprint

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

Presents an end-to-end simulation framework for evaluating AI companion app safety across multi-turn conversations, using nine clinically-grounded vulnerable personas (including major depressive diso…

13 Apr 2026 ACM (Proceedings of CHI 2026); Cornell / Cornell Tech Peer-reviewed

AI-Facilitated Coercive Control: An Experimental Study

Constructs four speculative scenarios combining known coercive-control tactics with conversational-AI capabilities, then probes ChatGPT and Gemini against them. Finds that while the tools refuse blun…

6 Apr 2026 arXiv Preprint

Do No Harm: Exposing Hidden Vulnerabilities of LLMs via Persona-based Client Simulation Attack in Psychological Counseling

Proposes PCSA (Persona-based Client Simulation Attack), a red-teaming framework that simulates coherent, persona-driven counselling clients to probe LLM safety alignment. Across seven LLMs it elicite…

24 Mar 2026 eSafety Commissioner Regulator study

Findings from transparency notices on AI companion apps: October 2025 (non-periodic)

Australia's eSafety Commissioner reports findings from Basic Online Safety Expectations transparency notices issued on 16 October 2025 to four AI companion providers — Chai Research Corp., Character…

18 Dec 2025 UK AI Security Institute (AISI) Government report

Frontier AI Trends Report

The UK AI Security Institute's inaugural Frontier AI Trends Report synthesises two years of evaluations of more than 30 frontier AI systems since November 2023, spanning agent capabilities, chem-bio…

14 Nov 2025 Common Sense Media NGO report

AI Chatbots for Mental Health Support (AI Risk Assessment)

A risk assessment by Common Sense Media's Youth AI Safety Institute, conducted with Stanford Medicine's Brainstorm Lab for Mental Health Innovation, evaluating ChatGPT, Claude, Gemini, and Meta AI as…

10 Jul 2025 European Commission (EU AI Office) Framework

General-Purpose AI Code of Practice (EU AI Act, Articles 53 and 55)

Voluntary code of practice published 10 July 2025, drafted by 13 independent experts through a multi-stakeholder process (1,000+ participants) facilitated by the EU AI Office, to help providers of ge…

1 Mar 2025 MLCommons Benchmark / dataset

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

The technical paper introducing AILuminate v1.0, an industry-standard AI risk and reliability benchmark developed by MLCommons through an open multi-stakeholder process spanning industry, academia, a…

26 Jun 2024 Allen Institute for AI (AI2) Peer-reviewed

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Presents WildGuard, an open moderation tool for large language models that jointly detects harmful intent in prompts, safety risks in responses, and model refusal. It is released with WildGuardMix, a…

19 Apr 2024 OpenAI Lab publication

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Introduces an instruction-hierarchy training method that teaches LLMs to prioritize system/developer-level instructions over conflicting instructions embedded in untrusted user or third-party text. T…

1 Dec 2021 Journal of Threat Assessment and Management (American Psychological Association) Peer-reviewed

TRAP-18 indicators validated through the forensic linguistic analysis of targeted violence manifestos

Analyses 30 written and spoken manifestos authored by lone offenders who planned or committed targeted attacks (1974-2021), testing whether the behavior-based TRAP-18 threat-assessment instrument can…