11 artifacts matching
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Presents an end-to-end simulation framework for evaluating AI companion app safety across multi-turn conversations, using nine clinically-grounded vulnerable personas (including major depressive diso…
AI-Facilitated Coercive Control: An Experimental Study
Constructs four speculative scenarios combining known coercive-control tactics with conversational-AI capabilities, then probes ChatGPT and Gemini against them. Finds that while the tools refuse blun…
Do No Harm: Exposing Hidden Vulnerabilities of LLMs via Persona-based Client Simulation Attack in Psychological Counseling
Proposes PCSA (Persona-based Client Simulation Attack), a red-teaming framework that simulates coherent, persona-driven counselling clients to probe LLM safety alignment. Across seven LLMs it elicite…
Findings from transparency notices on AI companion apps: October 2025 (non-periodic)
Australia's eSafety Commissioner reports findings from Basic Online Safety Expectations transparency notices issued on 16 October 2025 to four AI companion providers — Chai Research Corp., Character…
Frontier AI Trends Report
The UK AI Security Institute's inaugural Frontier AI Trends Report synthesises two years of evaluations of more than 30 frontier AI systems since November 2023, spanning agent capabilities, chem-bio…
AI Chatbots for Mental Health Support (AI Risk Assessment)
A risk assessment by Common Sense Media's Youth AI Safety Institute, conducted with Stanford Medicine's Brainstorm Lab for Mental Health Innovation, evaluating ChatGPT, Claude, Gemini, and Meta AI as…
General-Purpose AI Code of Practice (EU AI Act, Articles 53 and 55)
Voluntary code of practice published 10 July 2025, drafted by 13 independent experts through a multi-stakeholder process (1,000+ participants) facilitated by the EU AI Office, to help providers of ge…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
The technical paper introducing AILuminate v1.0, an industry-standard AI risk and reliability benchmark developed by MLCommons through an open multi-stakeholder process spanning industry, academia, a…
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Presents WildGuard, an open moderation tool for large language models that jointly detects harmful intent in prompts, safety risks in responses, and model refusal. It is released with WildGuardMix, a…
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Introduces an instruction-hierarchy training method that teaches LLMs to prioritize system/developer-level instructions over conflicting instructions embedded in untrusted user or third-party text. T…
TRAP-18 indicators validated through the forensic linguistic analysis of targeted violence manifestos
Analyses 30 written and spoken manifestos authored by lone offenders who planned or committed targeted attacks (1974-2021), testing whether the behavior-based TRAP-18 threat-assessment instrument can…