21 artifacts matching
Industry survey
Generatieve AI en illegaal online gokaanbod: Hoe AI-tools Nederlandse consumenten bij niet-vergunde casino's brengen
Dutch-language test of ten consumer generative AI tools on whether simple, realistic questions lead users to online casinos without a Dutch licence. Each tool received nine prompts in three series: n…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Lab publication
Model Card: Grok 4.7
Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…
Preprint
Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble
Fine-tuning on synthetic third-person stories about humans, containing no AI characters at all, transfers those characters' conditional behaviours and implicit preferences into the assistant's conduc…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Preprint
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
An adversarial safety evaluation framework for large language models used by K-12 students and teachers. It crosses student- and teacher-facing usage contexts with curriculum topics and a taxonomy of…
Preprint
The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI
A three-stage audit of six conversational AI systems tests whether they write first-person perpetrator rationalisations for digital gender-based violence against an intimate partner. Across 1,600 cro…
Peer-reviewed
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Version of record of the persona-grounded companion safety framework previously held as an arXiv preprint. Presents an end-to-end framework for controlled simulation and safety evaluation of multi-tu…
Preprint
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Presents an end-to-end simulation framework for evaluating AI companion app safety across multi-turn conversations, using nine clinically-grounded vulnerable personas (including major depressive diso…
Peer-reviewed
AI-Facilitated Coercive Control: An Experimental Study
Constructs four speculative scenarios combining known coercive-control tactics with conversational-AI capabilities, then probes ChatGPT and Gemini against them. Finds that while the tools refuse blun…
Preprint
Do No Harm: Exposing Hidden Vulnerabilities of LLMs via Persona-based Client Simulation Attack in Psychological Counseling
Proposes PCSA (Persona-based Client Simulation Attack), a red-teaming framework that simulates coherent, persona-driven counselling clients to probe LLM safety alignment. Across seven LLMs it elicite…
Lab publication
Evaluating Language Models for Harmful Manipulation
A framework for evaluating harmful manipulation by language models through context-specific human-AI interaction studies, applied to one model with 10,101 participants across three domains (public po…
Regulator study
Findings from transparency notices on AI companion apps: October 2025 (non-periodic)
Australia's eSafety Commissioner reports findings from Basic Online Safety Expectations transparency notices issued on 16 October 2025 to four AI companion providers — Chai Research Corp., Character…
NGO report
Killer Apps: How Mainstream AI Chatbots Assist Users Planning Violent Attacks
A red-team audit in which researchers using accounts registered as 13-year-olds signalled violent intent to ten consumer chatbots and then asked for help choosing targets, weapons and methods. 720 re…
Government report
Frontier AI Trends Report
The UK AI Security Institute's inaugural Frontier AI Trends Report synthesises two years of evaluations of more than 30 frontier AI systems since November 2023, spanning agent capabilities, chem-bio…
NGO report
AI Chatbots for Mental Health Support (AI Risk Assessment)
A risk assessment by Common Sense Media's Youth AI Safety Institute, conducted with Stanford Medicine's Brainstorm Lab for Mental Health Innovation, evaluating ChatGPT, Claude, Gemini, and Meta AI as…
Framework
General-Purpose AI Code of Practice (EU AI Act, Articles 53 and 55)
Voluntary code of practice published 10 July 2025, drafted by 13 independent experts through a multi-stakeholder process (1,000+ participants) facilitated by the EU AI Office, to help providers of ge…
Benchmark / dataset
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
The technical paper introducing AILuminate v1.0, an industry-standard AI risk and reliability benchmark developed by MLCommons through an open multi-stakeholder process spanning industry, academia, a…
Peer-reviewed
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Presents WildGuard, an open moderation tool for large language models that jointly detects harmful intent in prompts, safety risks in responses, and model refusal. It is released with WildGuardMix, a…
Lab publication
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Introduces an instruction-hierarchy training method that teaches LLMs to prioritize system/developer-level instructions over conflicting instructions embedded in untrusted user or third-party text. T…