158 artifacts matching
Peer-reviewed
Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…
Industry survey
Generatieve AI en illegaal online gokaanbod: Hoe AI-tools Nederlandse consumenten bij niet-vergunde casino's brengen
Dutch-language test of ten consumer generative AI tools on whether simple, realistic questions lead users to online casinos without a Dutch licence. Each tool received nine prompts in three series: n…
Lab publication
System Card: Claude Sonnet 5.5
System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…
Benchmark / dataset
Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark
Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross…
NGO report
Artificially Assisted Gun Violence: Chatbot Risks and Preventative Steps for Responsible AI Companies
White paper reviewing firearm-related harms linked to consumer chatbot use, including firearm suicide, attack planning, and illegal acquisition or modification of guns. It audits the published polici…
NGO report
Social network e chatbot: uno studio dei rischi per utenti vulnerabili e minori
Report of the CDCR (Cittadino Digitale Critico e Responsabile) project, funded by the Italian Ministry of Enterprises and Made in Italy. Part one is a legal analysis of how Facebook, Instagram, TikTo…
Benchmark / dataset
Evaluating AI Safety in Teen Conversations
Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Preprint
Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness
Introduces a browser extension that flags concerning chatbot behaviour (overconfidence, sycophancy, anthropomorphism, persuasive influence and related classes) inline in ChatGPT and Claude conversati…
Framework
Preparing AI chatbots to respond to patient distress and suicidality in high-risk healthcare settings
Comment describing the suicide-risk and distress safety architecture built for 'Suzy', a generative AI chatbot offering recovery, wellness and local-resource support to adults receiving medication tr…
Lab publication
Model Card: Grok 4.7
Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…
Preprint
Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity
A preregistered audit of six deployed assistants (Claude Opus 4.8, GPT-5.5, Grok 4.3, Gemma 4 31B IT, Mistral Small 3.2 24B, DeepSeek V4 Flash) using 7,500 scripted multi-turn conversations that rand…
Preprint
AI-based detection of suicidal ideation in text: model development and evaluation for a student mental health chatbot
Development and evaluation of a lightweight suicidal-ideation detection system intended for integration into Minder, a University of British Columbia mental-health chatbot for students. A fine-tuned…
Preprint
The Adaptation Dilemma: Cultural Fit Does Not Guarantee Safety in Mental-Health LLMs
Conceptual paper arguing that cultural fit and safety are distinct properties of mental-health conversations with general-purpose chatbots. It sorts interaction harms into two classes, imposition (th…
Framework
网络安全标准实践指南——人工智能应用安全指引 教育 (TC260-PG-20269A) [Practice Guide for Cybersecurity Standards: Security Guidelines for Artificial Intelligence Applications: Education]
Sector practice guide (23 pages) released together with the general AI Application Security Guidelines and companion guides for health and for broadcasting and online audiovisual services. It sets ge…
Benchmark / dataset
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper ev…
Preprint
When AI Says "I Am Unable to Answer": Understanding User Responses to AI Refusals
Experiment on how users respond to refusal-based safeguards over repeated interactions. 599 participants interacted with an AI system that never refused, refused infrequently, or refused frequently,…
NGO report
Perplexity: AI Risk Assessment
Same-rubric product review of Perplexity's AI answer engine against the Institute's eight AI principles and its severe-harm red lines, using test accounts registered as a 15-year-old. Perplexity rece…
Peer-reviewed
Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study
Prospective, conversation-level evaluation of jAImes, a retrieval-augmented Parkinson's disease information chatbot commissioned by Parkinson Stiftung Deutschland and deployed publicly in Germany, ac…
Framework
Safe Participation Framework: Opportunity and Safety for the Next Generation in the Age of AI
Microsoft publishes a three-pillar corporate framework for young people's safety across its AI and online products, stating it intends the framework to inform emerging regulatory frameworks. The pill…