6 artifacts matching
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
An update on our mental health work
A Google blog post announcing changes to Gemini's handling of mental-health-related conversations, including a redesigned 'Help is available' module developed with clinical experts and a new 'one-tou…
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
Thematic analysis of 800 cases of AI-perpetrated sexual conduct identified within 35,105 negative Google Play Store reviews of the Replika companion app. The study characterizes the contextual patter…
ShieldGemma: Generative AI Content Moderation Based on Gemma
Introduces ShieldGemma, a suite of content-moderation models built on Gemma 2 (roughly 2B to 27B parameters) that classify safety risks across four harm types in both user inputs and model outputs. T…
The Ethics of Advanced AI Assistants
A book-length treatment from Google DeepMind of the risks and opportunities of advanced AI assistants, with substantial chapters on anthropomorphism, appropriate human-AI relationships, manipulation…
User Experiences of Social Support From Companion Chatbots in Everyday Contexts: Thematic Analysis
One of the earliest peer-reviewed academic studies of a companion chatbot (Replika) specifically as a source of everyday social support. Combines a qualitative analysis of Google Play Store reviews w…