17 artifacts matching
Preprint
Tailored to you: longitudinal effects of personalising language models
Five-day study in which 992 Prolific participants completed a daily advice-seeking conversation with a language model, randomised to a non-personalised control, a memory-based personalisation conditi…
Preprint
Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority
Audit of Google and Baidu AI Overview systems on 1,920 health queries across 12 countries and four languages, measuring which sources anchor the generated health answers, how source localisation vari…
Benchmark / dataset
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…
Preprint
Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries
An audit of what three free consumer generative-search products (ChatGPT, Perplexity, Google AI Overview) cite when answering mental-health questions. Twenty English questions were run under two prom…
Benchmark / dataset
Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)
Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…
Preprint
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
An adversarial safety evaluation framework for large language models used by K-12 students and teachers. It crosses student- and teacher-facing usage contexts with curriculum topics and a taxonomy of…
NGO report
Google Search: AI Overview & AI Mode — AI Risk Assessment
A product risk assessment of the generative-AI features built into Google Search (AI Overview and AI Mode) as experienced by child and teen accounts, rated Unacceptable Risk against the assessing org…
Preprint
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
Lab publication
An update on our mental health work
A Google blog post announcing changes to Gemini's handling of mental-health-related conversations, including a redesigned 'Help is available' module developed with clinical experts and a new 'one-tou…
Lab publication
Evaluating Language Models for Harmful Manipulation
A framework for evaluating harmful manipulation by language models through context-specific human-AI interaction studies, applied to one model with 10,101 participants across three domains (public po…
Preprint
Examining Risks Through a Characterization of the AI Companion Application Ecosystem: A Stratified Sample from the Apple App Store and Google Play Store
Systematic characterization of AI companion apps on the Apple App Store and Google Play Store. The authors identified 489 unique apps advertising social or relational AI companionship, then ran scree…
Peer-reviewed
The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
Evaluates sycophancy in ten language models from OpenAI, Google and Anthropic under a four-turn escalatory pushback protocol on open-ended diagnostic cases (MedCaseReasoning) and clear-answer biomedi…
Peer-reviewed
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
Thematic analysis of 800 cases of AI-perpetrated sexual conduct identified within 35,105 negative Google Play Store reviews of the Replika companion app. The study characterizes the contextual patter…
NGO report
Gemini with Teen Protections: AI Risk Assessment
Product review of the version of Google Gemini automatically served to accounts aged 13 to 17, rated High Risk. The report characterises it as the adult product with additional content filters and ch…
Lab publication
ShieldGemma: Generative AI Content Moderation Based on Gemma
Introduces ShieldGemma, a suite of content-moderation models built on Gemma 2 (roughly 2B to 27B parameters) that classify safety risks across four harm types in both user inputs and model outputs. T…
Lab publication
The Ethics of Advanced AI Assistants
A book-length treatment from Google DeepMind of the risks and opportunities of advanced AI assistants, with substantial chapters on anthropomorphism, appropriate human-AI relationships, manipulation…
Peer-reviewed
User Experiences of Social Support From Companion Chatbots in Everyday Contexts: Thematic Analysis
One of the earliest peer-reviewed academic studies of a companion chatbot (Replika) specifically as a source of everyday social support. Combines a qualitative analysis of Google Play Store reviews w…