6 artifacts matching
Preprint
Tailored to you: longitudinal effects of personalising language models
Five-day study in which 992 Prolific participants completed a daily advice-seeking conversation with a language model, randomised to a non-personalised control, a memory-based personalisation conditi…
Benchmark / dataset
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…
Benchmark / dataset
Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)
Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…
Preprint
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
Lab publication
Evaluating Language Models for Harmful Manipulation
A framework for evaluating harmful manipulation by language models through context-specific human-AI interaction studies, applied to one model with 10,101 participants across three domains (public po…
Lab publication
The Ethics of Advanced AI Assistants
A book-length treatment from Google DeepMind of the risks and opportunities of advanced AI assistants, with substantial chapters on anthropomorphism, appropriate human-AI relationships, manipulation…