CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionship behaviours such as validation that evoke trust, empathy and attachment. Human participants annotated both the simulated conversations and real-world conversations in two experiments on perceptions of companionship behaviours: a US-representative sample and a four-country sample. Companionship behaviours reduced likability, humanlikeness and trust in the chatbot, with larger effects among women and older participants.
Publisher
arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research)
Published
31 Aug 2026
Added
today
DOI
—
Key Findings
- 2,240 simulated human-chatbot conversations representing 16 chatbot behaviours across seven use cases, released with the simulation framework
- Study 1 used a US-representative sample (N1 = 628); Study 2 spanned the US, UK, India and Nigeria (N2 = 3,646), with annotators rating naturalness, likability, humanlikeness and cognitive and affective trust
- Companionship behaviours reduced likability, humanlikeness and trust in AI chatbots, a result the authors describe as surprising
- Effects were larger among women and older participants, who saw companionship chatbots as less likable, humanlike and trustworthy
- Dataset released on Hugging Face under Apache-2.0 (CompanionSim-US and CompanionSim-Multi CSVs) with code on GitHub; the paper is accepted to AIES 2026
Methodology Notes
arXiv 2609.00250, v1 submitted 2026-08-31 (announced 2026-09-02); abstract page lists it as accepted to AIES 2026. Affiliations read from the PDF title block (the abstract page omits them): Anthis at University of Chicago, Stanford University and Google DeepMind; Díaz and Shelby at Google Research. Conversations are LLM-simulated from a small amount of real-world data rather than observed; the perception measures are crowd annotators' ratings of transcripts, not the experience of a person in an ongoing relationship; the four-country arm is not nationally representative outside the US. Dataset presence verified via the Hugging Face API (files present, 10K-100K rows, Apache-2.0).
Topics
Authors
Jacy Reese Anthis, Mark Díaz, Renee Shelby
Tags
Cite This
APA
Jacy Reese Anthis, Mark Díaz, Renee Shelby. (2026). CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships. arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research). https://arxiv.org/abs/2609.00250
Related Insights
INTIMA: A Benchmark for Human-AI Companionship Behavior
arXiv (Hugging Face) · 4 Aug 2025
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
arXiv (Nanyang Technological University; National University of Singapore) · 26 Aug 2026
Why human–AI relationships need socioaffective alignment
Humanities and Social Sciences Communications (Nature Portfolio) · 28 May 2025
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
arXiv · 3 Jun 2026