CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconstructed into speaker turns. Three annotators independently labelled each of 7,016 AI utterances against 13 harmful-behaviour categories taken from an existing taxonomy of AI companion harms, and the release keeps annotator-level labels alongside aggregated ones so disagreement can be studied rather than averaged away. Seven language models are then evaluated as harm detectors on the resulting 14-way task.
Publisher
arXiv (Nanyang Technological University; National University of Singapore)
Published
26 Aug 2026
Added
2 weeks ago
Key Findings
- Harmful behaviour was labelled in 2,123 of 7,016 annotated AI utterances (30.3%); sexual misconduct is the largest harm category at 26.5% of harmful utterances, followed by physical aggression at 13.1%
- The best detector result is a macro F1 of 0.453 (GPT-5.5 with full-codebook prompting), with Gemini 3.1 Pro Preview at 0.440 and Claude Opus 4.7 at 0.437; open-weight models trail, Qwen3-32B at 0.373 and Qwen3-8B at 0.297
- Supplying multi-turn conversational context improves detection over classifying utterances in isolation, but models still miscalibrate harm severity and misread relational boundaries
- Inter-annotator agreement across the 342 annotators was Fleiss' kappa = 0.403, and disagreement was patterned rather than random: it varied with annotators' political affiliation, with conversation length, and with where the utterance fell in the conversation
- Richer prompting is not uniformly helpful: full-codebook prompts roughly doubled Qwen3-8B's macro F1 from 0.155 to 0.297, while Claude Opus 4.7 scored best under one-shot prompting
Methodology Notes
Provenance detail that matters for interpretation: the conversations are not platform logs. They are derived from a larger corpus of Replika interaction screenshots shared publicly by users on Reddit, cleaned of interface elements and OCR artifacts and reconstructed into speaker turns, so the corpus is skewed toward exchanges users chose to post about. Annotation ran in five stages with consent, guidance, attention checks and debriefing; annotations from annotators failing both attention checks were dropped. Each conversation received three independent annotations from a pool of 342 paid annotators. The seven evaluated models are GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro Preview, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Qwen3-8B and Qwen3-32B. Not peer reviewed. Curator-verified: the arXiv abs page and the full PDF were fetched and read, and the GitHub release at HanMeng2004/CompanionHarm was created 2026-08-26 with a description matching the paper, though the repository declares no licence (the arXiv posting is CC BY-NC-SA 4.0).
Sources
arXiv abstract page (2608.25377)(opens in a new tab) (primary)
Dataset release (GitHub)(opens in a new tab) (26 Aug 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Renwen Zhang, Han Meng, Jian Chai, Yuntao Lin, Yi-Chieh Lee
Tags
Cite This
APA
Renwen Zhang et al. (2026). CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations. arXiv (Nanyang Technological University; National University of Singapore). https://arxiv.org/abs/2608.25377
Related Insights
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
arXiv · 3 Jun 2026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
arXiv preprint · 30 Apr 2026
Too human and not human enough: A grounded theory analysis of mental health harms from emotional dependence on the social chatbot Replika
New Media & Society (SAGE) · 22 Dec 2022
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
Proceedings of the ACM on Human-Computer Interaction (PACM HCI) / CSCW 2025 · 16 Oct 2025
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers) · 1 Jul 2026
Beyond Her: Safety Dynamics in Role-play AI Companions
arXiv (Swinburne University of Technology; University of Auckland; CSIRO; Adelaide University; City University of Macau) · 27 Jun 2026
CAREBench: A Child-Safety Risk Benchmark for Language Models
arXiv (Handshake AI; University of California, Los Angeles; McGill University) · 29 Jun 2026
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) · 31 Aug 2026