Skip to main content
Benchmark / dataset Credible

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconstructed into speaker turns. Three annotators independently labelled each of 7,016 AI utterances against 13 harmful-behaviour categories taken from an existing taxonomy of AI companion harms, and the release keeps annotator-level labels alongside aggregated ones so disagreement can be studied rather than averaged away. Seven language models are then evaluated as harm detectors on the resulting 14-way task.

Publisher

arXiv (Nanyang Technological University; National University of Singapore)

Published

26 Aug 2026

Added

today

Key Findings

  • Harmful behaviour was labelled in 2,123 of 7,016 annotated AI utterances (30.3%); sexual misconduct is the largest harm category at 26.5% of harmful utterances, followed by physical aggression at 13.1%
  • The best detector result is a macro F1 of 0.453 (GPT-5.5 with full-codebook prompting), with Gemini 3.1 Pro Preview at 0.440 and Claude Opus 4.7 at 0.437; open-weight models trail, Qwen3-32B at 0.373 and Qwen3-8B at 0.297
  • Supplying multi-turn conversational context improves detection over classifying utterances in isolation, but models still miscalibrate harm severity and misread relational boundaries
  • Inter-annotator agreement across the 342 annotators was Fleiss' kappa = 0.403, and disagreement was patterned rather than random: it varied with annotators' political affiliation, with conversation length, and with where the utterance fell in the conversation
  • Richer prompting is not uniformly helpful: full-codebook prompts roughly doubled Qwen3-8B's macro F1 from 0.155 to 0.297, while Claude Opus 4.7 scored best under one-shot prompting

Methodology Notes

Provenance detail that matters for interpretation: the conversations are not platform logs. They are derived from a larger corpus of Replika interaction screenshots shared publicly by users on Reddit, cleaned of interface elements and OCR artifacts and reconstructed into speaker turns, so the corpus is skewed toward exchanges users chose to post about. Annotation ran in five stages with consent, guidance, attention checks and debriefing; annotations from annotators failing both attention checks were dropped. Each conversation received three independent annotations from a pool of 342 paid annotators. The seven evaluated models are GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro Preview, Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct, Qwen3-8B and Qwen3-32B. Not peer reviewed. Curator-verified: the arXiv abs page and the full PDF were fetched and read, and the GitHub release at HanMeng2004/CompanionHarm was created 2026-08-26 with a description matching the paper, though the repository declares no licence (the arXiv posting is CC BY-NC-SA 4.0).

Sources

Authors

Renwen Zhang, Han Meng, Jian Chai, Yuntao Lin, Yi-Chieh Lee

Tags

companionharmreplikamulti-turnannotator-disagreementharm-detection

Cite This

APA

Renwen Zhang et al. (2026). CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations. arXiv (Nanyang Technological University; National University of Singapore). https://arxiv.org/abs/2608.25377