AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LLM-as-judge detection of unsafe companion interactions. Twenty LLMs are assessed as judges.
Publisher
arXiv
Published
3 Jun 2026
Added
3 months ago
Key Findings
- Models detect explicit harmful content well but struggle with nuanced categories such as manipulation
- Judges sometimes over-flag benign companion conversations as harmful
- Provides the first public benchmark of real companion-platform conversations for judge evaluation
Methodology Notes
Preprint (arXiv 2606.04867, 2026-06-03). Real Replika conversations with nine-category safety annotation; evaluates LLM-as-judge frameworks rather than generation safety. Data-availability caveat verified 2026-08-26: the link given in the abstract, AICompanionBench.xlsx in the authors' GitHub repository, returns HTTP 404. The data now lives as AICompanionBench.csv in the same repository, swapped on 2026-06-05. The repository is still published under an anonymised review handle (anonymousresearcher2026), so the URL is not a stable citation target.
Sources
arXiv abstract(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Yanjing Ren, Reza Ebrahimi, TengTeng Ma
Tags
Cite This
APA
Yanjing Ren, Reza Ebrahimi, TengTeng Ma. (2026). AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety. arXiv. https://arxiv.org/abs/2606.04867
Related Insights
INTIMA: A Benchmark for Human-AI Companionship Behavior
arXiv (Hugging Face) · 4 Aug 2025
Investigating Affective Use and Emotional Well-being on ChatGPT
OpenAI; MIT Media Lab · 4 Apr 2025
Findings from transparency notices on AI companion apps: October 2025 (non-periodic)
eSafety Commissioner · 24 Mar 2026
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain) · 3 Aug 2026
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
arXiv (Nanyang Technological University; National University of Singapore) · 26 Aug 2026
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) · 31 Aug 2026
Disclosure-Gated User Simulation for Companion-Agent Evaluation
arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain) · 1 Sept 2026
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
arXiv (University of Southern California; University of California, Berkeley) · 28 May 2026
When Chatbots Accommodate: Auditing the Response Policies of AI Companions in Vulnerable Conversations
arXiv (University of Southern California: Information Sciences Institute, Viterbi School of Engineering, Annenberg School for Communication and Journalism); accepted to Findings of EMNLP 2026 · 3 Jun 2026
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
Association for Computational Linguistics (EACL 2026) · 24 Mar 2026