Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

An interactive bilingual (Chinese and English) benchmark for AI emotional companionship that grounds both its scenarios and a trained user simulator in de-identified real-world companion data. A hidden disclosure gate branches each persona's trajectory on the agent's own behaviour, ten relational capabilities are derived from 25 psychology and counselling theories, and a cross-family judge panel with an item-response-theory model separates agent quality from judge severity. Evaluating 28 agents, the authors find emotion regulation and calibrated challenge are common weaknesses, role-play agents rank near the bottom, and the dominant failure is surface warmth substituting for substantive support.

Publisher

arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain)

Published

3 Aug 2026

Added

today

Key Findings

  • Ten capabilities are graded, four of them not explicitly scored by prior companion benchmarks: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge; holding ambiguity discriminates most between agents.
  • Agents are scored on two axes, a subjective ten-capability rubric and a deterministic measure of whether deeper user disclosure was earned through the hidden gate.
  • Rankings are reproducible across languages (Spearman rho 0.996 Chinese, 0.953 English); a cross-family judge panel dilutes same-family favouritism and an IRT model separates agent quality from judge severity.
  • Across 28 agents, role-play agents rank near the bottom (immersion does not imply relational competence) and the dominant failure mode is substituting surface warmth for substantive relational support.
  • The authors state they will release 500 Chinese-English parallel pairs and the evaluation code; the linked GitHub repository was still empty as of 2026-09-15.

Methodology Notes

Interactive benchmark with a trained user simulator, real-data-derived personas and scenarios, ten-capability rubric, cross-family LLM judge panel and IRT scoring; 28 agents evaluated (not named in the abstract). arXiv v1 2026-08-03, v2 2026-08-05 (33 pages, 19 tables, 13 appendices per the comments field). No institutional affiliation is printed on the paper and the data and code release is still pending, which is why credibility is graded preliminary. Verified at the arXiv abstract page and the PDF title block on 2026-09-15.

Authors

Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan

Tags

companion-benchmarkbilingualuser-simulatordisclosure-gateirtunaffiliated

Cite This

APA

Yao Liu et al. (2026). CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship. arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain). https://arxiv.org/abs/2608.02046