CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
An interactive bilingual (Chinese and English) benchmark for AI emotional companionship that grounds both its scenarios and a trained user simulator in de-identified real-world companion data. A hidden disclosure gate branches each persona's trajectory on the agent's own behaviour, ten relational capabilities are derived from 25 psychology and counselling theories, and a cross-family judge panel with an item-response-theory model separates agent quality from judge severity. Evaluating 28 agents, the authors find emotion regulation and calibrated challenge are common weaknesses, role-play agents rank near the bottom, and the dominant failure is surface warmth substituting for substantive support.
Publisher
arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain)
Published
3 Aug 2026
Added
today
Key Findings
- Ten capabilities are graded, four of them not explicitly scored by prior companion benchmarks: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge; holding ambiguity discriminates most between agents.
- Agents are scored on two axes, a subjective ten-capability rubric and a deterministic measure of whether deeper user disclosure was earned through the hidden gate.
- Rankings are reproducible across languages (Spearman rho 0.996 Chinese, 0.953 English); a cross-family judge panel dilutes same-family favouritism and an IRT model separates agent quality from judge severity.
- Across 28 agents, role-play agents rank near the bottom (immersion does not imply relational competence) and the dominant failure mode is substituting surface warmth for substantive relational support.
- The authors state they will release 500 Chinese-English parallel pairs and the evaluation code; the linked GitHub repository was still empty as of 2026-09-15.
Methodology Notes
Interactive benchmark with a trained user simulator, real-data-derived personas and scenarios, ten-capability rubric, cross-family LLM judge panel and IRT scoring; 28 agents evaluated (not named in the abstract). arXiv v1 2026-08-03, v2 2026-08-05 (33 pages, 19 tables, 13 appendices per the comments field). No institutional affiliation is printed on the paper and the data and code release is still pending, which is why credibility is graded preliminary. Verified at the arXiv abstract page and the PDF title block on 2026-09-15.
Authors
Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan
Tags
Cite This
APA
Yao Liu et al. (2026). CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship. arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain). https://arxiv.org/abs/2608.02046