Disclosure-Gated User Simulation for Companion-Agent Evaluation
A methods paper on a known failure of LLM-simulated users in companion-agent evaluation: the simulated user is too cooperative, so a system can score well by asking many questions rather than by earning the user's willingness to disclose. The authors specify a disclosure gate, a ladder of five ordered gates merged onto three observable depth layers that release information only in response to the agent's behaviour, train a user simulator to it, and show on CompanionBench's English corpus that the gate changes rankings across 12 systems beyond re-seeding noise while leaving per-system scores almost unchanged.
Publisher
arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain)
Published
1 Sept 2026
Added
today
Key Findings
- The disclosure gate is a ladder of five ordered gates merged onto three depth layers; gating behaviour is learned from a synthetic training branch while a real-data branch supplies how people speak and react, so the simulator need not be told at runtime which gate an item sits behind.
- When training no longer states per item which gate applies, the largest rank displacement across 12 systems under test exceeds the noise band from re-running the environment under a new seed, while per-system scores show no detectable change.
- Two acceptance criteria are stated for a simulator (order-preserving rankings and scale-stable absolute scores); of the candidates examined only the released simulator passes both, and its leaderboard correlates 0.993 with the benchmark's original simulator.
- Prompting a frontier model as the simulator barely moves the ranking but shifts every score upward, a change invisible to anyone checking rankings alone.
- The paper supplies the specification, ablations, human studies, negative controls and downstream sensitivity analysis that the original CompanionBench description covered in about four hundred words.
Methodology Notes
Specification and ablation study of a user-simulator component; environment is the English corpus of CompanionBench; 12 systems under test (not named in the abstract). arXiv v1 2026-09-01, cs.CL, cs.AI and cs.HC; no institutional affiliation is printed, hence the preliminary grade. Verified at the arXiv abstract page and the PDF title block on 2026-09-15.
Sources
Authors
Yao Liu, Yu He
Tags
Cite This
APA
Yao Liu, Yu He. (2026). Disclosure-Gated User Simulation for Companion-Agent Evaluation. arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain). https://arxiv.org/abs/2609.00982
Related Insights
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
arXiv (affiliations not stated on the paper; corresponding address is a Tsinghua alumni domain) · 3 Aug 2026
Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
Association for Computational Linguistics (Findings of ACL 2026); Graduate School of Data Science, Seoul National University · 1 Jul 2026
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
arXiv · 3 Jun 2026