When Chatbots Accommodate: What AI Companions Optimize for in Vulnerable Conversations
The authors argue that existing audits score reactions to pre-defined crisis prompts and so miss the decision policy governing sustained interaction. They introduce a paired taxonomy of user vulnerability and chatbot response for extended companion interactions, then apply Inverse Reinforcement Learning to roughly 48,000 turns of real-world user conversations with GPT-4.1, Character.AI and Replika to infer each platform's response policy. Each platform is found to favour a different response mode, and all three downweight the responses that introduce corrective friction.
Key Findings
- Policies inferred from about 48,000 turns of real-world conversations on three named platforms, rather than from scripted crisis prompts
- GPT-4.1 reaches for advice; Character.AI spreads its response across strategies with no dominant mode; Replika consistently asks questions and stays present
- All three platforms downweight the responses that introduce corrective friction
- GPT-4.1 probes less as conversations continue, and probes less when interacting with psychologically high-risk users
- Replika advises bonded users more and challenges them less
- Character.AI shows no committed engagement strategy on internal distress
- The authors state that the estimated policies are invisible to output-level audits
Methodology Notes
Inverse Reinforcement Learning applied to roughly 48,000 turns of real-world conversations across three platforms, plus a new paired vulnerability-response taxonomy. Conversation provenance, consent basis and sampling are described in the paper body rather than the abstract and were not read at curation time. The inferred policy is a model of observed behaviour, not access to any platform's actual objective, so it is an inference about what each system behaves as if it optimises. Preprint, not peer-reviewed, single version, v1 submitted 2026-06-03, cs.HC. The arXiv abstract renders the second platform name as 'this http URL' because arXiv mangles the domain-shaped string character.ai; the platform is Character.AI.
Sources
arXiv abstract (arXiv:2606.04431) (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Topics
Authors
Minh Duc Chu, Yifan Wu, Zhiyi Chen, Angel Hsing-Chi Hwang, Luca Luceri
Tags
Cite This
APA
Minh Duc Chu et al. (2026). When Chatbots Accommodate: What AI Companions Optimize for in Vulnerable Conversations. arXiv (preprint). https://arxiv.org/abs/2606.04431
Related Insights
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
arXiv preprint · 11 Aug 2026
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
arXiv · 2 Jan 2026
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
arXiv · 3 Jun 2026
Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design
Center for Democracy & Technology (CDT Research) · 29 May 2026
AI emotional support is better only when chosen, but shifts preferences even when it is not
arXiv (preprint) · 24 Aug 2026