Skip to main content
Preprint Preliminary

When Chatbots Accommodate: What AI Companions Optimize for in Vulnerable Conversations

The authors argue that existing audits score reactions to pre-defined crisis prompts and so miss the decision policy governing sustained interaction. They introduce a paired taxonomy of user vulnerability and chatbot response for extended companion interactions, then apply Inverse Reinforcement Learning to roughly 48,000 turns of real-world user conversations with GPT-4.1, Character.AI and Replika to infer each platform's response policy. Each platform is found to favour a different response mode, and all three downweight the responses that introduce corrective friction.

Publisher

arXiv (preprint)

Published

3 Jun 2026

Added

today

Key Findings

  • Policies inferred from about 48,000 turns of real-world conversations on three named platforms, rather than from scripted crisis prompts
  • GPT-4.1 reaches for advice; Character.AI spreads its response across strategies with no dominant mode; Replika consistently asks questions and stays present
  • All three platforms downweight the responses that introduce corrective friction
  • GPT-4.1 probes less as conversations continue, and probes less when interacting with psychologically high-risk users
  • Replika advises bonded users more and challenges them less
  • Character.AI shows no committed engagement strategy on internal distress
  • The authors state that the estimated policies are invisible to output-level audits

Methodology Notes

Inverse Reinforcement Learning applied to roughly 48,000 turns of real-world conversations across three platforms, plus a new paired vulnerability-response taxonomy. Conversation provenance, consent basis and sampling are described in the paper body rather than the abstract and were not read at curation time. The inferred policy is a model of observed behaviour, not access to any platform's actual objective, so it is an inference about what each system behaves as if it optimises. Preprint, not peer-reviewed, single version, v1 submitted 2026-06-03, cs.HC. The arXiv abstract renders the second platform name as 'this http URL' because arXiv mangles the domain-shaped string character.ai; the platform is Character.AI.

Sources

Authors

Minh Duc Chu, Yifan Wu, Zhiyi Chen, Angel Hsing-Chi Hwang, Luca Luceri

Tags

arxivinverse-reinforcement-learningcompanion-platformspolicy-auditreplikacharacter-ai

Cite This

APA

Minh Duc Chu et al. (2026). When Chatbots Accommodate: What AI Companions Optimize for in Vulnerable Conversations. arXiv (preprint). https://arxiv.org/abs/2606.04431