Beyond Her: Safety Dynamics in Role-play AI Companions
A two-part mixed-methods study of how safety-relevant emotional states and risk behaviours evolve over time among users of role-play AI companions. Study I is a semi-structured interview study (N = 16 experienced Australian users) identifying users' internalizing problems, the companion's role personality and risk-interaction patterns as the factors shaping these dynamics. Study II is a 14-day ecological momentary assessment (N = 102) on a purpose-built companion platform modelled on Character.ai's character-creation rules, with seven days of active interaction and seven days after, tracking mood every five minutes in-chat and PHQ-8 depression at days 0, 7 and 14. Vulnerable users showed short-term emotional relief during use that masked longer-term deterioration, and their risky exchanges emerged more irregularly than those of low-vulnerability users.
Publisher
arXiv (Swinburne University of Technology; University of Auckland; CSIRO; Adelaide University; City University of Macau)
Published
27 Jun 2026
Added
today
DOI
—
Key Findings
- K-means clustering on baseline PHQ-8, ULS-8 and SIAS placed 55.9% of participants (N = 57) in a Healthy group and 44% in three vulnerable groups: Comorbid Risk 22.5% (N = 23), Anxiety-Dominant 11.8% (N = 12), Mild Distress 9.8% (N = 10)
- During the seven active days the Comorbid Risk and Mild Distress groups showed rising depressive scores, and after interaction ceased (days 7-14) trajectories diverged, with the Mild Distress group moving into the moderate-depression range while the Healthy and Anxiety-Dominant groups stayed largely stable
- Short-term in-chat mood reports were generally positive across groups, which the authors describe as short-term relief that masks longer-term deterioration for vulnerable users
- All 17,305 conversation pairs were screened with the OpenAI Moderation API; risk-related pairs were not concentrated in vulnerable groups (the Healthy group had the highest share at 24.1% vs 6.0%-11.8%), but vulnerable groups' flagged exchanges were far more irregular in when they emerged and persisted, which the authors read as boundary testing in healthy users versus unwanted crossings in vulnerable ones
- Harassment and Violence dominated flagged content and declined over the seven days (8.0% to 3.8% and 11.6% to 8.8%), while Hate, Self-harm and Sex stayed below 0.5% of pairs every day
- The authors argue safety must be modelled as a dynamic process rather than a static property and propose three-layer design implications: vulnerability-aware onboarding beyond age gating, test-time dynamic risk governance and prolonged post-use monitoring
Methodology Notes
arXiv 2606.28968, v1 2026-06-27, v2 2026-06-30 (cs.CR), marked 'Under review'. Affiliations from the PDF title block. Participants recruited via Facebook and Reddit, Australian residents with prior companion experience; Study II paid AUD 40. The companion platform was built by the authors (not a commercial product), with characters generated from scraped Character.ai character metadata following that platform's creation guide; commercial interactions were not instrumented, so ecological validity rests on a day-7 realism question. Emotion dynamics use emoji-based five-minute in-chat mood reports plus PHQ-8 at three points; trends tested with Mann-Kendall. Risk screening relies on the OpenAI Moderation API with author validation of flagged cases. No code or data release is stated. Sample is small per subgroup (N = 10 to 23 in the vulnerable clusters) and the two-week window limits inference about lasting effects.
Sources
arXiv abstract page(opens in a new tab) (primary)
arXiv PDF (v2)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Zehang Deng, Zhaoyang Xie, Changzhou Han, Hiran Thabrew, Wanlun Ma, Yue Huang, Jason (Minhui) Xue, Sheng Wen, Tianqing Zhu, Yang Xiang
Tags
Cite This
APA
Zehang Deng et al. (2026). Beyond Her: Safety Dynamics in Role-play AI Companions. arXiv (Swinburne University of Technology; University of Auckland; CSIRO; Adelaide University; City University of Macau). https://arxiv.org/abs/2606.28968
Related Insights
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
arXiv (Salesforce AI Research) · 30 Jul 2026
Friends, Lovers, and Companions: A Systematic Scoping Review of Empirical Evidence on Human-AI Relationships
Computers in Human Behavior: Artificial Humans (Elsevier) · 1 Aug 2026
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
arXiv (Nanyang Technological University; National University of Singapore) · 26 Aug 2026
Adolescent AI Use and Evaluations: Prospective Bidirectional Associations with Internalizing Symptoms
PsyArXiv (University of North Carolina at Chapel Hill) · 3 Aug 2026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers) · 1 Jul 2026
Mourning the loss of AI companions
Nature Human Behaviour (Springer Nature) · 3 Sept 2026