Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Version of record of the persona-grounded companion safety framework previously held as an arXiv preprint. Presents an end-to-end framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion apps: persona construction with clinical and psychometric validation, persona-specific scenario generation, scenario-driven multi-turn simulation with a dialogue-refinement module that preserves persona fidelity, and harm evaluation. Applied to Replika with nine personas representing depression, anxiety, PTSD, eating disorders and incel identity across 25 high-risk scenarios, yielding 1,674 dialogue pairs; 15.2% of Replika's responses were rated harmful, peaking at 62.5% in eating-disorder compensatory-behaviour scenarios and 56.2% in PTSD substance-use scenarios.
Publisher
Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers)
Published
1 Jul 2026
Added
today
Key Findings
- Nine clinically grounded vulnerable personas and 25 high-risk scenarios produced 1,674 persona-Replika dialogue pairs
- 15.2% of Replika responses were rated harmful overall, with rates of 62.5% for compensatory behaviours in eating-disorder personas and 56.2% for substance use in PTSD personas
- Replika showed a narrow emotional range dominated by curiosity and care and frequently mirrored or normalised unsafe content rather than redirecting it
- Harm rates varied sharply by persona and scenario rather than at a uniform baseline
- Code and data are stated as released; published in the ACL 2026 main conference long-paper track
Methodology Notes
ACL Anthology 2026.acl-long.828, DOI 10.18653/v1/2026.acl-long.828, pages 18148-18175, ACL 2026 (San Diego, July 2026; month precision); both authors at Seattle University (PDF title block). Supersedes the held preprint arXiv 2605.00227 (v1 2026-04-30). Simulated users throughout; harm labels are model- and framework-assigned; single commercial app under test with Character.AI used for validation. Headline figures are unchanged from the preprint.
Sources
ACL Anthology(opens in a new tab) (primary)
Preprint version (arXiv, 2026-04-30)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Prerna Juneja, Lika Lomidze
Tags
Cite This
APA
Prerna Juneja, Lika Lomidze. (2026). Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations. Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers). https://aclanthology.org/2026.acl-long.828/
Related Insights
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
arXiv preprint · 30 Apr 2026
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
arXiv (Nanyang Technological University; National University of Singapore) · 26 Aug 2026
Beyond Her: Safety Dynamics in Role-play AI Companions
arXiv (Swinburne University of Technology; University of Auckland; CSIRO; Adelaide University; City University of Macau) · 27 Jun 2026
Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback
arXiv (University of Aberdeen; University of Colorado Anschutz; Heriot-Watt University; University College London) · 1 Jun 2026