Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

Version of record of the persona-grounded companion safety framework previously held as an arXiv preprint. Presents an end-to-end framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion apps: persona construction with clinical and psychometric validation, persona-specific scenario generation, scenario-driven multi-turn simulation with a dialogue-refinement module that preserves persona fidelity, and harm evaluation. Applied to Replika with nine personas representing depression, anxiety, PTSD, eating disorders and incel identity across 25 high-risk scenarios, yielding 1,674 dialogue pairs; 15.2% of Replika's responses were rated harmful, peaking at 62.5% in eating-disorder compensatory-behaviour scenarios and 56.2% in PTSD substance-use scenarios.

Publisher

Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers)

Published

1 Jul 2026

Added

today

Key Findings

  • Nine clinically grounded vulnerable personas and 25 high-risk scenarios produced 1,674 persona-Replika dialogue pairs
  • 15.2% of Replika responses were rated harmful overall, with rates of 62.5% for compensatory behaviours in eating-disorder personas and 56.2% for substance use in PTSD personas
  • Replika showed a narrow emotional range dominated by curiosity and care and frequently mirrored or normalised unsafe content rather than redirecting it
  • Harm rates varied sharply by persona and scenario rather than at a uniform baseline
  • Code and data are stated as released; published in the ACL 2026 main conference long-paper track

Methodology Notes

ACL Anthology 2026.acl-long.828, DOI 10.18653/v1/2026.acl-long.828, pages 18148-18175, ACL 2026 (San Diego, July 2026; month precision); both authors at Seattle University (PDF title block). Supersedes the held preprint arXiv 2605.00227 (v1 2026-04-30). Simulated users throughout; harm labels are model- and framework-assigned; single commercial app under test with Character.AI used for validation. Headline figures are unchanged from the preprint.

Authors

Prerna Juneja, Lika Lomidze

Tags

replikacharacter-aipersona-safetymulti-turn-simulationacl-2026version-of-record

Cite This

APA

Prerna Juneja, Lika Lomidze. (2026). Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations. Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers). https://aclanthology.org/2026.acl-long.828/