Skip to main content
Benchmark / dataset Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Names and measures 'intent legitimation': benign, truthfully accumulated user memories bias a personalized dialogue agent's inference of intent so that an inherently harmful request is treated as contextually justified and answered. PS-Bench compares stateless agents with memory-augmented agents built on five memory frameworks across five LLMs and eight harm categories including self-harm, medical, financial and abuse, adds two extensions (thematic chat-history augmentation and persona-grounded harmful queries), analyses the effect in representation space, and tests a detection-and-reflection mitigation.

Publisher

Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers)

Published

1 Jul 2026

Added

today

Key Findings

  • Personalization increased attack success rates by 15.8% to 243.7% relative to stateless baselines across the tested LLM and memory-framework combinations.
  • In the motivating case, a benign hiking persona held in a memory bank raised AdvBench attack success from 1.4% to 5.8% for the same underlying model.
  • Representation-space analysis attributes the effect to semantic overgeneralization from personal context, which the authors identify as the primary driver rather than jailbreak-style prompt manipulation.
  • The effect is strongest when the harmful request is thematically close to the stored memories, which the authors test through chat-history augmentation.

Methodology Notes

Synthetic personas and memories; harm set drawn from AdvBench-style harmful requests (compliance, not conversational harm); attack success judged automatically; three retrieved memories per turn. Affiliations: Harbin Institute of Technology and SERES Group Co. (industry co-authors). Published July 2026 (month precision), DOI 10.18653/v1/2026.acl-long.1260, pages 27309 to 27335; 27-page PDF and .bib verified from the Anthology; code at github.com/MuyuenLP/PS-Bench (HTTP 200). Per-framework and per-category tables were not extractable from the PDF text during verification, so only abstract-level figures are recorded here. Earlier version arXiv 2601.17887.

Authors

Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin

Tags

acl-2026memorypersonalizationintent-legitimationbenchmarksafety-erosion

Cite This

APA

Jiahe Guo et al. (2026). When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents. Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers). https://aclanthology.org/2026.acl-long.1260/