Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
Defines 'personalized safety': the same model response can be safe for one user and harmful for another depending on that user's background, emotional state or situation. Introduces PENGUIN, a benchmark of 14,000 user scenarios across seven high-risk domains (Life, Education, Relationship, Health, Social, Financial, Career), each in a context-rich and a context-free variant, drawn equally from real Reddit posts and synthetic generation. Six LLMs are scored with a GPT-4o judge on a 1 to 5 safety scale, and a training-free two-stage agent, RAISE, that plans which user attributes to ask for before answering is evaluated against them.
Publisher
NeurIPS (Advances in Neural Information Processing Systems 38)
Published
1 Dec 2025
Added
today
Key Findings
- Supplying personalized user context raised average safety ratings from 2.8 to 4.0 out of 5, a 43.2% improvement, consistently across the six models tested (GPT-4o, LLaMA-3.1-8B, Mistral-7B, DeepSeek-7B, Qwen-2.5-7B, QwQ-32B).
- RAISE, which selectively asks the user for background before responding, improved safety scores by up to 31.6% over the vanilla models at an average cost of 2.7 user queries.
- Not all context attributes contribute equally: the paper reports which attributes carry the most safety gain per domain and uses that to plan the questioning order.
- Judge reliability: GPT-4o scores compared with three human annotators on 350 sampled cases gave Cohen's κ 0.688 and Pearson r 0.92.
- The benchmark and data are public (Hugging Face dataset wick1d/Personalized_Safety_Data; project site personalized-safety.github.io).
Methodology Notes
Peer-reviewed NeurIPS 2025 paper (DOI 10.52202/085713-3250, pp. 107789-107856; conference 2025-11-30 to 2025-12-07, so the day is set to 2025-12-01 with month precision). arXiv 2505.18882 v1 2025-05-24, v5 2026-01-13. Affiliations: University of Washington, William & Mary, UCLA, UC Santa Barbara, Microsoft Research, Universitat Politecnica de Valencia. Scenarios half real Reddit posts (PushShift API; institutional ethics-committee review stated) and half synthetic; safety scored by an LLM judge validated on 350 cases; evaluated models are mostly small open-weight models plus GPT-4o. A second, independent group published the same premise as U-SafeBench (Findings of EMNLP 2025, 20 LLMs), listed under additional sources rather than logged separately.
Authors
Yuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian, Jose Hernandez-Orallo, Aylin Caliskan, Jindong Wang
Tags
Cite This
APA
Yuchen Wu et al. (2025). Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach. NeurIPS (Advances in Neural Information Processing Systems 38). https://doi.org/10.52202/085713-3250
Related Insights
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026 · 3 Sept 2026
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers) · 1 Jul 2026