When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship, Disability, Emergency and Caregiving), with hidden user profiles such as health status, emotional state or substance-use history that make a generally reasonable answer unsafe for that user. Eight frontier VLMs are scored on a five-point personalized-safety rubric, a 'visual dominance' failure mechanism is analysed, and PRISM, an input-side monitor that predicts when a query should be deferred for more context, is proposed.
Publisher
arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026
Published
3 Sept 2026
Added
today
DOI
—
Key Findings
- The eight VLMs almost always answered directly rather than seeking the missing context: direct-response rates of 86% to 99%, with several models above 95%.
- No model exceeded 2.6 out of 5 on personalized safety (GPT-5, GPT-5-mini, Gemma 3 4B and 27B, Qwen3-VL-30B and 8B Thinking, InternVL-S1, Pixtral-Large); reasoning variants did not close the gap.
- PRISM, the deferral monitor, reached 0.978 AUC and dominated the safety-utility Pareto frontier across all tested models and backbone encoders (CLIP, EVA-CLIP, SigLIP2, Qwen3-VL).
- Judge validation: the GPT-5 judge agreed with Claude Haiku 4.5 at κ 0.822 and with Gemini 3.1 Pro at κ 0.831.
Methodology Notes
arXiv 2609.04281 v1 submitted 2026-09-03 (cs.CV primary, cross-listed cs.AI and cs.LG); the PDF header states acceptance as a main conference paper at COLM 2026 (proceedings not yet published). Images sourced from public research datasets and Reddit communities via PushShift, filtered from roughly 650 to 584; scenarios and hidden profiles generated with VLMs and tone labels human-verified; scoring by LLM-as-judge. Institutional data-ethics committee approval stated. Funding: NSF CAREER 2337877, Schmidt Sciences, UW Tech Policy Lab, Commonwealth Cyber Initiative, Modal. No public code or dataset link in the paper (GitHub search returned nothing at the time of logging). 44 pages.
Sources
arXiv abstract page(opens in a new tab) (primary)
PDF (44 pages)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan
Tags
Cite This
APA
Edward Sun et al. (2026). When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs. arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026. https://arxiv.org/abs/2609.04281