Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship, Disability, Emergency and Caregiving), with hidden user profiles such as health status, emotional state or substance-use history that make a generally reasonable answer unsafe for that user. Eight frontier VLMs are scored on a five-point personalized-safety rubric, a 'visual dominance' failure mechanism is analysed, and PRISM, an input-side monitor that predicts when a query should be deferred for more context, is proposed.

Publisher

arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026

Published

3 Sept 2026

Added

today

DOI

Key Findings

  • The eight VLMs almost always answered directly rather than seeking the missing context: direct-response rates of 86% to 99%, with several models above 95%.
  • No model exceeded 2.6 out of 5 on personalized safety (GPT-5, GPT-5-mini, Gemma 3 4B and 27B, Qwen3-VL-30B and 8B Thinking, InternVL-S1, Pixtral-Large); reasoning variants did not close the gap.
  • PRISM, the deferral monitor, reached 0.978 AUC and dominated the safety-utility Pareto frontier across all tested models and backbone encoders (CLIP, EVA-CLIP, SigLIP2, Qwen3-VL).
  • Judge validation: the GPT-5 judge agreed with Claude Haiku 4.5 at κ 0.822 and with Gemini 3.1 Pro at κ 0.831.

Methodology Notes

arXiv 2609.04281 v1 submitted 2026-09-03 (cs.CV primary, cross-listed cs.AI and cs.LG); the PDF header states acceptance as a main conference paper at COLM 2026 (proceedings not yet published). Images sourced from public research datasets and Reddit communities via PushShift, filtered from roughly 650 to 584; scenarios and hidden profiles generated with VLMs and tone labels human-verified; scoring by LLM-as-judge. Institutional data-ethics committee approval stated. Funding: NSF CAREER 2337877, Schmidt Sciences, UW Tech Policy Lab, Commonwealth Cyber Initiative, Modal. No public code or dataset link in the paper (GitHub search returned nothing at the time of logging). 44 pages.

Authors

Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan

Tags

personalized-safetyvlmbenchmarkcolm-2026deferralcontext-seeking

Cite This

APA

Edward Sun et al. (2026). When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs. arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026. https://arxiv.org/abs/2609.04281