当视觉凌驾于认知之上:用于视觉语言模型个性化安全的视觉主导性与基于 deferral 的方法
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
浏览论文内容
中文总结 AI 辅助
针对VLMs在高风险场景中对特定用户存在个性化安全风险的问题,构建基准MPS-Bench,揭示其视觉主导机制,提出PRISM方法,实现0.978 AUC并优化安全-效用帕累托前沿。
中文摘要 AI 辅助
视觉语言模型(VLMs)正越来越多地被部署到高风险场景中,在这些场景里,一个在通用层面合理的响应,对于模型未知其医疗、情感或情境背景的特定用户而言,仍可能是不安全的。我们研究多模态系统中的个性化安全问题,并推出MPS-Bench——一个包含584张真实世界图像、12个高风险领域的5181个场景的基准,每个场景都配有隐藏的用户画像。对8个前沿VLMs的评估显示,它们几乎总是直接响应(占比86%-99%),而非寻求缺失的上下文,且在个性化安全指标上,没有一个模型超过2.6/5。为理解这些失败的原因,我们分析多模态交互并识别出视觉主导性:视觉信息在早期进入文本表征,并在多模态融合过程中抑制文本风险信号。因果干预揭示了一个两阶段机制:视觉影响首先在早期层被传递到文本流,随后通过这种被改变的文本表征塑造最终决策,使得后期内部补救不可靠。受此机制启发,我们提出PRISM——一个轻量级输入监控器,它使用双向跨模态调制来预测查询何时可能需要弃权(不执行)。PRISM的AUC达到0.978,且在所有测试模型中严格优于安全-效用帕累托前沿。
英文摘要
Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reasonable in general may still be unsafe for a particular user whose medical, emotional, or situational context is unknown to the model. We study this problem of personalized safety in multimodal systems and introduce MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile. Evaluating eight frontier VLMs, we find that they almost always respond directly (86-99%) rather than seek missing context, and none exceeds 2.6/5 on personalized safety. To understand why these failures arise, we analyze multimodal interactions and identify visual dominance: visual information enters text representations early and suppresses textual risk signals during multimodal fusion. Causal interventions reveal a two-stage mechanism in which visual affect is first transferred into the text stream in early layers and then shapes the final decision through this altered text representation, making late-stage internal remediation unreliable. Motivated by this mechanism, we propose PRISM, a lightweight input monitor that uses bidirectional cross-modal modulation to predict when a query is likely to require deferral. PRISM achieves 0.978 AUC and strictly dominates the safety-utility Pareto frontier across all tested models.
发表机构
- University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
- University of Washington(华盛顿大学)
- Microsoft Research Asia(微软亚洲研究院)
- William & Mary(威廉玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。