arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04281cs.CVcs.AIcs.LG

当视觉凌驾于认知之上:用于视觉语言模型个性化安全的视觉主导性与基于 deferral 的方法

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLMs在高风险场景中对特定用户存在个性化安全风险的问题,构建基准MPS-Bench,揭示其视觉主导机制,提出PRISM方法,实现0.978 AUC并优化安全-效用帕累托前沿。

中文摘要 AI 辅助

视觉语言模型(VLMs)正越来越多地被部署到高风险场景中,在这些场景里,一个在通用层面合理的响应,对于模型未知其医疗、情感或情境背景的特定用户而言,仍可能是不安全的。我们研究多模态系统中的个性化安全问题,并推出MPS-Bench——一个包含584张真实世界图像、12个高风险领域的5181个场景的基准,每个场景都配有隐藏的用户画像。对8个前沿VLMs的评估显示,它们几乎总是直接响应(占比86%-99%),而非寻求缺失的上下文,且在个性化安全指标上,没有一个模型超过2.6/5。为理解这些失败的原因,我们分析多模态交互并识别出视觉主导性:视觉信息在早期进入文本表征,并在多模态融合过程中抑制文本风险信号。因果干预揭示了一个两阶段机制:视觉影响首先在早期层被传递到文本流,随后通过这种被改变的文本表征塑造最终决策,使得后期内部补救不可靠。受此机制启发,我们提出PRISM——一个轻量级输入监控器,它使用双向跨模态调制来预测查询何时可能需要弃权(不执行)。PRISM的AUC达到0.978,且在所有测试模型中严格优于安全-效用帕累托前沿。

英文摘要

Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reasonable in general may still be unsafe for a particular user whose medical, emotional, or situational context is unknown to the model. We study this problem of personalized safety in multimodal systems and introduce MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile. Evaluating eight frontier VLMs, we find that they almost always respond directly (86-99%) rather than seek missing context, and none exceeds 2.6/5 on personalized safety. To understand why these failures arise, we analyze multimodal interactions and identify visual dominance: visual information enters text representations early and suppresses textual risk signals during multimodal fusion. Causal interventions reveal a two-stage mechanism in which visual affect is first transferred into the text stream in early layers and then shapes the final decision through this altered text representation, making late-stage internal remediation unreliable. Motivated by this mechanism, we propose PRISM, a lightweight input monitor that uses bidirectional cross-modal modulation to predict when a query is likely to require deferral. PRISM achieves 0.978 AUC and strictly dominates the safety-utility Pareto frontier across all tested models.

发表机构

  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of Washington(华盛顿大学)
  • Microsoft Research Asia(微软亚洲研究院)
  • William & Mary(威廉玛丽学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑