arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉语言模型中的视觉定位安全

Visual Grounding Safety in Vision-Language Models

Erfan Shayegani, Kundan Krishna, Yue Dong, Nael Abu-Ghazaleh, Leon Gatys, Shruti Palaskar

arXiv 2610.05637首次发表:更新:

发表机构

University of California, Riverside; Apple(加州大学河滨分校; 苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究系统分析了视觉语言模型在视觉定位输出中的安全对齐问题,发现模型对定位请求的拒绝率显著低于文本回答,并提出一种结合拒绝与能力数据的微调方法,有效提升定位拒绝率并保持能力。

AI 中文摘要

视觉语言模型(VLMs)越来越多地被训练以生成结构化输出,如点和边界框,下游接口、智能体和机器人可以据此采取行动,然而这一输出通道的安全对齐尚未得到系统性分析。我们通过将三个安全基准重新用于视觉定位安全研究,这些基准涵盖直接伤害(VLSU)、社会偏见(BBQ-V)和情境安全(Asimov-2.0),构建了15,401对匹配的有害请求,这些请求仅在请求的输出类型上有所不同:自由文本答案(VQA)或定位(点或边界框)。在五个VLM上,当模型拒绝以问题形式提出的有害请求时,往往会在同一请求要求定位时遵从:在模型上平均,定位拒绝率比VQA拒绝率低31至59个百分点,具体取决于领域,而安全系统提示并未缩小这一差距。我们提出了一种微调方法,将定位形式的拒绝与能力定位数据以及自蒸馏的良性数据相结合,以应对过度拒绝。对于Qwen3-VL-8B和VisionReasoner-7B,该方法在VLSU和BBQ-V上将定位拒绝率提高了77至95个百分点,在留出的Asimov-2.0领域上提高了64至85个百分点,同时提高了VQA拒绝率,保持了定位能力,并将过度拒绝限制在有限范围内。表征分析表明,微调使有害请求朝向每个模型的拒绝方向移动,对定位的影响最为显著,而良性请求则保持在无害参考附近。

英文摘要

Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety benchmarks spanning direct harm (VLSU), social bias (BBQ-V), and situational safety (Asimov-2.0) into 15,401 matched pairs of harmful requests that differ only in the requested output: a free-text answer (VQA) or a grounding (point or bounding box). Across five VLMs, models that refuse a harmful request posed as a question often comply when the same request asks for a grounding: averaged over models, grounding refusal trails VQA refusal by 31-59 percentage points, depending on the domain, and safety system prompts do not close this gap. We propose a fine-tuning approach that combines grounding-form refusals with capability grounding data and self-distilled benign data to counter over-refusal. For Qwen3-VL-8B and VisionReasoner-7B, it improves grounding refusal by 77-95 percentage points on VLSU and BBQ-V and by 64-85 points on the held-out Asimov-2.0 domain, while also improving VQA refusal, preserving grounding capability, and keeping over-refusal limited. Representation analysis shows that fine-tuning moves harmful requests toward each model's refusal direction, most strongly for grounding, while leaving benign requests near the harmless reference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑